Tailslayer — DRAM Hedged Read Library
Skill by ara.so — Daily 2026 Skills collection.
Tailslayer is a C++ library that reduces tail latency in RAM reads caused by DRAM refresh stalls. It replicates data across multiple independent DRAM channels with uncorrelated refresh schedules, issues hedged reads across all replicas simultaneously, and returns whichever result responds first — eliminating worst-case stall spikes from DRAM refresh cycles.
Works on AMD, Intel, and AWS Graviton using undocumented channel scrambling offsets.
How It Works
- Data is replicated N times, each copy placed on a different DRAM channel
- Each replica is monitored by a worker pinned to a separate CPU core
- When a read is triggered (via your signal function), all replicas are read simultaneously
- Whichever channel responds first wins; the result is passed to your work function
- DRAM refresh on one channel cannot stall all channels simultaneously → tail latency is eliminated
Installation
Copy the header into your project
Include in your code
Build the provided example
Key API
tailslayer::HedgedReader<T, SignalFn, WorkFn, SignalArgs, WorkArgs>
Template parameters:
Constructor optional parameters
Methods
Utilities
Minimal Usage Pattern
Passing Arguments to Signal and Work Functions
Use tailslayer::ArgList<...> to pass compile-time integer arguments:
Custom Channel Configuration
Override channel offset, channel bit, and replica count in the constructor:
Note: N-way (more than 2 replicas) hedging requires using the benchmark code in
discovery/benchmark/. The main library header currently exposes 2 channels by default.
Running Benchmarks
Channel-hedged read benchmark (N-way)
Flags:
DRAM refresh spike timing probe
This measures your DRAM's tREFI refresh interval and the worst-case stall duration — useful for calibrating expectations.
Platform Notes
Use --all in the benchmark to auto-detect the best channel bit for your system.
Common Patterns
Low-latency trading / event-driven read
Preloading a lookup table across channels
Troubleshooting
High latency still observed
- Verify you are using the correct
--channel-bitfor your CPU. Run benchmark with--all. - Ensure workers are pinned to isolated cores (use
isolcpus=kernel boot parameter). - Run with real-time scheduling:
sudo chrt -f 99 ./your_binary
Build errors — missing headers
- Confirm
include/tailslayer/hedged_reader.hppis on your include path. - Requires C++17 or later: add
-std=c++17to your compiler flags.
Workers don't start / deadlock
start_workers()is blocking. It launches threads and waits — your signal function must eventually return.- Ensure the signal function does not block indefinitely during testing.
Data corruption / wrong values
- Each
insert()replicates the value N times (one per channel). Logical indexing is handled internally — do not attempt to address replicas directly. - Do not modify inserted data after
insert()is called.
Platform not supported
- Tailslayer uses undocumented DRAM channel scrambling offsets. If your platform is not AMD, Intel, or Graviton, run the trefi_probe and benchmark tools to characterize refresh behavior before using the library in production.


