Crypto orderflow
2026
The pipeline
A socket to a research panel, in five processes. Each stage is separate so that none can stall the one in front of it, because the tape has to be written whatever compaction is doing.
The raw tape is about 150× the size of what is kept. Folding it to a per-minute footprint, with volume-at-price and buys and sells separated, holds the shape OHLCV cannot express at all. It costs the individual print sizes and their ordering, which is the one deliberate loss in the archive, and Binance publishes the full tape if it is ever wanted back.
The S3 line is arithmetic on the published rate rather than a bill I pay, since the archive lives on the box. It is the clearest way to say what the fold is worth: a day I keep costs about three pence a year to store, and the same day unfolded would cost four pounds.
What it captures
| Stream | What it is | Rows/day |
|---|---|---|
| footprint | the trade tape, folded to volume-at-price per minute | 10.1M |
| mark | mark / index price and funding, sampled to the minute | 460k |
| oi | open interest, polled every minute | 454k |
| bar | 1-minute klines, as the venue publishes them | 449k |
| liq | forced liquidations, per event | 26k |
| optrade / opt / optoi | option prints, surface and open interest | 40k |
| clock / conn / error | chrony offset, socket lifecycle, every failure | 9k |
Collection is deliberately broad and mechanical. 100 base assets, with every venue’s own spelling of them resolved from its instrument listing at connect time. What to trade is a question you answer by querying the archive afterwards, rather than by filtering what gets written.
Connecting to the public trade feed…
A public Binance feed, read straight from your browser. The archive’s own tape is not served publicly, but this is the same shape of row at the same rate: about 14 GB a day across 583 streams.
Running it is the hard part
Collecting for an hour is a script. Collecting for months, unattended, on hardware that is also doing something else, is the engineering.
- 70 offline tests plus 10 stress scenarios: kill the compactor mid-write, replay a venue, drift the clock.
- Every price, size and OI sits on one 1e-8 integer grid, and the footprint sums back to the venue’s own published minute volume (21.24 against 21.27 BTC).
- Blue/green rollout: the new colour must produce a row on every feed before the old one is stopped.
- Across 18 retained days, 1,064 socket drops and 10,190 DNS failures were absorbed without losing a row. The failure that did stop collection was a different kind, and it has its own section below.
- A second-hand i5-6600T with 4 cores and 8 GB, shared with everything else on the box.
- Capture is one event loop, so it is bounded by one core. The ceiling moves by adding a shard, and collectors are pinned to cores 0 and 1 by cpuset.
What it actually does, measured
Every ten seconds a watcher writes a row of throughput, wire delay, writer lag, memory and disk. This is 149,059 of those rows, covering 23 days of the running box. Percentiles rather than averages, because the p99 decides whether a machine copes.
| p50 | p95 | p99 | ||
|---|---|---|---|---|
| Ingest rate | 635 rows/s | 1,664 | 2,801 | across 583 venue-symbol streams |
| Wire delay | 142 ms | 166 | 237 | venue timestamp to local receipt, clock-corrected |
| Writer lag | 1.0 ms | 1.3 | 1.9 | the socket path never waits on the disk |
| WAL growth | 164 KB/s | 432 | 729 | about 14 GB a day of raw tape |
| Collector RSS | 343 MB | 472 | 477 | both shards, against a 512 MB limit each |
The lag row is the one that matters structurally. Capture writes JSONL and never blocks on the archive, so the p99 wait between a row arriving and being written is 1.9 ms and IO pressure sits at zero. CPU says the same thing from the other side: the two pinned cores run at 28% and 32% while the two the archive cannot touch run at 3%.
The write contract
Eleven streams, one Parquet file per stream per UTC day in Hive-style date partitions at zstd-19. Before any of its rows are allowed in, a venue adapter has to satisfy four rules.
A day of footprint is 7.0M to 15.2M rows and 37 to 116 MB. Partition keys live in the path rather than as columns, so nothing pays to store the stream and the date a second time on every row.
- One time axis
- `ts_exchange` in UTC milliseconds is canonical, taken from the venue’s own event time rather than the receipt time. `recv_delay_ms` carries the difference, which delta-encodes to about two bytes.
- Never floats
- price is an integer count of 1e-8 units on a fixed global scale. Quantity is an integer count of the instrument’s own stepSize, resolved from a registry, so a tick-size change is a registry event rather than an archive migration.
- Normalised enums
- aggressor side, liquidation status and stream names are the archive’s vocabulary, not each venue’s. A venue’s spelling of a base asset is resolved from its instrument listing at connect time.
- Keys, and where there are none
- each stream declares the natural key two collectors would agree on, so the blue/green overlap deduplicates exactly. `liq` declares none, because it carries no id and two liquidations can legitimately share every field, so both copies are kept. A duplicate is recoverable and a dropped row is not.
The six days that are missing
Live capture is the one kind of data engineering with no backfill. The research half can re-download six and a half years from a public archive whenever it likes. This half has exactly what it was listening for at the time.
The box rebooted on 17 August. The chrony sampler that feeds the archive its clock stream runs on the host rather than in the container, because chronyd binds loopback in the host namespace. It is a systemd user unit, and those do not come back after a reboot unless lingering is enabled for the account. It did not come back.
Refused to start. It will not write rows it cannot timestamp against a clock sample newer than 60 seconds, so it failed closed rather than filling the archive with data of unknown time quality. I would choose that again. The cost is a hole, and the alternative is a subtly wrong archive that looks complete.
Nothing was watching the watcher. Prometheus scrapes the host every 15 seconds and holds no alert rules, so the dead stream sat visible on a dashboard nobody was looking at. Enabling lingering on 23 August fixed it, and the sampler has been up since. The alerting gap is named in the limits below rather than quietly closed.
Read back
Every pane is built from captured rows and aggregated across whichever venues are ticked. It is the same view a commercial terminal sells, off my own data.
Below it, the earlier bar-level pipeline this grew out of: five exchanges aggregated in one call, with CVD reconstructed as a continuous series and funding smoothed from the premium index rather than shown as delayed settlement steps. The second image is the pane-by-pane comparison against velo.xyz.
The strategy
A market-neutral cross-sectional book on Binance perpetuals. The signal is exchange-reported taker-buy share, so aggressive flow is measured rather than inferred from a tick rule. Averaged over a week, held for two.
| First book | Re-engineered | |
|---|---|---|
| Net annual | 32.6% | 43.5% |
| Net Sharpe | 1.93 | 2.39 |
| Max drawdown | −21.2% | −16.7% |
| Calmar | 1.54 | 2.60 |
| Turnover | 55×/yr | 27.7×/yr |
| Equity, 6.5y | 7.59× | 15.16× |
The second column is not a risk overlay bolted onto the first. Holding 336h instead of 168h halves turnover while the impulse response is still positive, ensembling the formation window removes a parameter choice, and equal weighting stops the book spending risk on low-volatility names that earn nothing here. Positive in every one of seven calendar years, at +0.08 correlation to BTC.
How it was falsified
Each of these is a way the result could have been an artifact. All were run on the final configuration, not on the version that happened to survive them.
The delayed fill matters most. A strategy whose edge disappears when you fill a day late is measuring bid-ask bounce rather than information, and this one is unchanged. Three more sit behind the table: a deflated Sharpe against 500 trials (p = 0.9998), a block bootstrap with a 90% CI of [1.24, 2.65], and a delisting shock forcing all 144 delisted names to −90% on their final bar, which moved Sharpe from 1.930 to 1.917.
| Shuffle the signal across names | −0.39 | the edge is the signal, not the weighting |
|---|---|---|
| Deliberate lookahead | 8.73 | time alignment is correct |
| Fill a full day late | 2.29 | no dependence on execution speed |
| Survivors only | 2.22 | worse, so no survivorship inflation |
| Inverted signal | −2.74 | mirrors symmetrically |
| Out-of-sample, 2024+ | 4.22 | Calmar, against 2.32 in-sample |
What was killed, and by what
Four hypotheses that looked good enough to build. Each was rejected by a specific measurement rather than abandoned.
Long the cheapest funding, short the dearest. Sharpe 0.65 full-sample, but that is 2020–21 in disguise, and −1.13 since 2025. The long leg collects 188% annualised to hold names falling 238%, which is textbook adverse selection.
Rank IC +3.5 at t = 15, and a dollar long-short spread of t = −0.83. The top and bottom 1% of name-bars own +733% and −519% of the leg’s P&L, so the signal ranks the median correctly while P&L pays the mean.
Gross Sharpe 1.49 full-sample at 6h formation, and entirely 2021. Gross −0.28 across 2024–25.
Vol targeting, a trend overlay, dispersion gating, a de-grossing cap, a drawdown circuit breaker. Five of six lowered Calmar, because they cut drawdown by cutting the strategy. What worked instead was structural: hold twice as long, ensemble the formation window, drop inverse-vol weighting.
The lesson that cost the most time is that rank IC and dollar P&L can disagree in sign. A signal with t = 22 on rank IC produced a portfolio with t = −0.8. Any research process that stops at IC will ship strategies that lose money.
Status and limits
The archive has been live since 2026-08-04. The research is at the end of its validation phase: the book is specified, falsified and packaged for paper trading, and it has no live track record. These are the things I would want asked about it.
- No alerting: Prometheus scrapes every 15 seconds and holds no alert rules, so a stopped stream is visible rather than announced. That is how the August gap ran to six days. The dead-man switch restarts a silently dead container, but nothing pages a human when the host-side dependency is what died.
- One box, no replica: the permanent Parquet is 108 MB a day, and the only off-box copy is one snapshot taken by hand. It is the irreplaceable half, and the next thing to fix.
- The backtest is single-venue: Binance only. The archive is the fix in progress, but it holds weeks of history rather than years, so it cannot backtest anything yet.
- The edge has thinned: 2021 Sharpe 2.99 against 0.99 across 2024–25. The planning number is the walk-forward 1.08, not the headline.
- No live track record: a paper-trading harness exists, with five books from $1k to $1m and no API keys anywhere, but it has not run long enough to mean anything. Impact is modelled with a square-root law rather than measured, so capacity beyond about $50m is an estimate.