Ancestree
2026
Four lines of yours, everything else recorded
Your code writes files at native speed into a scratch directory. At block exit the
artifact is chunked, hashed, compressed and committed in one transaction. A step
that raises keeps its partial output flagged healthy=False, a step
that wrote nothing is discarded with a warning, and a run killed outright is
adopted as an unhealthy node the next time the store opens.
Rules are optional and enforced at creation rather than logged afterwards. Declare rules={"model": ["clean"]} and an illegal transition raises
when it happens.
store = ancestree.LineageStore(
root="./my_project",
rules={"clean": ["ingest"], "model": ["clean"]},
)
with store.create_node(step_type="ingest") as node:
df = do_process()
df.to_csv(node / "raw.csv")
node.add_meta("rows", len(df))
store.find(accuracy=lambda a: a and a > 0.9)
store.lineage(best_model) # full ancestry, oldest first
store.serve_graph() # the explorer, on localhostWhere the time goes
Five things happen to an artifact at block exit. Two of them are the entire budget.
Chunking and compression are the whole budget, and hashing is free. The parameter that matters is average chunk size. At 32 KiB a delta base sits exactly on zlib’s 32,256-byte dictionary window, so it cannot be seen in full and the tail of every delta degenerates to literals. Halving it to 16 KiB stores 14% less and ingests 22% faster.
Three ways of not storing it twice
Each layer catches what the one above it missed. The top one is free, and the bottom one is the only part that costs anything.
A rerun whose step type, parents, metadata and artifact bytes are all identical rebinds onto the node that already exists. Nothing is stored and nothing is inserted.
Artifacts are split into content-defined chunks addressed by SHA-256. A chunk already in the pool is a lookup rather than a write, so copying a dataset into three branches stores it once.
A chunk resembling one already stored is kept as a zlib delta against it, found through the resemblance index. Depth is capped at one, and a delta is only kept when it beats plain compression by a margin.
The same 1% of bytes, edited in four places
Twelve revisions of one 4 MB CSV. Insert, delete and append concentrate the edit at a point, so the boundary algorithm re-syncs within a chunk or two and everything after it hashes identically. That is what content-defined chunking buys and what a fixed-block scheme cannot do. Scattered overwrites leave almost every chunk differing by a byte or two, so the ratio has to be earned by deltas instead.
The headline 3.93× is therefore the ratio for one corpus rather than a
guarantee. The same edit budget swings it from 3.9× to 26× depending on where
the edits land, and store.stats() reports it on your own data.
Every node, and everything recorded about it
A real exported store, laid out by generation and coloured by step type. Click a node for what was recorded when it ran.
This is the static snapshot (store.export_graph()), one
self-contained HTML file. The live explorer adds search, node diffs and a sortable
runs table, all answered by SQL. Provenance is captured on every node without
being asked for: user, Python, platform, git commit, branch, and whether the
worktree was dirty.
Three cost classes, three orders of magnitude
The write path is the only thing that scales with data volume, and it is paid at block exit rather than while your code runs. Everything on the query side is indexed.
| get(node_id) | 0.011 |
|---|---|
| find(run_id=…) selective | 0.020 |
| lineage(node) | 0.103 |
| create_node, metadata only | 8.4 |
|---|---|
| of which git provenance | 7.3 |
| find() over 3,000 nodes | 27.2 |
| ingest 4 MB of CSV | 352 |
|---|---|
| export_metadata(), 3,000 nodes | 515 |
| ingest a 48 MB artifact | 5,009 |
Times in milliseconds, medians of repeated runs after a warm-up. The numbers that do not flatter it are in the same table in the repo: ingest runs at roughly 11 MB/s, so a hundred-megabyte artifact belongs at a step boundary rather than inside a loop.
Guarantees, checks, and what runs on every push
A lineage store holds somebody else’s work, so what matters is what it promises across versions and what it verifies rather than assumes.
- Every store stamps its format version at creation. Open checks it and refuses anything ancestree did not write, without touching the file.
- Ten structural keys are reserved, among them node_id, generation, healthy and content_hash, and add_meta raises rather than shadowing a fact lineage depends on.
- Every chunk is re-hashed against its digest on read and every artifact against its own SHA-256, so silent corruption surfaces as an error rather than as data.
- Partial work is evidence. A step that raises is committed and flagged healthy=False, searchable with find(healthy=False).
- ruff, mypy --strict and pytest across 3.9 to 3.14, with 92.97% line coverage reported to Codecov.
- ruff is pinned to an exact version, because its default rule set widens between releases and would turn CI red without a code change.
Audited against its own documentation
806 tests in 14 modules, written by treating every falsifiable sentence in the docs as an assertion, then adding hostile inputs, seeded property tests, real SIGKILLs, multi-process concurrency and 1,000-node scale runs.
SQLite runs `incremental_vacuum` as a stepped statement and the cursor was never consumed, so CPython finalised it after a single step. No error, and an emptied store needed about 319 calls to release its space. Checkpointing the WAL first and then driving the pragma to exhaustion took 1,417,216 bytes on disk down to 114,688.
Under WAL journalling a live store is three files, and the committed data sat in the one the docs did not name. Copying `ancestree.db` out of an open store produced a valid, openable, empty store. Fixed with `store.backup()` on SQLite’s online backup API, counting the WAL in `stats()`, and correcting the documentation.
The orphan-scratch sweep that runs on every open treated an unseeded directory as litter, and a node was created before it was seeded. The concurrent-write test hit it naturally. Each node is now assembled in a staging directory and renamed into place atomically, after which four processes can write with no coordination.
The value was iterated, so an 8-character node id became a list of eight characters and matched no node, returning an empty result rather than an error. Every other node-accepting method already resolved ids, records and handles interchangeably.
Nothing turned up in the lineage DAG, the metadata envelope, the chunker or either deduplication layer. All four defects were in the operational plumbing, and two of them could lose data. 36 of 40 documented claims held exactly as written when audited, and all 40 hold now. Tests that pinned the broken behaviour were inverted rather than deleted, so a regression fails loudly.
0.1 was a directory per node
It worked, and it carried a hand-rolled index to make it work: a snapshot, a journal, a reconcile pass, a background packer with fork handling, and a GC lock file. 0.2 deletes all of it and lets SQLite be the index.
The break is deliberate and total. There is no migration, and a 0.1 store is refused rather than half-read.
| 0.1 | 0.2 | |
|---|---|---|
| Hand-rolled index, journal and reconcile pass | deleted | SQLite is the index |
| Background packer and its fork handling | deleted | packed at block exit |
| Average chunk size | 32 KiB | 16 KiB, 14% less stored and 22% faster |
| Provenance capture | 3 git subprocesses | 2, concurrent, nodes 2.2× faster |
| Chunker hot loop | 16.6 MB/s | 26.8 MB/s, boundaries identical |
| Selective find over 3,000 nodes | 0.40 ms | 0.04 ms |
What it is not for
A single SQLite file buys simplicity and crash safety. These are the consequences, and they are in the shipped documentation rather than only here.
- Megabyte-scale: a 16 KiB artifact is roughly three times bigger in the store than as a plain file, because a database has pages and indexes whether you use them or not. The two ratios converge by a few megabytes.
- One writer: many readers, one writer, local disk only. SQLite locking over NFS is unreliable, and heavy parallel writing is not what this is for.
- No migrations, by design: a store records its format version and ancestree refuses anything it did not write. To read an old store, keep the version that wrote it. Every version stays on PyPI.
- One file is the whole store: there is no side index to rebuild, so a corrupt database is real data loss. WAL journalling, an integrity check and `meta.json` sidecars are the mitigations.
- Above 64 MiB: the chunker cuts at fixed offsets to keep huge ingests at C speed. Exact dedup survives, shift-resilience does not.