All projects

Ancestree

2026

interface

Four lines of yours, everything else recorded

Your code writes files at native speed into a scratch directory. At block exit the artifact is chunked, hashed, compressed and committed in one transaction. A step that raises keeps its partial output flagged healthy=False, a step that wrote nothing is discarded with a warning, and a run killed outright is adopted as an unhealthy node the next time the store opens.

Rules are optional and enforced at creation rather than logged afterwards. Declare rules={"model": ["clean"]} and an illegal transition raises when it happens.

store = ancestree.LineageStore(
    root="./my_project",
    rules={"clean": ["ingest"], "model": ["clean"]},
)

with store.create_node(step_type="ingest") as node:
    df = do_process()
    df.to_csv(node / "raw.csv")
    node.add_meta("rows", len(df))

store.find(accuracy=lambda a: a and a > 0.9)
store.lineage(best_model)      # full ancestry, oldest first
store.serve_graph()            # the explorer, on localhost
ingest

Where the time goes

Five things happen to an artifact at block exit. Two of them are the entire budget.

Chunking and compression are the whole budget, and hashing is free. The parameter that matters is average chunk size. At 32 KiB a delta base sits exactly on zlib’s 32,256-byte dictionary window, so it cannot be seen in full and the tail of every delta degenerates to literals. Halving it to 16 KiB stores 14% less and ingests 22% faster.

per 4 MB of CSV
chunk 149.0 ms Gear rolling hash finds content-defined boundaries
compress 147.8 ms zlib-6 on every chunk; a chunk that does not shrink is kept verbatim
fingerprint 17.1 ms super-features, so a delta base can be found without comparing every chunk to every other
hash 1.5 ms SHA-256 of each chunk’s plaintext, which is the content address
delta 0.8 ms zlib with the base chunk as preset dictionary, when it beats plain compression
storage

Three ways of not storing it twice

Each layer catches what the one above it missed. The top one is free, and the bottom one is the only part that costs anything.

Node reuse the whole step

A rerun whose step type, parents, metadata and artifact bytes are all identical rebinds onto the node that already exists. Nothing is stored and nothing is inserted.

136× faster 5.3 ms against a 720 ms cold write
Exact chunk dedup bytes, across artifacts

Artifacts are split into content-defined chunks addressed by SHA-256. A chunk already in the pool is a lookup rather than a write, so copying a dataset into three branches stores it once.

10 copies = 1 ten 8 MB writes cost 8 MB, even on incompressible data
Delta storage near-duplicates

A chunk resembling one already stored is kept as a zlib delta against it, found through the resemblance index. Depth is capped at one, and a delta is only kept when it beats plain compression by a margin.

2.15× less for 1.15× the ingest time, which is why it defaults to on

The same 1% of bytes, edited in four places

Twelve revisions of one 4 MB CSV. Insert, delete and append concentrate the edit at a point, so the boundary algorithm re-syncs within a chunk or two and everything after it hashes identically. That is what content-defined chunking buys and what a fixed-block scheme cannot do. Scattered overwrites leave almost every chunk differing by a byte or two, so the ratio has to be earned by deltas instead.

The headline 3.93× is therefore the ratio for one corpus rather than a guarantee. The same edit budget swings it from 3.9× to 26× depending on where the edits land, and store.stats() reports it on your own data.

insert 23.19×
delete 26.27×
append 25.14×
scattered overwrite 3.94×
explorer

Every node, and everything recorded about it

A real exported store, laid out by generation and coloured by step type. Click a node for what was recorded when it ran.

This is the static snapshot (store.export_graph()), one self-contained HTML file. The live explorer adds search, node diffs and a sortable runs table, all answered by SQL. Provenance is captured on every node without being asked for: user, Python, platform, git commit, branch, and whether the worktree was dirty.

measured

Three cost classes, three orders of magnitude

The write path is the only thing that scales with data volume, and it is paid at block exit rather than while your code runs. Everything on the query side is indexed.

Microseconds put them anywhere
get(node_id)0.011
find(run_id=…) selective0.020
lineage(node)0.103
Milliseconds fine per step, not per row
create_node, metadata only8.4
of which git provenance7.3
find() over 3,000 nodes27.2
Scales with your data budget for it
ingest 4 MB of CSV352
export_metadata(), 3,000 nodes515
ingest a 48 MB artifact5,009

Times in milliseconds, medians of repeated runs after a warm-up. The numbers that do not flatter it are in the same table in the repo: ingest runs at roughly 11 MB/s, so a hundred-megabyte artifact belongs at a step boundary rather than inside a loop.

shipping

Guarantees, checks, and what runs on every push

A lineage store holds somebody else’s work, so what matters is what it promises across versions and what it verifies rather than assumes.

The contract is the format version
  • Every store stamps its format version at creation. Open checks it and refuses anything ancestree did not write, without touching the file.
  • Ten structural keys are reserved, among them node_id, generation, healthy and content_hash, and add_meta raises rather than shadowing a fact lineage depends on.
Quality is checked on the way out
  • Every chunk is re-hashed against its digest on read and every artifact against its own SHA-256, so silent corruption surfaces as an error rather than as data.
  • Partial work is evidence. A step that raises is committed and flagged healthy=False, searchable with find(healthy=False).
Every push, six Pythons
  • ruff, mypy --strict and pytest across 3.9 to 3.14, with 92.97% line coverage reported to Codecov.
  • ruff is pinned to an exact version, because its default rule set widens between releases and would turn CI red without a code change.
audit

Audited against its own documentation

806 tests in 14 modules, written by treating every falsifiable sentence in the docs as an assertion, then adding hostile inputs, seeded property tests, real SIGKILLs, multi-process concurrency and 1,000-node scale runs.

High compact() reclaimed one 4 KiB page per call fixed

SQLite runs `incremental_vacuum` as a stepped statement and the cursor was never consumed, so CPython finalised it after a single step. No error, and an emptied store needed about 319 calls to release its space. Checkpointing the WAL first and then driving the pragma to exhaustion took 1,417,216 bytes on disk down to 114,688.

High “Back it up by copying one file” silently backed up nothing fixed

Under WAL journalling a live store is three files, and the committed data sat in the one the docs did not name. Copying `ancestree.db` out of an open store produced a valid, openable, empty store. Fixed with `store.backup()` on SQLite’s online backup API, counting the WAL in `stats()`, and correcting the documentation.

Medium Opening a store could delete another session’s in-flight node fixed

The orphan-scratch sweep that runs on every open treated an unseeded directory as litter, and a node was created before it was seeded. The concurrent-write test hit it naturally. Each node is now assembled in a staging directory and renamed into place atomically, after which four processes can write with no coordination.

Low find(parent_id=…) matched nothing for a bare id fixed

The value was iterated, so an 8-character node id became a list of eight characters and matched no node, returning an empty result rather than an error. Every other node-accepting method already resolved ids, records and handles interchangeably.

Nothing turned up in the lineage DAG, the metadata envelope, the chunker or either deduplication layer. All four defects were in the operational plumbing, and two of them could lose data. 36 of 40 documented claims held exactly as written when audited, and all 40 hold now. Tests that pinned the broken behaviour were inverted rather than deleted, so a regression fails loudly.

rewrite

0.1 was a directory per node

It worked, and it carried a hand-rolled index to make it work: a snapshot, a journal, a reconcile pass, a background packer with fork handling, and a GC lock file. 0.2 deletes all of it and lets SQLite be the index.

The break is deliberate and total. There is no migration, and a 0.1 store is refused rather than half-read.

0.10.2
Hand-rolled index, journal and reconcile passdeletedSQLite is the index
Background packer and its fork handlingdeletedpacked at block exit
Average chunk size32 KiB16 KiB, 14% less stored and 22% faster
Provenance capture3 git subprocesses2, concurrent, nodes 2.2× faster
Chunker hot loop16.6 MB/s26.8 MB/s, boundaries identical
Selective find over 3,000 nodes0.40 ms0.04 ms
limits

What it is not for

A single SQLite file buys simplicity and crash safety. These are the consequences, and they are in the shipped documentation rather than only here.

  • Megabyte-scale: a 16 KiB artifact is roughly three times bigger in the store than as a plain file, because a database has pages and indexes whether you use them or not. The two ratios converge by a few megabytes.
  • One writer: many readers, one writer, local disk only. SQLite locking over NFS is unreliable, and heavy parallel writing is not what this is for.
  • No migrations, by design: a store records its format version and ancestree refuses anything it did not write. To read an old store, keep the version that wrote it. Every version stays on PyPI.
  • One file is the whole store: there is no side index to rebuild, so a corrupt database is real data loss. WAL journalling, an integrity check and `meta.json` sidecars are the mitigations.
  • Above 64 MiB: the chunker cuts at fixed offsets to keep huge ingests at C speed. Exact dedup survives, shift-resilience does not.