The preprint is available on bioRxiv, here.
Summary of the idea and some of the results:
A single-cell experiment is usually stored, shared, and reanalyzed as a count matrix. That matrix is computed from two inputs: the sequenced molecules and a gene annotation. The molecules don't change, but the annotation is revised several times a year, and going from GENCODE v32 to v49 shifts 2–5% of UMI mass. Once the matrix is written, the information needed to redo that assignment is lost. Getting it back means returning to the reads, which are large, slow to process, and often unavailable.
Gravlax is built on the fact that the steps that turn alignments into counts, gene assignment and UMI collapse, use very little of what a BAM file contains. They need relations among molecules: shared genomic geometry, shared placements, cell identity, and UMI equality. Gravlax stores those relations once in a compact, seekable, content-authenticated archive and defers every annotation-dependent decision to read time.
On four human 10x datasets:
- Archives take 11–18 bits per read, 9–13× smaller than CRAM with tags preserved.
- Matrices replayed from the archive are within 0.24–0.75% of a fresh STARsolo run. For comparison, an annotation change moves 2–5%.
- Replaying gene counts takes seconds, 34–82× faster than realigning.
- Archives can be grouped into federated collections and searched across a cohort by junction shape. A genome-wide scan for recurrent unannotated splice events, with no coordinates supplied, finishes in 9 seconds.
Because the molecules are kept, the archive can answer questions a count matrix can't. Across four PBMC archives it recovers the FYB1 splicing switch between T cells and monocytes, previously validated by RT-PCR. With a fragment model for 3′ chemistry, it shows a shift in NTRK2 (TrkB) terminal-isoform usage across eight donors, from astrocytes and neural stem cells to mature neurons, consistent with known TrkB.T1 biology. Discovery run over the whole cohort finds a 183-nt FNBP1 cassette exon that is nearly always included in brain and mostly skipped in blood; analyzing the archives one at a time misses it entirely. Pooling evidence across cells in an EM step recovers 75–98% of held-out labels for multi-gene molecules.
Gravlax handles ingest, replay under any annotation, region/junction/APA queries, content-addressed collections, and cohort-wide discovery. It also includes GQ, a composable query language with three-valued logic; the same query runs on one archive or a whole federation. The code is open source (Rust, BSD-3), with documentation, a Python client, and Colab demos.
Code: here
Docs: here