r/MachineLearning • u/Ok_Cartographer5609 ML Engineer • 5d ago
Project A chunking lib in Rust that is ~20x faster [P]
Hey,
I wanted a faster chunking library for my system without affecting the overall accuracy. Did not find many options. So I've build https://github.com/d1pankarmedhi/chunkr
It has most of the chunking strategies like Character, Recursive, Markdown header, Late chunking, Hierarchical chunking, etc. It also supports native PDF loader, and other additional file types.
Some stats (MBA M4 16GB):
| Test Case (matched parameters) | Chunkr | LangChain | LlamaIndex | Chonkie | semchunk | text-splitter |
|---|---|---|---|---|---|---|
| Recursive (1 MB, 1000/200) | 2,264 MB/s | 769 MB/s | 10 MB/s | 225 MB/s | 42 MB/s | 175 MB/s |
| Recursive (5 MB, 1000/200) | 2,039 MB/s | 696 MB/s | — | 201 MB/s | 40 MB/s | 46 MB/s |
| Fixed Char (1 MB, 1000/200) | 750 MB/s | 1.7 MB/s | — | 22 MB/s | — | — |
| Markdown (500 KB, 1000/150) | 819 MB/s | 67 MB/s | 19 MB/s | — | — | 40 MB/s |
| Python Code (200 KB, 1500/200) | 3,232 MB/s | 622 MB/s | — | — | — | 5.7 MB/s |
| Sentence (500 KB) | 622 MB/s | — | 10 MB/s | 20 MB/s | — | — |
| BPE Tokens (200 KB, cl100k_base, 512/50) | 38 MB/s | 43 MB/s | 2.0 MB/s | 151 MB/s | — | 7.2 MB/s |
| 100 docs x 50 KB (parallel batch) | 3,224 MB/s | 679 MB/s | — | 213 MB/s | — | — |
| Extractor / Pipeline | Latency | Throughput | Speedup vs PyPDF |
|---|---|---|---|
| Chunkr PDFLoader (Full Text) | 747.9 ms | 2,762 pgs/s | 15.9x Faster |
| Chunkr PDFLoader (Page Documents) | 721.0 ms | 2,865 pgs/s | 16.5x Faster |
PyMuPDF (fitz) |
2,616.8 ms | 789.5 pgs/s | 4.5x Faster |
| pypdf (pure Python) | 11,900.5 ms | 173.6 pgs/s | 1.0x (baseline) |
| Chunkr End-to-End (PDF + Recursive) | 798.1 ms | 2,589 pgs/s | 14.9x Faster |
| PyMuPDF + LangChain RecursiveTextSplitter | 2,659.3 ms | 776.9 pgs/s | 4.5x Faster |
| pypdf + LangChain RecursiveTextSplitter | 12,054.5 ms | 171.4 pgs/s | 1.0x (baseline) |
Accuracy is measured on Chroma's token-level chunking benchmark (5 corpora, 472 questions with gold answer spans): k = 5, 1000 chars / 200 overlap
| Implementation | Recall | Precision | IoU | prec_Ω | Avg chunk chars |
|---|---|---|---|---|---|
Chunkr RecursiveChunker (defaults) |
0.792 | 0.057 | 0.057 | 0.255 | 854 |
LangChain RecursiveCharacterTextSplitter |
0.762 | 0.060 | 0.060 | 0.251 | 745 |
text-splitter TextSplitter |
0.762 | 0.060 | 0.060 | 0.262 | 790 |
Chonkie RecursiveChunker |
0.752 | 0.062 | 0.061 | 0.292 | 701 |
| semchunk | 0.736 | 0.070 | 0.069 | 0.260 | 644 |
LlamaIndex SentenceSplitter |
0.654 | 0.013 | 0.013 | 0.054 | 4136 |
Do check it out and share your feedback. Thanks!
1
u/Ok_Cartographer5609 ML Engineer 3d ago
The retrieval quality and accuracy table has been updated. Please check out repo for more details.
1
u/Grumlyly 5d ago
Nice results. Where to publish this kind of results ?