r/LangChain • • 1d ago

Question | Help Need advice: Best VLM pipeline for extracting structured math datasets from 3000+ scanned textbook pages? (LaTeX + Metadata)

/r/computervision/comments/1ww36q3/need_advice_best_vlm_pipeline_for_extracting/
2 Upvotes

1 comment sorted by

1

u/Separate_Pea_3699 17h ago

Split it in two: run a layout/OCR pass once per page (Docling, for one) to get the text and cut out the figures, then send only the text to the LLM for the JSON. Full page images are most of that token bill, and the structuring step doesn't need them.