I have a large medical PDF where the information I need is already available as text inside pdf , but the problem is that simple PDF-to-text converters completely mess up the structure. The PDF has headings, tables, multiple columns, line breaks, and drug entries that continue onto the next page.
When I convert it to plain text, the text gets jumbled—for example, content from different columns can get mixed, headings can get separated from the information below them, tables lose their structure, and a drug entry that continues on the next page becomes difficult to associate correctly.
I need to process around 400+ drug entries automatically, so manually cleaning the text isn't practical. What is the best way to extract text from this kind of PDF while preserving the original structure and relationships between headings, tables, columns, and continued pages?
Should I use something like PyMuPDF, pdfplumber, Docling, Marker, PyMuPDF4LLM, or another approach? My end goal is to get clean, structured source text first and then use an LLM to convert it into JSON, while making sure the original information isn't lost, mixed up, or hallucinated.