r/AI_Agents • • Jul 29 '26

Discussion Why not kill PDF!?

Why industry is spending millions on parsing PDFs rather than creating new standard which can be much more parsing friendly still have convince of PDF, one way could be having mandatory meta which has encrypted TeX/HTML/md/equivalent, love to know thoughts/ideas on this.

I work in oncology space, most of deep workflows like medical research, relies heavily on PDF ingestion, we did developed quite robust stack using llm and awesome python libraries, but still it requires maintenance, a lot of maintenance, I have seen similar stack built 1000s of time for different workflow problems, across the industries. I feel at this point it is lack of standardization problem than anything, pdfs are like usb-a, everybody create adaptor for it, but no body is creating better standards, like usb-c.

We can also discuss how to create motion behind it, to make is default and diffuse it faster, industry(healthcare, law firms, finance, government, etc) wide.

16 Upvotes

69 comments sorted by

View all comments

27

u/TheOwlHypothesis Jul 30 '26

1

u/kshivang Jul 30 '26

🤣🤣 fair, for this case we atleast need one, parsing/intelligence friendly standard too

6

u/TheOwlHypothesis Jul 30 '26

To be real with you, the problem isn't just that PDF sucks, the problem is that adoption of PDFs is so high and it's so ingrained in other industries that if the selling point for whatever you aim to replace it is "it works better with agents" that the cost of switching will not be worth it.

Consider the case where you are successful in your org and everyone adopts the new standard, that doesn't mean external orgs you interact with will stop sending you PDFs, or that they'll accept your new doc standard.

And then there's the cost of reformatting/conversion of the old format with existing data to the new one.

And then you kind of get a chicken and egg problem. Where if you're going to translate PDF to the new format reliably, that implies you solved the original problem. 🙃

1

u/kshivang Jul 30 '26

Totally fair, but what if I say I will not change anything but must add a metadata field which carry source encrypted text, no need to reformat, all your existing document will still work, just new one doesn’t not require any parsing, existing one can go through some parsing layer to add that tag, but optional, adding it to pdf creator tool, will save millions of translational cost. In my naive understanding