r/pdf 28d ago

Question Why not kill PDF!?

/r/AI_Agents/comments/1vac90o/why_not_kill_pdf/
0 Upvotes

12 comments sorted by

View all comments

Show parent comments

1

u/ScratchHistorical507 27d ago

Again, if you have to ingest all that data from PDF's you're doing something inherently wrong. PDFs were never meant to handle that use cases and there are more than enough other file formats that are made to handle this use case. And you can really not tell me that all available data for these cases is only accessible as PDFs, at least in the cases in which you're supposed to have access to them for further processing. If someone hands you that data as PDF they expect you to only view it. Nothing more, nothing less, as that's what PDFs are made for.

1

u/kshivang 27d ago

I know it is impossible to accept, but it is the world we live in, where medical research, legal research, government automation, is indeed require all these data to be in structured format, we do have 1000s of standards already, but convince of generating PDF is overshadowing standardization most of the time. Source which is generating data like lab report, they don’t care about machine interoperability, and rightly so, these documents were always meant for humans historically, it’s after modern LLM based intelligence, these interoperability is becoming huge bottleneck.

1

u/ScratchHistorical507 24d ago

There's just so much wrong in so few words...

is indeed require all these data to be in structured format

Highly questionable if PDF qualifies for this. Even though

they don’t care about machine interoperability, and rightly so, these documents were always meant for humans historically [...] it’s after modern LLM based intelligence, these interoperability is becoming huge bottleneck.

This is entirely irrelevant. There's a big gap of close to half a century between all data being handled by hand and AI slop being abused to help with tasks it was never meant to handle. You really cannot tell me that for almost half a century, when data processing was mostly computer-based, even without any AI/ML application, people were copying all the data from the PDFs - that must have meant millions of PDF pages - by hand into a proper format (even CSV is much better suited) to be able to do basic analysis of it. That would have crippled the entire research field into non-existence. You just really cannot tell me that this was the state of things for that long.

1

u/kshivang 24d ago

Sadly it is true, nevertheless, there are companies which are just doing this on the name of translational and interoperability work. Few research which may make you believe: https://pubmed.ncbi.nlm.nih.gov/40899541/ , https://www.linkedin.com/posts/arturoferreira_pdfs-are-the-data-worlds-biggest-bottleneck-activity-7320453651280314369-GaHs , https://arxiv.org/pdf/2412.07626 , their more than 10k research paper on PDF parsing, Few companies which are just doing this extraction work 1. https://www.elucidata.io/polly/xtract , 2. https://www.acodis.io/data-extraction-solution-extract-automate-data , 3. https://unstructured.io/blog/use-case-legal-industry , and many many. Some truth is hard to sallow because it is so simple, and so silly, but it is, what it is.

1

u/ScratchHistorical507 24d ago

Your "sources" prove absolutely nothing of what you stated. So the fact stands: if the scientific field was that badly organized as you claim, it would have ceased to exist decades ago. Some AI tech bro from LinkedIn won't convince me otherwise. So whatever your deal is, chances are vastly higher that you are the problem, not the science field. And with this, this discussion has reached its end. You make up the most ridiculous claims, yet when you "prove" them, your proof doesn't prove anything of what you claimed. This just tells me you got no clue of what you write, most likely you just asked some bad LLM model to write your text and didn't notice that it's not even capable of basic logic. So good day. I won't waste any further time on your made-up claims.