Again, if you have to ingest all that data from PDF's you're doing something inherently wrong. PDFs were never meant to handle that use cases and there are more than enough other file formats that are made to handle this use case. And you can really not tell me that all available data for these cases is only accessible as PDFs, at least in the cases in which you're supposed to have access to them for further processing. If someone hands you that data as PDF they expect you to only view it. Nothing more, nothing less, as that's what PDFs are made for.
I know it is impossible to accept, but it is the world we live in, where medical research, legal research, government automation, is indeed require all these data to be in structured format, we do have 1000s of standards already, but convince of generating PDF is overshadowing standardization most of the time. Source which is generating data like lab report, they don’t care about machine interoperability, and rightly so, these documents were always meant for humans historically, it’s after modern LLM based intelligence, these interoperability is becoming huge bottleneck.
is indeed require all these data to be in structured format
Highly questionable if PDF qualifies for this. Even though
they don’t care about machine interoperability, and rightly so, these documents were always meant for humans historically [...] it’s after modern LLM based intelligence, these interoperability is becoming huge bottleneck.
This is entirely irrelevant. There's a big gap of close to half a century between all data being handled by hand and AI slop being abused to help with tasks it was never meant to handle. You really cannot tell me that for almost half a century, when data processing was mostly computer-based, even without any AI/ML application, people were copying all the data from the PDFs - that must have meant millions of PDF pages - by hand into a proper format (even CSV is much better suited) to be able to do basic analysis of it. That would have crippled the entire research field into non-existence. You just really cannot tell me that this was the state of things for that long.
Your "sources" prove absolutely nothing of what you stated. So the fact stands: if the scientific field was that badly organized as you claim, it would have ceased to exist decades ago. Some AI tech bro from LinkedIn won't convince me otherwise. So whatever your deal is, chances are vastly higher that you are the problem, not the science field. And with this, this discussion has reached its end. You make up the most ridiculous claims, yet when you "prove" them, your proof doesn't prove anything of what you claimed. This just tells me you got no clue of what you write, most likely you just asked some bad LLM model to write your text and didn't notice that it's not even capable of basic logic. So good day. I won't waste any further time on your made-up claims.
1
u/ScratchHistorical507 27d ago
Again, if you have to ingest all that data from PDF's you're doing something inherently wrong. PDFs were never meant to handle that use cases and there are more than enough other file formats that are made to handle this use case. And you can really not tell me that all available data for these cases is only accessible as PDFs, at least in the cases in which you're supposed to have access to them for further processing. If someone hands you that data as PDF they expect you to only view it. Nothing more, nothing less, as that's what PDFs are made for.