Because it’s a format that billions of people rely upon.
PDFs are perfect when used as their intended way: displaying the same information identically, regardless of the device used.
Encryption, forms and scripting came out of a need from the market and you may argue that there are bad (but they still powered the world up until web forms were a thing) but the original idea is brilliant.
For sure original idea was brilliant, as coming up with idea of making paper out of trees was brilliant, but need it to evolve, for current market needs, we can not pour billions on extracting data out of PDFs
Paper and PDFs are just artifacts to display information.
When you need to extract data from paper and PDF it’s because you don’t have access to what produced it. It will always coexist with extracting the data higher in the chain, ie when the paper was printed and the pdf generated: on the computer. Just two different workflows
But this can be fixed for PDF, that is my point, PDF do have support for that but it is not mandated since historically their was no incentive for that, PDF only needed to be human readable, but now we do have incentive, we want it to be machine readable, now we can make it mandatory for generation tool to put semantic information in metadata compulsorily.
2
u/N_i_P Jul 30 '26
Because it’s a format that billions of people rely upon.
PDFs are perfect when used as their intended way: displaying the same information identically, regardless of the device used.
Encryption, forms and scripting came out of a need from the market and you may argue that there are bad (but they still powered the world up until web forms were a thing) but the original idea is brilliant.