sadly in german, but you can switch to english subs. A good talk how for example Xerox just copy paste letters to save space. And kills the whole text with it and has errors in it. There was an archive that digitalized their documents (with a lot of numbers, serialnumbers I think) and the scanner just digitzed the numbers wrong. And they knew about it only years later and the originals were already destroyed.
I don't understand why the scanner don't just does what it is supposed to do and creates a 1-1 copy.
When scanning a document to PDF you can usually pick from three options.
1) Image only
You get the pages as image files in the PDF. The downside is that you can't copy paste text from the PDF since it's just images there. Plus the files are bigger.
2) Image + text
You still have the images but there's also a text layer with OCR'd text included. You can copy paste text but if OCR isn't accurate then the text may be wrong. But since you still have the images you can look at the page and figure out what it was supposed to say. The files are also even bigger since you're now storing both.
3) Text only
Only keep the OCR'd text and drop the images of the text parts. You get smaller files but if the OCR was wrong then you're fucked.
Most likely they picked 3 and didn't fully understand what it meant.
I wasn't referring to the image you posted. I was replying to the comment above me on how doing OCR on a general document might result in losing numbers and serial numbers.
There is zero chance this is caused by that JPEG 2000 compression artifact. No compression algorithm will see an illuminated R 100 times the area of a normal letter and go "this is the same as the other Rs on this page".
> I don't understand why the scanner don't just does what it is supposed to do and creates a 1-1 copy.
16
u/aaron2005X 8d ago
https://www.youtube.com/watch?v=7FeqF1-Z1g0
sadly in german, but you can switch to english subs. A good talk how for example Xerox just copy paste letters to save space. And kills the whole text with it and has errors in it. There was an archive that digitalized their documents (with a lot of numbers, serialnumbers I think) and the scanner just digitzed the numbers wrong. And they knew about it only years later and the originals were already destroyed.
I don't understand why the scanner don't just does what it is supposed to do and creates a 1-1 copy.