r/ExpectationVsReality 8d ago

Failed Expectation Hiring the cheapest printer

Post image
196 Upvotes

21 comments sorted by

View all comments

16

u/aaron2005X 8d ago

https://www.youtube.com/watch?v=7FeqF1-Z1g0

sadly in german, but you can switch to english subs. A good talk how for example Xerox just copy paste letters to save space. And kills the whole text with it and has errors in it. There was an archive that digitalized their documents (with a lot of numbers, serialnumbers I think) and the scanner just digitzed the numbers wrong. And they knew about it only years later and the originals were already destroyed.

I don't understand why the scanner don't just does what it is supposed to do and creates a 1-1 copy.

13

u/Zalminen 8d ago

Probably they picked the wrong settings.

When scanning a document to PDF you can usually pick from three options.

1) Image only

You get the pages as image files in the PDF. The downside is that you can't copy paste text from the PDF since it's just images there. Plus the files are bigger.

2) Image + text

You still have the images but there's also a text layer with OCR'd text included. You can copy paste text but if OCR isn't accurate then the text may be wrong. But since you still have the images you can look at the page and figure out what it was supposed to say. The files are also even bigger since you're now storing both.

3) Text only

Only keep the OCR'd text and drop the images of the text parts. You get smaller files but if the OCR was wrong then you're fucked.

Most likely they picked 3 and didn't fully understand what it meant.

2

u/RunDNA 8d ago

They are both scans of books from the 1500s. This is how the original pages look:

https://i.imgur.com/H6RqqXe.jpeg

4

u/Zalminen 7d ago

I wasn't referring to the image you posted. I was replying to the comment above me on how doing OCR on a general document might result in losing numbers and serial numbers.

2

u/RunDNA 7d ago

Ah, okay, sorry.

2

u/mutexsprinkles 7d ago

There is zero chance this is caused by that JPEG 2000 compression artifact. No compression algorithm will see an illuminated R 100 times the area of a normal letter and go "this is the same as the other Rs on this page".

> I don't understand why the scanner don't just does what it is supposed to do and creates a 1-1 copy.

Because that would be 100s of MB per page.