@VikParuchuri - You might be wondering why you should even bother with the

Vik Paruchuri
Vik Paruchuri@VikParuchuri
Parsing PDFs has slowly driven me insane over the last year. Here are 8 weird edge cases to show you why PDF parsing isn't an easy problem. 🧵
image not captured
Vik Paruchuri
Vik Paruchuri@VikParuchuri
PDFs have a font map that tells you what actual character is connected to each rendered character, so you can copy/paste. Unfortunately, these maps can lie, so the character you copy is not what you see. If you're unlucky, it's total gibberish.
image not captured
Read 7 more tweets
Vik Paruchuri
Vik Paruchuri@VikParuchuri
PDFs can have invisible text that only shows up when you try to extract it. "Measurement in your home" is only here once...or is it?
image not captured
Vik Paruchuri
Vik Paruchuri@VikParuchuri
Math is a whole can of worms. Remember the font map problem? Well, math is almost always random characters - here we get some strange Tamil/Amharic combo.
image not captured
Vik Paruchuri
Vik Paruchuri@VikParuchuri
Math bounding boxes are always fun - see how each formula is broken up into lots of tiny sections? Putting them together is a great time!
image not captured
Vik Paruchuri
Vik Paruchuri@VikParuchuri
Once upon a time, someone decided that their favorite letters should be connected together into one character - like ffi or fl. Unfortunately, PDFs are inconsistent with this, and sometimes will totally skip ligatures - very ecient of them.
image not captured
Vik Paruchuri
Vik Paruchuri@VikParuchuri
Not all text in a PDF is correct. Some PDFs are digital, and the text was added on creation. But others have had invisible OCR text added, sometimes based on pretty bad text detection. That's when you get this mess:
image not captured
Vik Paruchuri
Vik Paruchuri@VikParuchuri
Overlapping text elements can get crazy - see how the watermark overlaps all the other text? Forget about finding good reading order here.
image not captured
Vik Paruchuri
Vik Paruchuri@VikParuchuri
I've been showing you somewhat nice line bounding boxes. But PDFs just have character positions inside - you have to postprocess to join them into lines. In tables, this can get tricky, since it's hard to know when a new cell starts:
image not captured
Vik Paruchuri
Vik Paruchuri@VikParuchuri
2025-08-12
You might be wondering why you should even bother with the text inside PDFs. The answer is that a lot of PDFs have good text, and it's faster and more accurate to just pull it out. This is what we do with marker - https://github.com/datalab-to/marker - we only OCR if the text is bad.

View on X →