Back to glossary
Documents

Native vs scanned PDF

Also known as: PDF nativo vs escaneado

Definition

A native (digital) PDF contains selectable text and structured metadata generated digitally. A scanned PDF is just an image of a physical document, with no text directly extractable.

LATAM context

Modern LATAM banks generally deliver native PDFs from online banking. Some older formats or re-scanned printouts (when a customer prints a statement and scans it again) still need OCR for extraction.

Concrete example

A BBVA México PDF downloaded from Net Cash is native: you can select and copy the text. A PDF you received over WhatsApp after someone printed and re-scanned it is a scanned PDF, and it needs OCR.

How it shows up on your bank statement

It's the first thing worth checking before processing a batch of statements. A native PDF carries its text inside the file and reads exactly; a scanned one is an image, and every character has to be reconstructed. The same bank can deliver both formats depending on the channel: the one downloaded from online banking is usually native, and the one printed at a branch stops being native the moment you scan it.

How does finO$ handle this?

finO$ automatically detects the PDF type: native PDFs go straight to structured parsing (faster, 99%+ accuracy). Scanned ones go through OCR first (97%+ accuracy on good-quality scans).

Related terms

Need to convert bank statements into structured data?

Just one file? Use the free converter to convert your bank statement to Excel.

Frequently asked questions about Native vs scanned PDF

How do I tell a native PDF from a scanned one?

Try selecting text, or search for a word with Ctrl+F. If the search finds it, the PDF is native. If it finds nothing even though the word is visible on screen, the content is an image.

Does a native PDF guarantee a perfect extraction?

It guarantees the characters, not the structure. The text is there, but the table might not have declared columns, so you still have to figure out which column each number belongs to. It's the difference between reading correctly and understanding correctly.

Can I convert a scanned PDF into a native one?

You can add a text layer via OCR, and the file becomes searchable. But that layer inherits the recognition's own errors: the result looks like a native PDF without having its accuracy.