Back to glossary
Banking / Technical LATAM

OCR

Also known as: Optical Character Recognition, Reconocimiento Óptico de Caracteres

Definition

OCR (Optical Character Recognition) is the technology that converts images of text (scans, photos) into editable, searchable digital text. It's what makes it possible to process scanned bank statement PDFs.

LATAM context

Across LATAM, many regional banks and some older formats deliver scanned PDFs rather than digital-native ones. OCR is what's needed to extract data from those documents. Modern accuracy on Spanish-language text exceeds 98% on good-quality scans.

Concrete example

A scanned 2018 PDF from Banco del Bajío gets run through OCR first (turning the image into text) and then through structured parsing (interpreting the column layout as transactions).

How it shows up on your bank statement

OCR doesn't show up on the statement. It's what's needed to read it when the PDF is an image. If you open the file and can't select the text with your cursor, that PDF needs OCR. Typical case: a statement printed at a branch and then scanned, or a photo taken with a phone.

How does finO$ handle this?

finO$ applies OCR automatically when it detects a scanned or low-quality PDF. For native PDFs (most modern banks) it skips OCR entirely and goes straight to structured parsing, which improves both speed and accuracy.

Need to convert bank statements into structured data?

Just one file? Use the free converter to convert your bank statement to Excel.

Frequently asked questions about OCR

How do I know if my statement needs OCR?

Open the PDF and try selecting a number with your cursor. If you can select and copy it, the file has a text layer and doesn't need OCR. If your selection grabs an image box instead, the document is a photo of the statement and needs character recognition.

How reliable is OCR with dollar amounts?

It depends heavily on the source quality. A straight, well-lit scan reaches very high accuracy; a crooked photo of a crumpled page doesn't. The typical errors aren't random: confusing 0 and O, 1 and 7, and misreading thousands separators: exactly the mistakes that do the most damage in an amount.

Does OCR understand the table structure, or just the letters?

Recognizing characters and reconstructing the table are two different problems. An OCR engine can read every number correctly and still assign it to the wrong column if the table has no visible borders. That's why post-extraction validation matters so much: checking that balances add up row by row.