Home · Free tools · Glossary · OCR

What is OCR?

OCR — optical character recognition — turns an image of text into text a computer can search, copy and calculate on. A photograph of a receipt is pixels; after OCR, the total is a number. Modern systems go further and infer layout, which is what makes "photo of a table" become "table".

The distinction that saves the most time

A PDF does not necessarily need OCR. PDFs come in two kinds: ones containing real text characters, and ones containing a picture of a page. Try to select a word in a PDF viewer — if it highlights, the characters are already there and can be copied exactly, with no model and no error rate. Only a scanned page needs OCR.

People routinely run scanned-document workflows on text PDFs and inherit an error rate they did not need to have.

Four things that still defeat it

  • Faded thermal paper. Receipts fade fastest where the ink was densest — often the total. Nothing recovers information that is no longer on the paper.
  • A warped baseline. A photo of a curled or angled document scrambles decimal points more than it scrambles letters, because a misplaced dot is still a plausible number.
  • Tables without ruling lines. Column boundaries inferred from whitespace break when one cell's content is long.
  • Handwriting, still, outside of narrow cases like forms with boxed characters.

The practical consequence: how you capture the image matters more than which OCR you use. Flat, evenly lit, straight on, as soon as possible.

Where it runs, and why that matters

OCR good enough for documents runs as a model, and models run somewhere. Free web converters run it on their servers, which means your document is uploaded. Some applications run a local model. Neither the tools on this site nor TableDI 2 does OCR at all: they read files that already contain text — a spreadsheet, a CSV, a PDF with a text layer — and say so when a file turns out to be a scan.

Questions people ask

Does a PDF need OCR?

Only if it is a scan. Try selecting text in a PDF viewer — if it highlights, the characters are real and can be copied exactly. Running OCR on a text PDF adds an error rate for nothing.

How accurate is OCR on receipts?

Highly variable, and driven more by capture than by software. A flat, evenly lit, recent receipt reads well; a curled, faded, angled one produces plausible wrong numbers, which is worse than failing.

Can OCR read handwriting?

Partially, and unreliably outside narrow cases like forms with boxed characters. Treat any handwritten figure as a draft to be checked.

Doing this every month?

In TableDI 2 you do it once, then save it as a job. Next month you drop in the new files and run it again.

macOS, Apple silicon and Intel; Windows is in progress (what to do meanwhile). Free is not a trial — no account, no card.

Last reviewed 2026-09-11 by the TableDI team. Something wrong on this page? Tell us — it is one inbox, read by the people who build TableDI.