Home · Free tools · Glossary · Data extraction

What is data extraction?

Data extraction is the process of pulling specific, structured fields out of a source that was not organized for you — an invoice PDF, a photograph of a table, a web page, an email body, a log file. The output is rows and columns; the input is something that merely contains the information.

What separates it from neighbouring terms

  • Not the same as export. If the source already offers CSV, you are exporting, not extracting — and you should always check for an export before extracting anything, because an export is exact and extraction is inference.
  • Not the same as scraping. Scraping is one source of extraction (web pages). Extraction covers documents, images and messages too.
  • Not the same as the E in ETL. In ETL, "extract" usually means reading from a database or API that already has structure. Document extraction is the harder problem: the structure has to be inferred.

The part that gets glossed over

Extraction accuracy is almost always quoted at the field level — "99% accurate" — and read at the document level. Those are very different numbers. An invoice with 20 fields at 99% per-field accuracy has roughly an 82% chance of being entirely correct. On a hundred invoices, about eighteen have something wrong in them.

That is why every serious extraction workflow ends with a reconciliation step — comparing an extracted total against a number you already know. If a tool's pitch has no reconciliation step in it, the pitch is incomplete.

Where extracted data goes wrong, in order of frequency

  1. Decimal separators. 1.234,56 read as 1.23 is silent and catastrophic.
  2. Dates. 03/04/2026 is March or April depending on where the document came from.
  3. Table continuation. Rows lost at a page break, or a repeated header read as data.
  4. Signs. Credits and debits, parentheses for negatives, refunds.

Related on this site

Invoice data extraction compares the practical routes and reads text PDFs in the browser; what is OCR covers photographs and scans; and PDF statements covers the case where the source is a bank or broker statement.

Questions people ask

Is data extraction the same as web scraping?

No — scraping is one kind of data extraction, the kind whose source is a web page. Extraction also covers PDFs, images, emails and log files.

How accurate is data extraction?

Accuracy is usually quoted per field and read per document, and the gap is large: 99% per field across 20 fields means roughly 82% of documents are fully correct. Always reconcile an extracted total against a number you already have.

Do I need AI for data extraction?

Only when the structure has to be inferred — a scanned document, an inconsistent layout, free text. If the source has an export, an API, or a PDF with real selectable text, a rule or a copy-paste is both cheaper and exact.

Doing this every month?

In TableDI 2 you do it once, then save it as a job. Next month you drop in the new files and run it again.

macOS, Apple silicon and Intel; Windows is in progress (what to do meanwhile). Free is not a trial — no account, no card.

Last reviewed 2026-09-11 by the TableDI team. Something wrong on this page? Tell us — it is one inbox, read by the people who build TableDI.