Home · Free tools · Articles · Extract tables from PDF

How to extract tables from a PDF

Ranked by accuracy rather than convenience, because the fast route is the one that produces numbers that are wrong in ways nobody notices. The first question decides everything: does this PDF contain real text, or a picture of text?

First: which kind of PDF is it

Open it and try to select a number with your cursor. If it highlights, the characters are real. If nothing happens, the page is an image and you are in OCR territory — a different problem with a different error rate. This one check changes which of the routes below are even available, and most guides skip it.

Route 1 — copy and paste (exact, when it works)

For a text PDF, select the table and paste it. No model, no error rate: the characters transfer exactly. The mess afterwards is structural — everything lands in one column, or columns split on the wrong spaces.

Fix it with a split step rather than by hand: paste into a sheet, then split on the delimiter that actually separates the columns. Our CSV viewer will show you what you pasted; a text-to-columns step in your spreadsheet does the reshaping. Tedious, and still the most accurate route available.

Route 2 — a library that parses the PDF structure

Tools like Tabula and Camelot, or Python libraries, read the PDF's internal layout — text positions and ruling lines — and reconstruct the table. Exact on the characters, since nothing is being recognized, and good on tables with visible ruling lines. They struggle where columns are implied by whitespace only, and they need you to be comfortable running them.

Route 3 — an online converter

Fast, free, and uploads your document. Fine for a public report. Not fine for a bank statement, a payslip or anything under an NDA — and worth reading the retention policy before deciding, treating anything unstated as unbounded.

Route 4 — a model reads it

The only option for scanned pages, and the most forgiving on messy layouts, because it infers structure the way a person would. It is also the only route that can be confidently wrong: a misread decimal produces a plausible number, not an error.

Neither the tools on this site nor TableDI 2 take this route: they read PDFs that already contain text — the PDF table reader here does it in your browser — and say so when a page turns out to be a scan.

The check that makes any of this safe

Whichever route: reconcile a total against a number you already know before building anything on the result. Row count against the document, a column total against the stated total. It takes thirty seconds and it catches the decimal errors, the page-break losses and the repeated-header rows that all four routes produce.

A worked version of this for statements specifically: PDF financial statements to Excel.

Questions people ask

What is the most accurate way to get a table out of a PDF?

Copy and paste, if the PDF contains real text — nothing is being recognized, so the characters transfer exactly. Everything else, including AI, introduces an error rate. Check first by trying to select text in a viewer.

How do I know if a PDF is scanned?

Try to select a word. If it highlights, the text is real; if nothing selects, the page is an image and needs OCR.

Is it safe to use an online PDF-to-Excel converter?

It depends entirely on the document. The file is uploaded to their server, so treat it as publishing. For a public report, fine. For a statement, payslip or anything confidential, use a route that does not upload.

Doing this every month?

In TableDI 2 you do it once, then save it as a job. Next month you drop in the new files and run it again.

macOS, Apple silicon and Intel; Windows is in progress (what to do meanwhile). Free is not a trial — no account, no card.

Last reviewed 2026-09-11 by the TableDI team. Something wrong on this page? Tell us — it is one inbox, read by the people who build TableDI.