Home · Free tools · Glossary · Data pipeline

What is a data pipeline?

A data pipeline moves data from where it is produced to where it is used, on a schedule, without a person doing it. ETL and ELT are pipeline shapes; so is a nightly export dropped in a folder, and so is a scheduled job that fetches an API and appends rows to a table.

The four things that break

  1. The source changes shape. A column is renamed upstream and the pipeline keeps running, writing blanks. This is the most common failure and the hardest to notice.
  2. Duplicates on re-run. A job that appends rather than upserts turns one retry into a double-counted day.
  3. Time zones and late data. Records that arrive after the window closed either vanish or land in the wrong day.
  4. Silence. A pipeline that fails loudly is a nuisance. A pipeline that fails quietly is a quarter of wrong reports.

Notice that three of the four produce plausible wrong numbers rather than errors. That is why the useful investment is not more pipeline, it is a check on the output — row counts against expectation, a total against a known figure — that runs every time.

The smallest pipeline that counts

A pipeline does not require a warehouse or an orchestrator. A scheduled fetch from an API that appends to a table and recalculates the columns built on it is a pipeline. It has all the same failure modes, at a scale where one person can actually see them.

TableDI 2 stops one step short of this, on purpose. TableDI 2 is a desktop app for file work you redo every period: it keeps your sources, rules and delivery as a job, so next month you drop in the new files and run it again. Everything runs on your own machine. Nothing runs on a schedule: a person starts each period, the preflight checks the inputs, and the exceptions list is the check on the output described above.

Questions people ask

What is the difference between a data pipeline and ETL?

ETL is a specific shape of pipeline — extract, transform, load. "Pipeline" is the general term and covers anything that moves data on a schedule, including a nightly file drop.

What is the most common way a pipeline fails?

An upstream schema change that the pipeline survives — it keeps running and writes blanks or defaults. It produces plausible wrong numbers rather than an error, which is why output checks matter more than error handling.

Do I need an orchestration tool?

Only at a scale where many jobs depend on each other. One scheduled fetch feeding one table is a pipeline too, and adding an orchestrator to it buys nothing but maintenance.

Doing this every month?

In TableDI 2 you do it once, then save it as a job. Next month you drop in the new files and run it again.

macOS, Apple silicon and Intel; Windows is in progress (what to do meanwhile). Free is not a trial — no account, no card.

Last reviewed 2026-09-11 by the TableDI team. Something wrong on this page? Tell us — it is one inbox, read by the people who build TableDI.