← Browse skills

Data

Check the data pipeline ran

Confirm the numbers people are about to use are fresh and complete, by checking freshness and volume rather than job status.

By Toolspoke

Skill procedure

A green job is not fresh data. Jobs succeed against empty inputs, partial loads and yesterday's partition every day of the week, and the dashboards read exactly the same either way. Check the data, not the runner.

Steps

  1. Read the last run: did it finish, when, and how long did it take. A run that succeeded in a quarter of its usual time processed a fraction of its usual rows.
  2. Check freshness at the source, per table: the newest timestamp in each table the models read. Compare to now and to what the table's own cadence should be. This catches the case where the pipeline is healthy and the extraction upstream stopped.
  3. Check volume against the same weekday last week, not against yesterday. Weekly seasonality makes Monday against Sunday useless. A day at half the usual row count is a partial load whatever the status says.
  4. Check the tests that guard meaning: uniqueness on keys, not-null on the columns joins rely on, and referential checks between fact and dimension. These are the ones that produce silently wrong numbers instead of errors.
  5. Look at the one or two metrics people will actually read and see whether they moved in a way the business would explain. A revenue figure at zero for a region is a pipeline problem until proven otherwise.
  6. Say it in one line in the channel that uses the data: fresh as of a timestamp, volumes normal or not, and anything not to trust today. Say it every day, including the days it is fine — a report that only appears when something is wrong trains people to ignore the silence.

Stop and ask

  • Before re-running a job that writes to production tables, especially a full refresh: it costs money and can leave tables empty while it runs.
  • Before telling people the data is fine when a test failed. Say which test, and what it means for which number.