A test catches a null where one shouldn’t be. It doesn’t catch a table that stopped updating three days ago, or row counts that quietly dropped by half for no obvious reason. Data observability vs data quality testing is the difference between checking rules you already thought of and noticing something broke that you didn’t.
What data-quality testing actually covers
Tests like uniqueness, not-null, and referential integrity check specific, predefined rules against the data. They’re precise and reliable for exactly what they’re written to catch. Their limitation is built into how they work: a test only catches a failure mode someone already anticipated and wrote a check for.
What data observability adds
Observability monitors the shape of the data itself over time, row counts, freshness, schema, distribution, and flags anomalies without needing a specific rule written in advance. If a table’s row count usually grows by a few thousand rows a day and suddenly grows by zero, observability can flag that as unusual, even though no test was written for “row count dropped to zero.”
A real example: where testing already does real work
The retail analytics warehouse and Olist analytics engineering project both run 30-plus automated dbt tests on every build. Those tests catch a real and specific category of failure well: known rules being violated. What they don’t automatically catch is a source system going silent, or a gradual distribution shift that never technically breaks a not-null or uniqueness rule.
Why the gap matters
A pipeline can pass every existing test while still being quietly wrong, a source feed stopped updating, but the last successful load still satisfies every constraint that was ever written for it. Data-quality testing alone has no way to notice that absence, since nothing about a stale table necessarily violates a rule.
Where observability fits alongside testing, not instead of it
- Testing catches known, specific rule violations reliably and immediately.
- Observability catches unknown or gradual anomalies that no one wrote a rule for.
- Freshness checks, part of an observability layer, catch a source that stopped updating, which most standard dbt tests wouldn’t flag on their own.
- Volume anomaly detection catches a row count that dropped or spiked well outside its normal historical range.
A quick checklist
- Would your current tests catch a source that simply stopped sending new data?
- Do you track expected freshness for each table, or only whether existing rows meet their constraints?
- If row counts dropped by half tomorrow with no rule technically violated, would anyone notice quickly?
- Are you relying entirely on tests for failure modes you haven’t specifically anticipated yet?
FAQ
Does observability replace the need for data-quality tests?
No. They cover different failure modes. Tests remain the precise, reliable check for known rules; observability adds coverage for what wasn’t anticipated.
Is data observability only relevant at large scale?
It becomes more valuable at scale, where manual review of every table isn’t feasible, but even a simple freshness check on a small project catches a real category of failure testing alone misses.
Can dbt tests be extended toward observability?
To a degree, freshness checks and custom tests can approximate some observability functions, though dedicated observability tools typically track anomalies more automatically, without a rule being written for each one in advance.

