dbt Tests vs Data Observability: Where Each One Stops

dbt tests check assertions you wrote, at build time, on tables dbt owns. Every word in that sentence is a boundary. Here is what lives outside each one.

dbt Tests vs Data Observability: Where Each One Stops

A dbt test is an assertion you write about a model, evaluated when dbt runs, against a table dbt built. Data observability is continuous monitoring of tables in your warehouse regardless of what produced them or whether anyone wrote a rule.

Those are not competing products. They are different sentences, and the interesting part is which words differ: you write, when dbt runs, and dbt built. Each of those is a boundary, and incidents live outside boundaries.

This post is about where each one stops. Not which is better, because that question has no answer, and most teams doing this well run both.

What do dbt tests actually catch?

They catch violations of assertions you thought to write. That is a real and useful category, and it is smaller than most teams believe.

# models/marts/schema.yml
models:
  - name: fct_orders
    columns:
      - name: order_id
        tests:
          - unique
          - not_null
      - name: status
        tests:
          - accepted_values:
              values: ['pending', 'shipped', 'delivered']
      - name: customer_id
        tests:
          - relationships:
              to: ref('dim_customers')
              field: customer_id

Four tests, and each is genuinely valuable. unique and not_null catch a broken join fanning out rows. accepted_values catches an upstream system introducing a status nobody told you about. relationships catches orphaned foreign keys.

Notice what they have in common. Every one encodes something a human knew: that order_id should be unique, that there are exactly three statuses, that customers should exist. dbt tests are the externalization of your team's knowledge about the data, which is why they're worth writing and why they can't be the whole strategy.

They run when dbt runs. On a project running once nightly, that means your assertions are checked once a day, at a time you chose, on tables that had already been built.

Where does "you write" stop?

At the edge of what you thought of.

Nobody writes a test asserting that the row count should be within a normal range for a Tuesday, because expressing that requires knowing the distribution, and the distribution changes. Nobody writes a test asserting that a column's average shouldn't shift 30% overnight. Nobody writes a test for the column that got added last week, because they didn't know it was coming.

Here is the shape of the gap, using a real incident type. An upstream vendor changes a currency field from dollars to cents. Every dbt test passes: order_total is still not null, still numeric, still positive. The relationship tests still hold. Your revenue numbers are 100x too high and the build is green.

The only signal available is that the values look nothing like yesterday's, and no assertion you would plausibly write catches that. It needs a baseline, meaning a model of what normal looked like, which is the thing dbt tests structurally do not have.

The reverse gap is real too. An observability tool watching order_total sees the distribution shift and alerts. It has no idea that three statuses are the valid set, because that fact exists only in a business rule somebody knows. That's why you don't need to write data tests overstates it as a slogan: you need fewer, and the ones that remain are the ones encoding knowledge.

What breaks dbt test Observability
Join fans out, duplicate keys Yes, unique Yes, via row count anomaly
New enum value from upstream Yes, accepted_values No, unless it shifts a distribution
Orphaned foreign keys Yes, relationships No
Currency changes units No Yes, distribution shift
Source stopped updating No Yes, freshness
Column added or type changed No Yes, schema monitoring
Volume drops 40% on a Tuesday No Yes, baseline
Business rule violated Yes, if you wrote it No

Read that table as two columns of roughly equal length, because that's the actual finding. Neither one dominates.

Where does "when dbt runs" stop?

At the gap between runs, which is most of the day.

If dbt runs nightly at 02:00, an ingestion failure at 19:04 is invisible for seven hours. Worse, the 02:00 run succeeds, because stale data is still data, and every test passes on it. Your first signal is a person opening a dashboard at 09:00.

This is the boundary that surprises people, because "we have tests" feels like continuous coverage and is actually a snapshot on a schedule. Tests are a build-time gate. Monitoring is a clock.

The other half of this boundary is severity. dbt tests default to error, which fails the build, but plenty of tests get set to warn to keep pipelines green:

      - name: status
        tests:
          - accepted_values:
              values: ['pending', 'shipped', 'delivered']
              config:
                severity: warn        # build stays green, nobody reads the log

A warning in a dbt log is not an alert. It is a line in output nobody reads, and it is the most common way teams have coverage on paper and none in practice. I covered that failure mode in detail in silent dbt test failures, and it's worth auditing your project for severity: warn before you conclude anything about your coverage.

There's a middle option worth knowing:

              config:
                severity: error
                error_if: ">100"      # fail only if more than 100 rows violate
                warn_if: ">0"

That gets you a build that fails on real breakage and warns on noise, which is better than blanket warn and requires you to pick a threshold.

Where does "dbt built" stop?

At the edge of your dbt project, which is smaller than your warehouse.

Raw tables from Fivetran or Airbyte. Reverse ETL outputs. Airflow tasks that write directly. Stored procedures. An analyst's scheduled query. Anything predating dbt adoption. None of it has a schema.yml entry to hang a test on.

Declaring raw tables as dbt sources extends the boundary somewhat, and dbt source freshness is genuinely useful:

dbt source freshness    # checks loaded_at_field against warn_after / error_after

Two caveats. It covers only the sources you declared, and declaring sources is a chore that decays. And it runs when you run it, so it's the same build-time snapshot rather than a clock.

Count your own boundary. List the tables in your warehouse, mark which ones dbt produces, and take the ratio. In warehouses older than about three years I have not seen that number above 80%, and the tables outside are disproportionately the ingestion tables where incidents start.

Is dbt itself changing this?

Somewhat, and it's worth tracking.

Fivetran completed its merger with dbt Labs on June 1, 2026. dbt Core v2.0 alpha open-sources the Fusion engine runtime under Apache 2.0, moving capabilities that were commercial into the free distribution.

The strategic read: the merged company owns both ingestion and transformation. Monitoring across that boundary, exactly the gap this post describes, is the obvious thing for them to build, and they now have both ends of the pipe. If that ships well, the "dbt tests stop at dbt's edge" argument weakens considerably.

It hasn't shipped, so plan for the warehouse you have. I'd just avoid making a five-year bet that this boundary stays where it is today.

What does each approach cost to run?

Cost gets left out of this comparison and it shouldn't be, because both sides have a real bill and they are shaped differently.

dbt tests cost warehouse compute, per run, forever. Every test is a query. A suite of 400 tests is 400 queries on every invocation, and the expensive ones are relationships tests, which join two tables in full. On a nightly schedule that's 12,000 queries a month you are paying for whether or not any of them ever fail.

Put rough numbers on it. If your test suite adds 8 minutes to a nightly build on a medium warehouse, that's roughly 4 hours of compute a month for the tests alone. Whether that's $50 or $400 depends on your platform and cluster size, but it is not zero and almost nobody measures it.

The uncomfortable part is the distribution. If 360 of those 400 tests are not_null on columns that have never once been null, you are paying 90% of that bill for tests that have never fired and, on current evidence, never will.

Monitoring costs a subscription plus lighter queries. Freshness and row-count checks are metadata queries or cheap aggregates rather than full scans, and they run on their own schedule instead of inside your build. The subscription is the visible cost; the compute is usually smaller than people expect and the build time is unaffected.

dbt tests Monitoring
Direct cost $0 license Subscription
Compute Every test, every run, full scans on relationships Cheap aggregates on a schedule
Build time Adds minutes to every invocation None, runs outside the build
Marginal cost per new table Whatever tests you write for it Per-table price or included
Cost when nothing is wrong Same as when something is wrong Same

The practical move is to audit rather than to switch. Find the 10 slowest tests in your suite, check when each last failed, and delete the ones that never have and never could. That is usually worth more than any tooling decision in this post, and it takes an afternoon.

How should you actually split the work?

The split that works, in my experience, is by what kind of knowledge the check encodes.

Write a dbt test when the check encodes something a human knows and a machine cannot infer. Valid status values. Foreign key relationships. A business invariant like "refunds are never positive." These are irreplaceable and there are fewer of them than your project probably has.

Use monitoring for anything defined relative to normal. Freshness, volume, null rates, distributions, schema changes. These need history, which is the thing tests don't have, and writing them by hand means hardcoding a threshold that is wrong within a quarter.

Practical consequence: most teams should have fewer dbt tests than they do. A project with 400 tests usually has 40 that encode knowledge and 360 that are not_null on columns that have never been null, running nightly, costing warehouse compute and attention. Deleting those is not a loss of coverage if something else watches the same tables continuously.

The uncomfortable version: a large test suite can be worse than a small one, because it produces enough noise that people stop reading the output, and then the 40 that mattered fail silently alongside the 360 that didn't.

What about the third option nobody mentions?

There is a middle path between "write an assertion" and "learn a baseline," and it gets skipped because it doesn't belong to either camp: the contract.

A dbt test asserts something about a model you own. A contract asserts something about an interface between two teams, and it is enforced at the boundary rather than inside your project. The difference matters when the thing breaking is somebody else's output.

The concrete version: your accepted_values test on status fires at 02:00, hours after the upstream team deployed the change that introduced a fourth status. The test worked exactly as designed and you still found out late, because the test lives downstream of the problem. A contract on that interface fails in the upstream team's pipeline, at the moment they try to ship it, which is when it is cheap to fix.

# an ODCS contract fragment: the interface, not the model
schema:
  - name: orders
    properties:
      - name: status
        logicalType: string
        quality:
          - rule: validValues
            mustBe: ['pending', 'shipped', 'delivered']
      - name: order_id
        logicalType: string
        required: true
        unique: true

Same three assertions as the dbt test at the top of this post. Different owner, different enforcement point, different failure time.

I'm not going to pretend contracts are free. They require an upstream team that agrees to be bound by one, which is an organizational fact rather than a technical one, and most of the reason contracts get discussed more than they get adopted. If the team producing your source data is a vendor, or another company's API, there is no contract to negotiate.

Where they work: two internal teams, one producing and one consuming, with a history of breaking each other. Where they don't: everywhere the producer has no incentive to care. Worth knowing the option exists, because a lot of teams reach for more dbt tests when what they actually have is an interface problem.

What should you check this week?

Four things, in order, none taking more than an hour.

  1. Grep for severity: warn. Every one is a test you believe you have and don't. Decide for each: promote to error, add an error_if threshold, or delete it.
  2. Count non-dbt tables. The ratio tells you how much of your warehouse your test suite can possibly cover. If it's 60%, no amount of test writing fixes the other 40%.
  3. Check whether dbt source freshness runs at all. Many projects declare sources and never run the command, which is coverage on paper only.
  4. Look at your last three incidents. For each, ask whether a dbt test could plausibly have caught it. If the answer is repeatedly no, more tests is the wrong investment.

One caveat on step 4, because it can be read as an argument against testing generally. The incidents you remember are the ones that got through. Your dbt tests have almost certainly caught things that never became incidents precisely because the build went red and somebody fixed it before anyone downstream noticed. That work is invisible by construction. So the question is not "have tests prevented anything," it is "are the failures that reach production of a kind tests could catch." Those are different questions and only the second one tells you where to invest.

That last one is the whole post compressed. Teams reach for more tests because tests are the tool they know. If the incidents are freshness, schema drift, and upstream units changing, then writing tests is effort spent on the boundary you already cover well, and the gap stays exactly where it was.

Frequently asked questions

Do dbt tests replace data observability?

No. dbt tests check assertions you wrote, when dbt runs, on tables dbt built. Observability watches tables continuously regardless of who wrote a rule or what produced the table. They cover different failure classes and the overlap is small. The category distinction is covered in data observability vs data quality.

Does data observability replace dbt tests?

Also no. Business rules like valid status values or foreign key relationships encode human knowledge that no baseline can infer. Those belong in tests and always will.

How many dbt tests should I have?

Fewer than you probably do. Keep the ones encoding knowledge a machine can't derive, and let monitoring handle anything defined relative to normal. A suite of 400 tests where 360 are not_null on never-null columns is cost without coverage.

What is severity warn in dbt and why is it a problem?

severity: warn lets a failing test emit a log line without failing the build. It looks like coverage and behaves like nothing, because a warning in CI output is not an alert. Audit your project for it.

What is error_if in a dbt test?

A threshold that fails the build only past a certain number of violating rows, usually paired with warn_if for smaller counts. It's the middle ground between a blanket warn and a build that fails on a single stale row.

Can dbt tests catch freshness problems?

Only via dbt source freshness, only on sources you explicitly declared, and only when you run it. It's a build-time check on a schedule rather than a clock, so a source that breaks after the run is invisible until the next one.

Why do my dbt tests pass when the data is wrong?

Because most tests check shape rather than meaning. Stale data has the right shape. A currency field switching from dollars to cents is still numeric, still not null, still positive. Catching that needs a baseline, not an assertion.

Should I test raw source tables?

You can, once they're declared as dbt sources, and freshness checks there are worth having. It doesn't extend to tables outside dbt, and declaring sources is a chore that decays, so the covered set drifts below the real set over time.

Does the Fivetran and dbt Labs merger change this?

Not yet. The merged company owns both ingestion and transformation, so monitoring across that boundary is a natural thing for them to build, and dbt Core v2.0 already open-sourced the Fusion runtime under Apache 2.0. Plan for the warehouse you have rather than the roadmap.

What is the cheapest way to close the gap?

Start with schema and freshness monitoring on your ingestion tables, since that's where most incidents originate and where dbt tests have the least reach. Volume and distribution monitoring is the next layer.

Do I need lineage to make this work?

Not to detect problems, but to triage them. When one source breaks and thirty downstream tables go stale, lineage is what collapses that into one incident with a named cause. Uploading a dbt manifest is the cheap way to get it.

What is the difference between a dbt test and a data contract?

A dbt test asserts something about a model you own, and it fails in your pipeline after the bad data has landed. A contract asserts something about an interface between two teams and fails in the producer's pipeline, before the data ships. Same assertions, different enforcement point.

Do dbt tests cost money to run?

Yes, in warehouse compute, on every invocation. relationships tests are the expensive ones because they join two tables in full. A large suite running nightly is a real recurring bill that almost nobody measures.

How do I know my test suite is actually protecting anything?

Break something on purpose in a staging environment and see whether the pipeline notices, and how loudly. A test suite nobody has ever seen fail correctly is an untested test suite.