What Is Data Observability — And Why Monitoring Isn't Enough


In June 2023, the city health department of a major U.S. metro submitted its quarterly syndromic surveillance data to the CDC — on time, with zero pipeline errors in the logs. Six weeks later, the CDC flagged a 22% anomaly in respiratory illness counts for the reporting period. The culprit: a source EMR vendor had silently changed a field encoding, causing a category of cases to map to a null bucket rather than the correct diagnosis group. Every pipeline monitoring check had passed. No alerts had fired. The data was wrong for six weeks before anyone caught it.
This is the monitoring gap. And it’s exactly what data observability is designed to close.
Data monitoring and data observability are not synonyms, though they are used as if they are. The confusion is costly. Monitoring watches your pipelines. Observability watches your data. The difference determines whether you find out about a problem from your pipeline logs or from your executive team three weeks after a board presentation.
This article defines each precisely, maps the five pillars of data observability against what monitoring tools actually catch, gives a real failure pattern for every gap, and ends with a practical implementation path for teams that want both.
What Data Monitoring Actually Does — And What It Misses
Data monitoring is infrastructure observability applied to data pipelines. It watches for process failure: did the job run? Did it finish in an acceptable time window? Did it error out? Did the connection drop? These are operational questions, and monitoring tools — whether that’s Airflow alerts, Datadog pipeline checks, dbt test pass/fail counts, or CloudWatch metrics — answer them well.
Here is exactly what monitoring catches:
Pipeline job failures — a task that errored, timed out, or was retried beyond threshold
SLA breaches — a job that ran longer than expected or finished outside the delivery window
Schema breaking changes — when a downstream model fails because an upstream field was dropped or renamed
Row count anomalies (crude) — when a table that usually loads 50,000 rows loads zero, most monitoring setups will alert
Infrastructure health — compute, memory, disk, connection pool saturation
Here is what monitoring does not catch:
A pipeline that runs successfully but loads subtly wrong data
A field that switches from meaningful values to placeholder defaults because a source system changed its export logic
A provider that stops submitting, causing their patient population to silently disappear from aggregates
A distribution shift — the age distribution of your patient cohort drifting by 8 years over 90 days because a new EMR integration skews toward older patients
A lineage break — a dashboard pulling from a table that is now two joins removed from the certified source, with a transformation in between that nobody documented
The core limitation is architectural: monitoring assumes a correctly running pipeline produces correct data. Observability does not make that assumption. It treats data correctness as a separate question from pipeline correctness, requiring separate instrumentation.
The 5 Pillars of Data Observability — And the Gap Monitoring Leaves at Each One
Data observability frameworks have converged on five dimensions that must be continuously measured to actually know whether your data is trustworthy. Each one maps directly to a class of failure that monitoring alone cannot surface.
Pillar 1: Volume
What it measures: Whether the expected amount of data arrived — not just whether any data arrived.
What monitoring catches: A table that loaded zero rows when 50,000 were expected. Most orchestration tools will flag this.
What monitoring misses: A table that loaded 47,200 rows when 50,000 were expected. That’s a 5.6% shortfall — statistically significant for a health surveillance cohort, invisible to a monitoring alert that only fires at zero.
Real failure pattern: A state Medicaid agency ingests weekly claims from 14 regional payers. One payer migrates to a new claims clearinghouse mid-quarter and their file drops to 60% of normal volume. The pipeline runs successfully. No alert fires. The undercounting propagates into the monthly PMPM report presented to leadership. It surfaces during a CMS data validation review three months later.
What observability adds: Statistical volume baselines per source, per time period, per data entity type. Alerts when observed volume falls outside two standard deviations of the historical window. Volume checks at the source level, not just the aggregate.
Pillar 2: Freshness
What it measures: Whether data is as current as it should be — whether it arrived on schedule and reflects the expected time horizon.
What monitoring catches: A job that did not run at all (SLA miss).
What monitoring misses: A job that ran on schedule but loaded data that was already 72 hours stale because the upstream source system had a silent backlog.
Real failure pattern: A hospital system’s ADT feed (Admissions, Discharges, Transfers) runs nightly. The HL7 interface engine develops a queue backlog after a configuration change. Messages are still delivered, but with a 3-day lag. The pipeline shows green. Clinicians pulling the “current census” report are looking at patient status that is 3 days old. Nobody knows until a nurse notices a patient listed as admitted who was discharged two days ago.
What observability adds: Maximum-timestamp monitoring on every data entity (the most recent updated_at or event_date in the loaded data). If the newest record is older than the expected freshness window — 2 hours for ADT, 24 hours for claims, 7 days for provider directories — an alert fires. Pipeline ran successfully is not the same as data is fresh.
Pillar 3: Schema
What it measures: Whether the structure of your data — fields, types, cardinality, relationships — is what downstream consumers expect.
What monitoring catches: A hard schema break — a field dropped or renamed that causes a downstream dbt model or SQL query to error out.
What monitoring misses: Soft schema drift — a field that changes data type silently (integer to string, or a date field that starts receiving null for new records while historical records retain values). These changes do not break pipelines. They silently corrupt downstream models.
Real failure pattern: A provider credentialing system updates its API. The npi field, previously always a 10-digit string, now returns null for providers registered after the API update date while still returning values for previously-registered providers. No schema validation fires because the column still exists and the type is still varchar. 8 months of new provider records load with null NPIs. This is only discovered during a network adequacy audit when regulators ask for a provider count by specialty.
What observability adds: Column-level null rate tracking over time. Cardinality monitoring (the number of distinct values in a column). Type distribution checks. Field-level change detection that fires when a column’s statistical profile shifts meaningfully, even when the column itself doesn’t disappear.
Pillar 4: Distribution
What it measures: Whether the statistical distribution of values in key fields — age, diagnosis category, geographic region, claim amount, test result — is consistent with what you’d expect based on historical patterns.
What monitoring catches: Nothing. Distribution is entirely outside the scope of standard monitoring tools.
What monitoring misses: Everything in this category. A pipeline can run perfectly and deliver data whose value distributions have shifted in ways that invalidate every downstream analysis that assumes historical distributions.
Real failure pattern: A behavioral health program integrates a new EHR vendor that serves a rural hospital network. The new population is older, with different comorbidity profiles, and predominantly uninsured. Their records enter the main data warehouse without any distribution flags. Six months of trend analyses now show a shift in program demographics that is real — because the population genuinely changed — but untagged, causing a state health equity report to misattribute the demographic shift to program success rather than integration expansion. The correction requires rerunning six months of analyses.
What observability adds: Automated statistical profiling of key fields on every pipeline run. Z-score alerts when a metric’s distribution shifts beyond a configurable threshold. Cohort-level distribution monitoring (new vs. returning patients, by payer type, by region) that separates structural changes from data quality changes.
Pillar 5: Lineage
What it measures: The complete provenance of every piece of data — where it came from, what transformations it passed through, which downstream reports and dashboards consume it.
What monitoring catches: A job that failed. The job’s dependencies (if the orchestrator tracks them).
What monitoring misses: The full downstream impact of a data change or a data quality issue. Which of your 47 dashboards are affected when a source field changes? Which reports used data that was later found to be incorrect? Who needs to be notified?
Real failure pattern: A data team fixes a calculation error in a claims cost model. The fix is applied to the source table. Three weeks later, an analyst notices that a KPI dashboard used by the finance team is showing numbers inconsistent with the corrected model. Investigation reveals the dashboard pulls from a materialized view that was not in any documented lineage, created by a former analyst, refreshed daily by a scheduled job nobody documented. The incorrect numbers had been in that dashboard for the full three weeks after the fix.
What observability adds: Automated lineage tracking that maps every query, every transformation, every dashboard connection to its source. When a table changes, every downstream consumer is identified automatically. When a data quality issue is found, the blast radius is known immediately — not discovered three weeks later by a finance analyst.
Observability vs Monitoring: Side-by-Side
Dimension | Data Monitoring | Data Observability |
Primary question | Did the pipeline run? | Is the data correct? |
What it watches | Jobs, connections, SLAs | Volume, freshness, schema, distribution, lineage |
When it alerts | On process failure | On data quality deviation |
What it misses | Subtle data drift, silent errors | Infrastructure-level failures (needs both) |
Tooling examples | Airflow, Datadog, dbt tests, CloudWatch | Monte Carlo, Acceldata, Vexdata, Great Expectations |
Alert mechanism | Binary (pass/fail) | Statistical (deviation from baseline) |
Who owns it | Data engineering / DevOps | Data engineering + data quality / analytics |
Blast radius visibility | Job dependencies only | Full downstream lineage |
Distribution detection | None | Yes — z-score / threshold-based |
Freshness at data level | Rarely | Yes — max timestamp per entity |
How to Implement: A Practical Stack
The most common question data teams ask after understanding the observability gap is: do we have to replace our monitoring tooling? The answer is no. Observability and monitoring are complementary, not competing. The practical implementation path layers them.
Layer 1 — Keep your existing monitoring. Airflow, dbt tests, and infrastructure alerts handle process correctness. Don’t remove them. Add to them.
Layer 2 — Instrument volume and freshness baselines. Before adding a dedicated observability tool, instrument the simplest checks in your existing SQL or dbt layer: row counts by source per run, max timestamp on ingested data vs. expected freshness window. These two checks alone catch the majority of silent data quality failures.
-- Volume baseline check
SELECT
source_system,
DATE(load_ts) AS load_date,
COUNT(*) AS row_count,
AVG(COUNT(*)) OVER (PARTITION BY source_system ORDER BY DATE(load_ts) ROWS 30 PRECEDING) AS rolling_30d_avg
FROM raw.events
GROUP BY 1, 2
HAVING row_count < rolling_30d_avg * 0.85 -- alert if more than 15% below average-- Freshness check
SELECT
source_system,
MAX(event_timestamp) AS most_recent_record,
CURRENT_TIMESTAMP - MAX(event_timestamp) AS data_lag
FROM raw.events
GROUP BY 1
HAVING data_lag > INTERVAL '26 hours' -- alert if data is more than 26 hours staleLayer 3 — Add schema drift detection. Track column-level null rates and cardinality on a per-run basis. Store them. Alert when they shift. This can be implemented in dbt as a custom generic test or handled by a dedicated observability platform.
Layer 4 — Distribution monitoring for critical fields. Prioritize the fields that feed executive reporting, federal submissions, or clinical decisions. Profile their distributions weekly. Flag outlier shifts.
Layer 5 — Lineage. This is the hardest to implement manually and the strongest argument for a dedicated observability platform. Automated lineage requires intercepting queries at the compute layer (BigQuery, Snowflake, Redshift job history APIs) and building the graph from actual execution, not documentation.
Vexdata’s pipeline validation layer sits at Layers 2–4: automated volume, freshness, schema, and distribution checks run on every pipeline execution, with configurable thresholds and an alert routing layer that sends results to Slack, email, or your incident management platform. Teams using Vexdata alongside existing Airflow or dbt monitoring have typically reduced manual pre-submission validation from 2–3 weeks to under 48 hours.
Conclusion
Data monitoring answers the question your pipeline tooling is built to answer: did the job run correctly? Data observability answers the question your stakeholders are actually asking: is this data right?
The five pillars — volume, freshness, schema, distribution, and lineage — define the complete surface of data trustworthiness. None of them are fully covered by monitoring alone. The organizations that have shifted from reactive data quality (find the problem after the board meeting) to proactive data quality (catch it before the pipeline run completes) have implemented all five, in sequence, with tooling matched to their current stack.
You do not need to rebuild your monitoring infrastructure. You need to add the observability layer that sits above it.
Ready to see where your pipeline’s blind spots are? Take the Vexdata Data Fitness Assessment — a 10-minute evaluation that maps your current validation coverage against the five pillars and identifies the highest-risk gaps in your pipeline.




Comments