How Much of Your Attribution Data Is Modeled vs Observed?
Every conversion column stacks two things into one number: orders the system observed end-to-end and orders it estimated with a model. Vendors do not tell you the mix. Coverage is the fraction observed before any model touched it: observed orders divided by real orders from your system of record. Below roughly 70% coverage, arguing about attribution models is arguing about how to slice a pie that is missing a third of its filling. The fix at that point is collection, not model choice. Ask any vendor: what percentage of these conversions was observed end-to-end, can you show observed and modeled in separate columns, and what is your eligibility threshold for modeling.
One column, two very different things
Every attribution report you read presents a single conversion number as if it were one kind of thing. It isn't. It is a blend of two things stacked into one column: conversions the system actually watched happen, where a real user, a real session, and a real order got tied together end to end, and conversions it estimated using a model trained on the users it could see.
The vendor does not tell you the mix. The column looks identical whether it is 95% observed or 45% observed. A "1,240 conversions" cell in your dashboard could be 1,240 real stitched journeys or it could be 600 real ones and 640 the model filled in. Same cell, same font, same confidence.
So the buyer question that actually matters is not "which attribution model should I use?" It is what fraction of your orders the system observed with a complete journey, before any model touched the number. Call that fraction coverage.
THE METRIC
Coverage % = orders observed with a full click-to-conversion journey ÷ total actual orders you fulfilled.
The denominator is the thing no attribution platform has and you do: your real order count from the system of record (Shopify, Stripe, your billing database, your CRM's closed-won). The numerator is what the attribution system could genuinely see and stitch, stripped of modeled fill-in.
The argument has a hard operational edge, and it is the reason to compute this at all. Below roughly 70% coverage, model choice is noise. If a third of your orders never entered the observed set, then last-touch versus data-driven versus Markov is a debate about how to slice a pie that is missing a third of its filling. You are re-weighting credit across journeys the system never saw. The fix at that point is better collection, not a better model: server-side capture, identity resolution, first-party conversion feeds. Fix the denominator gap first, then argue about models.
Here is the uncomfortable part. You cannot buy a coverage number today, because no vendor publishes one. You have to compute it yourself, and the arithmetic is embarrassing for most stacks.
The vendors combine the two on purpose
This is how the dominant tools are built, not an accident of your setup.
Google's own Ads documentation states plainly that "in the 'Conversions' column, Google reports both modeled and observed conversions," and that modeled conversions are reported "with the same granularity as observed conversions." There is no standard-report mechanism to pull the two apart. The blend is the product.
GA4 goes one step further into unhelpfulness. Its behavioral-modeling documentation says Analytics "seamlessly integrates modeled data and observed data in your reports" and offers only a data-quality icon so you can "see when modeled data is integrated." That icon tells you when modeling is present. It never tells you how much. There is no accuracy figure and no way to quantify the modeled share of any metric. You get a flag, not a fraction.
You have probably heard the number 70% attached to modeling accuracy. It is worth correcting the record, because it is almost always mis-cited. The figure comes from a 2021 Google Ads blog post, which said conversion modeling through Consent Mode "recovers more than 70% of ad-click-to-conversion journeys lost due to user cookie consent choices." Three things about that claim get lost. It is specific to Google Ads Consent Mode modeling, not GA4. It is a recovery rate (how much of the lost signal it fills back), not an accuracy rate (how right each modeled conversion is). And GA4 behavioral modeling has no published accuracy number at all. So when someone tells you GA4 is "70% accurate," they are quoting a different product's recovery statistic and calling it precision.
WHAT THE VENDORS ACTUALLY TELL YOU
Question: can you see the modeled share of your conversion number?
- Google Ads: modeled + observed in one column
- GA4: an icon showing when data is modeled
- No published accuracy figure for GA4 modeling
The gap is bigger than "modeling," and some of it is silent
Modeling at least tries to fill the hole. The worse case is when the hole is left open.
Modeling only turns on above hard traffic thresholds. Google Ads consent-mode modeling requires a daily ad-click threshold of 700 ad clicks over a 7-day period, per country and domain grouping. GA4 behavioral modeling requires the property to collect at least 1,000 events per day with analytics_storage='denied' for at least 7 days, plus roughly a thousand daily consented users, sustained across the prior 28 days. <!-- VERIFY: cite the specific Google Ads Help and GA4 behavioral-modeling docs pages for these exact thresholds (700 clicks / 1,000 events / consented-user counts) before publish; Google has revised these numbers over time. --> Most bootstrapped and mid-market advertisers sit below these thresholds in at least some country and domain buckets. Below the threshold, the gap is not even modeled. It is just missing, and nothing on the dashboard tells you so.
There is a second, structural hole at the collection layer. On iOS, App Tracking Transparency opt-in rates have stabilized around 25% globally, which means roughly three-quarters of iOS users decline cross-app tracking. <!-- VERIFY: attribute the ~25% global ATT opt-in figure to a named source (e.g. AppsFlyer or Flurry/Adjust benchmark, with year) before publish. -->
(Reported e-commerce category figures run somewhat higher in some datasets, but they vary by vendor and year, so treat any specific category percentage as vendor-dependent. The directional claim holds either way.) A material fraction of mobile journeys never generates the deterministic signal attribution needs. That is a collection-layer hole, not a modeling problem, and no model choice recovers it.
Put those together and the observed set shrinks from three directions at once: consent-declined users who may or may not get modeled, iOS users who are structurally invisible, and cross-device journeys the platform never stitched. Coverage is what survives all three.
Three worked examples, three different companies
The arithmetic is the content here. The brand names are scaffolding. These are illustrative constructions, not case studies.
A WORKED EXAMPLE · NORTHLARK (DTC SKINCARE)
Northlark does 2,000 orders a month. Shopify, the system of record, says 2,000. GA4 and Google Ads together report 2,180 conversions across channels, the usual over-count from double-attribution and modeling. Now pull the observed set: sessions where GA4 stitched a full journey to a purchase event with analytics_storage='granted' and no modeling flag. That set is 1,240 orders.
Coverage = 1,240 / 2,000 = 62%. The other 38% is a mix of consent-declined users (behaviorally modeled), iOS ATT holes, and cross-device journeys the platform never saw.
Verdict: below 70%. Northlark should stop A/B-testing attribution models this quarter. The lever is server-side conversion capture plus first-party consent recovery. At 62% observed, a data-driven model is redistributing credit across the 62% and guessing at the rest.
A WORKED EXAMPLE · MERIDIAN B2B (SaaS, $40K ACV)
Meridian closes 48 deals a quarter, per Salesforce, on a 90-day sales cycle. Ad platforms and GA4 claim credit, in aggregate, for 71 conversions (multi-touch double-counting across Google, LinkedIn, and organic). The observed-end-to-end set, a tracked lead that maps to a specific closed-won account with an intact touch history, is 19 deals.
Coverage = 19 / 48 = 40%. The gap here is the offline conversion boundary, not ATT. The purchase happens in Salesforce, weeks after the last trackable click, on a different device, often via a buying committee where the researcher and the signer are different people.
Verdict: at 40% coverage, per-channel MTA credit for a B2B pipeline is close to fiction. The fix is offline conversion import and server-side CRM-to-attribution feeds, not a new weighting model. Until deals flow back into the observed set keyed on account, Meridian is modeling a two-fifths sample and calling it channel performance.
Now the company for whom model choice is a legitimate, high-value debate.
HARBOR GOODS · HIGH-VOLUME MARKETPLACE
Setup: 60,000 orders a month, mostly desktop web, strong consent rates.
- System of record: 60,000
- Observed end-to-end: 51,600
- Coverage = 86%
Harbor should invest in modeling sophistication. Northlark and Meridian should not, yet. Coverage tells you which company you are.
Computing your own coverage
The vendors will not hand you the number, so you reconcile against the one source they cannot touch: your system of record.
Take total real orders for a fixed window from Shopify, Stripe, or your CRM. Then count only the conversions your attribution system observed with a complete, consented, single-identity journey. Where a platform does expose a modeled flag (Google's data-quality icon, Meta's modeled-events labels), treat that disclosure as a floor on the modeled share, not the whole of it.
-- Coverage = observed orders / real orders, per week.
-- orders = your system of record (Shopify / Stripe / CRM closed-won)
-- attribution_events = your attribution system's conversion records
-- is_modeled / consent flags are your best available approximations;
-- where the platform won't expose them, treat modeled share as unknown-and-present.
SELECT
date_trunc('week', o.ordered_at) AS week,
COUNT(DISTINCT o.order_id) AS real_orders,
COUNT(DISTINCT ae.order_id) FILTER (
WHERE ae.journey_complete
AND ae.consent = 'granted'
AND NOT ae.is_modeled
AND ae.single_identity
) AS observed_orders,
ROUND(
COUNT(DISTINCT ae.order_id) FILTER (
WHERE ae.journey_complete
AND ae.consent = 'granted'
AND NOT ae.is_modeled
AND ae.single_identity
)::numeric
/ NULLIF(COUNT(DISTINCT o.order_id), 0),
3
) AS coverage
FROM orders o
LEFT JOIN attribution_events ae
ON ae.order_id = o.order_id
GROUP BY 1
ORDER BY 1 DESC;
This is estimation, not a clean readout, and it should be. The whole point is that the vendors make it estimation by refusing to publish the split. If your coverage lands above 70% and holds week to week, model work is worth doing. If it sits below, the SQL just told you where the money is: collection.
What coverage is not
This audience is senior, so the honest limits matter as much as the metric.
Coverage is not accuracy, and observed is not true. A 90%-coverage number can still be wrong about causality. The largest field study on this settles it: 15 randomized advertising experiments at Facebook, roughly 500 million user-experiment observations and 1.6 billion impressions, found that common observational measurement approaches, the family that includes last-touch and multi-touch attribution, often fail to accurately measure the true effect of advertising relative to randomized experiments (Gordon, Zettelmeyer, Bhargava, and Chapsky, Marketing Science 38(2):193-225, 2019). High coverage means the model is working from real journeys instead of fabricated ones. It does not mean the credit assignment is causally correct. MTA measures correlation across observed touchpoints. Incrementality requires a holdout or geo experiment.
The 70% line is a heuristic, not a constant. Defend it as a decision rule: below about 70%, the modeled-plus-missing fraction is large enough that it likely dominates the delta between attribution models, so collection improvements have higher expected value than model improvements. The exact cut depends on how the missing data is distributed. If the unobserved 40% is a random sample of your orders, models degrade gracefully. If it is systematically one channel or one device class, which it usually is (mobile, paid social, cross-device), then even 80% coverage can mislead badly. Coverage is necessary but not sufficient. The distribution of the gap matters as much as its size.
And coverage does not fix over-counting. It addresses under-observation, the journeys the system missed. It does nothing about the separate problem of multiple platforms each claiming the same order. A brand can simultaneously have 62% coverage and 2,180 reported conversions against 2,000 real orders. Those are two different distortions. Name both, fix both, and do not confuse one for the other. (For the over-counting side of the ledger, see count your conversions before crediting channels.)
Three questions to ask any vendor
Whether you are auditing your current stack or evaluating a new one, these three questions do most of the work. If a vendor cannot answer the first with a number, or the second with a yes, that is itself the finding.
A tool that can show observed and modeled separately, on data that was collected server-side so the observed set is as large as it can honestly be, hands you a coverage number instead of hiding one. That is the difference between reasoning about your own data and accepting a blend you cannot decompose. The model debate is real and worth having. It is just the second debate, and only for the companies whose coverage earns it.
Key Takeaways
- ✓A conversion column blends observed conversions (a real journey the system stitched) with modeled ones (estimated), and vendors do not publish the split
- ✓Coverage % = orders observed with a full journey / total real orders from your system of record (Shopify, Stripe, CRM)
- ✓Below about 70% coverage, model choice is noise: you are re-weighting credit across journeys the system never saw
- ✓GA4's data-quality icon tells you when data is modeled, never how much, and there is no published accuracy figure for GA4 behavioral modeling
- ✓High coverage makes attribution less fictional, not causal: even fully observed MTA measures correlation, not incremental lift
How do I calculate my coverage number if my attribution tool won't tell me what's modeled?▼
Is 70% a real threshold or just a rule of thumb?▼
If Google already models the missing conversions, why should I care about coverage?▼
Does higher coverage mean my attribution is finally accurate?▼
What do I actually do if my coverage is below 70%?▼
Isn't modeled data better than dropping those conversions to zero?▼
How mature is your marketing measurement?
The free Measurement Maturity Assessment shows where you stand, where you're exposed, and what to fix first. 10 questions, 3 minutes.
Take the AssessmentReady to try server-side attribution?
Set up in 10 minutes. Free up to 30K records/month.