How Much of Your Attribution Data Is Modeled vs Observed?

· 13 min read

Every conversion column stacks two things into one number: orders the system observed end-to-end and orders it estimated with a model. Vendors do not tell you the mix. Coverage is the fraction observed before any model touched it: observed orders divided by real orders from your system of record. Below roughly 70% coverage, arguing about attribution models is arguing about how to slice a pie that is missing a third of its filling. The fix at that point is collection, not model choice. Ask any vendor: what percentage of these conversions was observed end-to-end, can you show observed and modeled in separate columns, and what is your eligibility threshold for modeling.

One column, two very different things

Every attribution report you read presents a single conversion number as if it were one kind of thing. It isn't. It is a blend of two things stacked into one column: conversions the system actually watched happen, where a real user, a real session, and a real order got tied together end to end, and conversions it estimated using a model trained on the users it could see.

The vendor does not tell you the mix. The column looks identical whether it is 95% observed or 45% observed. A "1,240 conversions" cell in your dashboard could be 1,240 real stitched journeys or it could be 600 real ones and 640 the model filled in. Same cell, same font, same confidence.

So the buyer question that actually matters is not "which attribution model should I use?" It is what fraction of your orders the system observed with a complete journey, before any model touched the number. Call that fraction coverage.

THE METRIC

Coverage % = orders observed with a full click-to-conversion journey ÷ total actual orders you fulfilled.

The denominator is the thing no attribution platform has and you do: your real order count from the system of record (Shopify, Stripe, your billing database, your CRM's closed-won). The numerator is what the attribution system could genuinely see and stitch, stripped of modeled fill-in.

The argument has a hard operational edge, and it is the reason to compute this at all. Below roughly 70% coverage, model choice is noise. If a third of your orders never entered the observed set, then last-touch versus data-driven versus Markov is a debate about how to slice a pie that is missing a third of its filling. You are re-weighting credit across journeys the system never saw. The fix at that point is better collection, not a better model: server-side capture, identity resolution, first-party conversion feeds. Fix the denominator gap first, then argue about models.

Here is the uncomfortable part. You cannot buy a coverage number today, because no vendor publishes one. You have to compute it yourself, and the arithmetic is embarrassing for most stacks.

The vendors combine the two on purpose

This is how the dominant tools are built, not an accident of your setup.

Google's own Ads documentation states plainly that "in the 'Conversions' column, Google reports both modeled and observed conversions," and that modeled conversions are reported "with the same granularity as observed conversions." There is no standard-report mechanism to pull the two apart. The blend is the product.

GA4 goes one step further into unhelpfulness. Its behavioral-modeling documentation says Analytics "seamlessly integrates modeled data and observed data in your reports" and offers only a data-quality icon so you can "see when modeled data is integrated." That icon tells you when modeling is present. It never tells you how much. There is no accuracy figure and no way to quantify the modeled share of any metric. You get a flag, not a fraction.

You have probably heard the number 70% attached to modeling accuracy. It is worth correcting the record, because it is almost always mis-cited. The figure comes from a 2021 Google Ads blog post, which said conversion modeling through Consent Mode "recovers more than 70% of ad-click-to-conversion journeys lost due to user cookie consent choices." Three things about that claim get lost. It is specific to Google Ads Consent Mode modeling, not GA4. It is a recovery rate (how much of the lost signal it fills back), not an accuracy rate (how right each modeled conversion is). And GA4 behavioral modeling has no published accuracy number at all. So when someone tells you GA4 is "70% accurate," they are quoting a different product's recovery statistic and calling it precision.

WHAT THE VENDORS ACTUALLY TELL YOU

Question: can you see the modeled share of your conversion number?

WHAT THEY DISCLOSE
  • Google Ads: modeled + observed in one column
  • GA4: an icon showing when data is modeled
  • No published accuracy figure for GA4 modeling
WHAT YOU CAN'T GET
0
standard-report ways to read the modeled percentage of any given number. The split is not exposed.

The gap is bigger than "modeling," and some of it is silent

Modeling at least tries to fill the hole. The worse case is when the hole is left open.

Modeling only turns on above hard traffic thresholds. Google Ads consent-mode modeling requires a daily ad-click threshold of 700 ad clicks over a 7-day period, per country and domain grouping. GA4 behavioral modeling requires the property to collect at least 1,000 events per day with analytics_storage='denied' for at least 7 days, plus roughly a thousand daily consented users, sustained across the prior 28 days. <!-- VERIFY: cite the specific Google Ads Help and GA4 behavioral-modeling docs pages for these exact thresholds (700 clicks / 1,000 events / consented-user counts) before publish; Google has revised these numbers over time. --> Most bootstrapped and mid-market advertisers sit below these thresholds in at least some country and domain buckets. Below the threshold, the gap is not even modeled. It is just missing, and nothing on the dashboard tells you so.

There is a second, structural hole at the collection layer. On iOS, App Tracking Transparency opt-in rates have stabilized around 25% globally, which means roughly three-quarters of iOS users decline cross-app tracking. <!-- VERIFY: attribute the ~25% global ATT opt-in figure to a named source (e.g. AppsFlyer or Flurry/Adjust benchmark, with year) before publish. -->
(Reported e-commerce category figures run somewhat higher in some datasets, but they vary by vendor and year, so treat any specific category percentage as vendor-dependent. The directional claim holds either way.) A material fraction of mobile journeys never generates the deterministic signal attribution needs. That is a collection-layer hole, not a modeling problem, and no model choice recovers it.

Put those together and the observed set shrinks from three directions at once: consent-declined users who may or may not get modeled, iOS users who are structurally invisible, and cross-device journeys the platform never stitched. Coverage is what survives all three.

Three worked examples, three different companies

The arithmetic is the content here. The brand names are scaffolding. These are illustrative constructions, not case studies.

A WORKED EXAMPLE · NORTHLARK (DTC SKINCARE)

Northlark does 2,000 orders a month. Shopify, the system of record, says 2,000. GA4 and Google Ads together report 2,180 conversions across channels, the usual over-count from double-attribution and modeling. Now pull the observed set: sessions where GA4 stitched a full journey to a purchase event with analytics_storage='granted' and no modeling flag. That set is 1,240 orders.

Coverage = 1,240 / 2,000 = 62%. The other 38% is a mix of consent-declined users (behaviorally modeled), iOS ATT holes, and cross-device journeys the platform never saw.

Verdict: below 70%. Northlark should stop A/B-testing attribution models this quarter. The lever is server-side conversion capture plus first-party consent recovery. At 62% observed, a data-driven model is redistributing credit across the 62% and guessing at the rest.

A WORKED EXAMPLE · MERIDIAN B2B (SaaS, $40K ACV)

Meridian closes 48 deals a quarter, per Salesforce, on a 90-day sales cycle. Ad platforms and GA4 claim credit, in aggregate, for 71 conversions (multi-touch double-counting across Google, LinkedIn, and organic). The observed-end-to-end set, a tracked lead that maps to a specific closed-won account with an intact touch history, is 19 deals.

Coverage = 19 / 48 = 40%. The gap here is the offline conversion boundary, not ATT. The purchase happens in Salesforce, weeks after the last trackable click, on a different device, often via a buying committee where the researcher and the signer are different people.

Verdict: at 40% coverage, per-channel MTA credit for a B2B pipeline is close to fiction. The fix is offline conversion import and server-side CRM-to-attribution feeds, not a new weighting model. Until deals flow back into the observed set keyed on account, Meridian is modeling a two-fifths sample and calling it channel performance.

Now the company for whom model choice is a legitimate, high-value debate.

HARBOR GOODS · HIGH-VOLUME MARKETPLACE

Setup: 60,000 orders a month, mostly desktop web, strong consent rates.

THE NUMBERS
  • System of record: 60,000
  • Observed end-to-end: 51,600
  • Coverage = 86%
VERDICT
86%
Above 70%. This is the buyer for whom last-touch versus an algorithmic model actually changes spend decisions, because it re-weights real journeys instead of papering over missing ones.

Harbor should invest in modeling sophistication. Northlark and Meridian should not, yet. Coverage tells you which company you are.

Computing your own coverage

The vendors will not hand you the number, so you reconcile against the one source they cannot touch: your system of record.

Take total real orders for a fixed window from Shopify, Stripe, or your CRM. Then count only the conversions your attribution system observed with a complete, consented, single-identity journey. Where a platform does expose a modeled flag (Google's data-quality icon, Meta's modeled-events labels), treat that disclosure as a floor on the modeled share, not the whole of it.

sql
-- Coverage = observed orders / real orders, per week. -- orders = your system of record (Shopify / Stripe / CRM closed-won) -- attribution_events = your attribution system's conversion records -- is_modeled / consent flags are your best available approximations; -- where the platform won't expose them, treat modeled share as unknown-and-present. SELECT date_trunc('week', o.ordered_at) AS week, COUNT(DISTINCT o.order_id) AS real_orders, COUNT(DISTINCT ae.order_id) FILTER ( WHERE ae.journey_complete AND ae.consent = 'granted' AND NOT ae.is_modeled AND ae.single_identity ) AS observed_orders, ROUND( COUNT(DISTINCT ae.order_id) FILTER ( WHERE ae.journey_complete AND ae.consent = 'granted' AND NOT ae.is_modeled AND ae.single_identity )::numeric / NULLIF(COUNT(DISTINCT o.order_id), 0), 3 ) AS coverage FROM orders o LEFT JOIN attribution_events ae ON ae.order_id = o.order_id GROUP BY 1 ORDER BY 1 DESC;

This is estimation, not a clean readout, and it should be. The whole point is that the vendors make it estimation by refusing to publish the split. If your coverage lands above 70% and holds week to week, model work is worth doing. If it sits below, the SQL just told you where the money is: collection.

What coverage is not

This audience is senior, so the honest limits matter as much as the metric.

Coverage is not accuracy, and observed is not true. A 90%-coverage number can still be wrong about causality. The largest field study on this settles it: 15 randomized advertising experiments at Facebook, roughly 500 million user-experiment observations and 1.6 billion impressions, found that common observational measurement approaches, the family that includes last-touch and multi-touch attribution, often fail to accurately measure the true effect of advertising relative to randomized experiments (Gordon, Zettelmeyer, Bhargava, and Chapsky, Marketing Science 38(2):193-225, 2019). High coverage means the model is working from real journeys instead of fabricated ones. It does not mean the credit assignment is causally correct. MTA measures correlation across observed touchpoints. Incrementality requires a holdout or geo experiment.

The 70% line is a heuristic, not a constant. Defend it as a decision rule: below about 70%, the modeled-plus-missing fraction is large enough that it likely dominates the delta between attribution models, so collection improvements have higher expected value than model improvements. The exact cut depends on how the missing data is distributed. If the unobserved 40% is a random sample of your orders, models degrade gracefully. If it is systematically one channel or one device class, which it usually is (mobile, paid social, cross-device), then even 80% coverage can mislead badly. Coverage is necessary but not sufficient. The distribution of the gap matters as much as its size.

And coverage does not fix over-counting. It addresses under-observation, the journeys the system missed. It does nothing about the separate problem of multiple platforms each claiming the same order. A brand can simultaneously have 62% coverage and 2,180 reported conversions against 2,000 real orders. Those are two different distortions. Name both, fix both, and do not confuse one for the other. (For the over-counting side of the ledger, see count your conversions before crediting channels.)

Three questions to ask any vendor

Whether you are auditing your current stack or evaluating a new one, these three questions do most of the work. If a vendor cannot answer the first with a number, or the second with a yes, that is itself the finding.

QUESTION 1
The number
What percentage of the conversions in this report were observed end-to-end, versus modeled or estimated?
QUESTION 2
The separation
Show me observed and modeled in separate columns. Can your product even do that?
QUESTION 3
The threshold
What is your eligibility threshold for modeling to kick in, and what happens to my numbers on the days I fall below it?

A tool that can show observed and modeled separately, on data that was collected server-side so the observed set is as large as it can honestly be, hands you a coverage number instead of hiding one. That is the difference between reasoning about your own data and accepting a blend you cannot decompose. The model debate is real and worth having. It is just the second debate, and only for the companies whose coverage earns it.

Key Takeaways

  • A conversion column blends observed conversions (a real journey the system stitched) with modeled ones (estimated), and vendors do not publish the split
  • Coverage % = orders observed with a full journey / total real orders from your system of record (Shopify, Stripe, CRM)
  • Below about 70% coverage, model choice is noise: you are re-weighting credit across journeys the system never saw
  • GA4's data-quality icon tells you when data is modeled, never how much, and there is no published accuracy figure for GA4 behavioral modeling
  • High coverage makes attribution less fictional, not causal: even fully observed MTA measures correlation, not incremental lift
How do I calculate my coverage number if my attribution tool won't tell me what's modeled?
Reconcile against your system of record. Take total real orders from Shopify, Stripe, or your CRM for a fixed window. Then count only the conversions your attribution system observed with a complete, consented, single-identity journey. In GA4, exclude anything flagged by the data-quality or modeling indicator. In Google Ads, treat the modeled portion as present-but-unknown. Coverage equals observed divided by real orders. It is an estimate, and the fact that you are forced to estimate it is the point.
Is 70% a real threshold or just a rule of thumb?
A decision heuristic, not a law of nature. Below it, the missing-plus-modeled fraction is usually large enough to swamp the difference between attribution models, so collection work has higher expected value than model work. How the gap is distributed across channels and devices matters as much as its size. Treat 70% as 'stop and check collection first,' not as a pass or fail line.
If Google already models the missing conversions, why should I care about coverage?
Because you cannot see how much is modeled, cannot verify the model's accuracy (Google publishes none for GA4 behavioral modeling), and cannot turn it off. Below the eligibility thresholds, modeling may not run at all, leaving silent gaps. And modeled conversions are estimates layered on top of your real ones inside a single number that looks observed. You are making budget decisions on a blend you cannot decompose.
Does higher coverage mean my attribution is finally accurate?
No. Higher coverage means the model is reasoning from real journeys instead of fabricated ones, which is a necessary condition for trustworthy attribution, not a sufficient one. Even fully observed attribution measures correlation across touchpoints, not causal lift. For incremental truth you still need holdout or geo experiments. Coverage makes your attribution less fictional. It does not make it causal.
What do I actually do if my coverage is below 70%?
Fix collection before touching models. In order of usual impact: move conversion capture server-side so you are not dependent on the browser and consent layer; wire your system of record back into attribution (offline conversion import for B2B, first-party purchase events for DTC); implement identity resolution so cross-device and returning-user journeys stitch instead of fragmenting; recover consented first-party signal where you legitimately can. Then, and only then, is it worth arguing about which model to run.
Isn't modeled data better than dropping those conversions to zero?
Yes. Consent-mode and conversion modeling exist because deterministic tracking genuinely cannot see consent-declined and cross-device users, and a decent model beats counting those conversions as zero. The complaint is not that modeling is bad. It is that modeling you cannot see, cannot quantify, and cannot turn off, blended into a number labeled as if it were observed, defeats your ability to reason about your own data. The ask is transparency and separability, not abolition.
Holly Mehakovic
Holly Mehakovic

Co-Founder, mbuzz

Holly Mehakovic is Co-Founder of mbuzz. With 10+ years in marketing including roles at Westpac, Avon, and Forebrite, she's obsessed with making measurement actually useful.

Harvard Extension School Forebrite Westpac Avon

How mature is your marketing measurement?

The free Measurement Maturity Assessment shows where you stand, where you're exposed, and what to fix first. 10 questions, 3 minutes.

Take the Assessment

Ready to try server-side attribution?

Set up in 10 minutes. Free up to 30K records/month.