Questions to Ask an Attribution Vendor (+ Scorecard)
Most vendor-evaluation content is a feature checklist. This is the opposite: five questions you drive during the demo that expose where the numbers come from and whether you can leave with your own data. Ask for the coverage number and its denominator, whether you can see and edit the model, how the vendor resolves a conversion two channels both claim, whether it has any commercial relationship with the ad platforms it grades, and whether you can export raw event-level data on day one and the day you cancel. Each has a good answer, a bad answer, and a tell that the salesperson is dodging. None of it measures incrementality. The audit finds the least-biased, most auditable allocation layer, not a truth machine.
The scorecard, not the checklist
Most "how to choose an attribution vendor" content is a feature checklist written by people who sell attribution vendors. It compares dashboards and integration counts and asks whether the tool "fits your stack." That is the wrong instrument. It tells you what the tool looks like, not where its numbers come from.
This is a weighted scorecard instead, backed by a five-question demo audit. You score seven categories from 1 to 5, weight them, and total to 100. The weighting is the opinion: the categories that decide whether the numbers are trustworthy at all carry the most, because a beautiful dashboard on biased data is still biased.
The idea that carries the whole thing: a vendor is only as independent as its data supply chain. Any tool that ingests platform-reported conversions and re-serves them through a prettier model inherits the platform's bias by construction. Dedupe logic and a data-driven model can rearrange that bias across channels, but they cannot remove it. You cannot de-bias a number downstream of the party that profits from inflating it. Call that pattern independence laundering. The scorecard exists to catch it.
One guardrail up front, because this audience will otherwise stop reading. No attribution vendor measures incrementality. Attribution allocates credit for conversions that already happened. It does not tell you which conversions would not have happened without the ad. That is a separate discipline: holdout and geo experiments, and marketing mix modeling at the macro level. A high score finds the least-biased, most-auditable allocator, the thing your incrementality tests and MMM triangulate against. It does not crown a truth machine.
How to score
Run the five-question demo audit below, then score each criterion on this scale. Weighted points for a criterion are (your score ÷ 5) × that criterion's share of its category weight. Total the seven categories to 100.
THE 1–5 SCALE
The seven categories
Each category holds two or three scored criteria. Score every criterion 1–5; the category weight is split across its criteria.
1. Data completeness (coverage)
25
1.1 Names a real coverage number. States coverage as observed ÷ actual with a defined denominator (conversions, not sessions). 5 = defines it and shows the figure; 1 = "~100%", no denominator.
1.2 Reconciles live against your backend. Will put observed conversions next to your Shopify/Stripe/CRM count for one week, in the demo. 5 = runs it live; 1 = refuses.
1.3 Server-side / first-party collection. Captures events your servers send, not just client-side pixels. 5 = server-side SDK/tag; 1 = client-side only or reads platform APIs.
2. Model transparency
20
2.1 You can see the model. Shows how credit is assigned; the logic is inspectable. 5 = readable rules / visible weights; 1 = "proprietary, can't view".
2.2 You can compare or edit it. Multiple models side-by-side on the same data, or a way to fork and edit the rules, with a visible diff and a defensible default. 5 = compare + edit + default; 1 = one fixed hidden model.
3. Data ownership / export
15
3.1 Raw event-level export. Raw events, not aggregated reports, in a queryable format from day one. 5 = warehouse sync (BigQuery/Snowflake) you control; 1 = aggregated CSV only.
3.2 Export on exit. Full raw export the day you cancel, no premium gate. 5 = guaranteed on exit; 1 = data lost or gated on cancellation.
4. Independence / incentives
15
4.1 No platform commercial ties. No revenue-share, data-licensing, referral or co-marketing deal with any ad platform it grades. 5 = none, disclosed; 1 = independent branding + undisclosed ties.
4.2 Doesn't treat platform conversions as truth. Builds the count from your transaction system, treats platform-reported conversions as one signal to reconcile, not ground truth. 5 = your backend is the denominator; 1 = platform conversions as source of truth.
5. Setup reality
10
5.1 Honest about setup effort. States real setup time (hours to days for a tag/SDK) and what breaks. 5 = specific and honest; 1 = promises "zero setup" (means client-side or platform APIs).
6. Pricing honesty
10
6.1 Published, readable pricing. Real prices on the page, no forced demo to see a number. 5 = public tiered pricing; 1 = "contact us" only.
7. Support & docs
5
7.1 Real docs and responsive support. Public docs, working examples, a human who answers. 5 = strong public docs; 1 = docs gated, support is a sales rep.
Reading the total
One override on the total: any score of 1 in Data completeness or Independence is a hard disqualifier whatever the total says. A vendor that launders platform conversions or won't define its coverage denominator has failed the audit even if it scores well everywhere else.
The rest of this article is the demo audit that produces the scores: the five questions, in the order to ask them, with the good answer, the bad answer, and the tell for each.
Question 1: Demand a coverage number (scores category 1)
Start with coverage because a beautiful model on partial data is a confident guess. Ask what fraction of conversions the tool actually observes, then refuse the first answer, because the first answer is always near 100 percent.
The number is meaningless without a denominator. Coverage of what: sessions, known users, or conversions? Those are different questions with wildly different answers. A tool can observe 98 percent of sessions and still miss a third of conversions if the conversions happen where its tracking does not reach.
The move that turns the question into a real audit: ask them to show you observed conversions versus your own backend conversion count for one week, live in the demo. You know how many orders, signups, or bookings you had last week. Make them put their observed number next to yours. The gap is your real coverage. A vendor that will not run that reconciliation in front of you has just told you the coverage number is a marketing figure.
Coverage matters for a specific reason, and it is not the reason vendors imply. Higher coverage means less missing data. It says nothing about whether the ad caused the conversion. Server-side, first-party collection lowers your dependence on platform-reported conversions and captures roughly 30 to 40 percent more touchpoints than client-side tracking loses to ad blockers and iOS privacy features. That is real, and it improves both the coverage number and the dedupe that depends on it. It does not make the allocation causal. Keep the two ideas apart, or the coverage question quietly mutates into "high coverage equals correct attribution," which is false.
Question 2: Can you see the model, and can you change it
Every vendor sells a model. The question is whether you are allowed to look at it.
Ask two things: can I see the model, and can I edit the weights? A good answer shows you the logic and lets you touch it. A bad answer is some version of "our algorithm is proprietary; you cannot view or edit the weighting." That leaves you holding a number you cannot audit and cannot reconcile against your own MMM or your holdout tests.
A WORKED EXAMPLE — THE BLACK-BOX TELL
A vendor sells a proprietary data-driven model. You ask whether you can see it and change the weights. The answer: the algorithm is proprietary, no viewing, no editing. Now you have a number you cannot interrogate.
Six months later your finance team's marketing mix model says paid social contributes about $400k incremental per quarter. The vendor attributes $1.1M to paid social. That is a 2.75× gap, and because the model is sealed you have no way to find out where it comes from. You cannot see that it is crediting a 30-day-click window, and you cannot reset that window to 7 days to watch the allocation move.
Compare a vendor whose model you can inspect and fork, ideally in a readable rules language or a SQL-like DSL. You open the model, see the window, change it, and the number reconciles or it does not, but either way you can check its work. A black box you cannot reconcile against your incrementality read is a liability, not an asset. Checking the work is the entire point of paying for a third-party layer.
There is an honest cost on this side too, so do not oversell editability as pure virtue. A fully editable model can be misconfigured by a buyer into something worse than a sensible default. Pair the question with a second one: does it ship with a defensible default, and does it show you what changed when you edit it? Editability plus a good default and a visible diff is the thing you want. Editability with no guardrails is just a different way to get the wrong number.
Question 3: How dedupe actually works
When Meta claims a $180 purchase under a 1-day view-through window and Google claims the same $180 purchase under a 30-day click window, somebody has to decide who wins. That decision is dedupe, and it silently moves credit between channels. Ask how the vendor makes it.
There are three broad answers. Deterministic matching resolves the conflict on a shared identifier: an email, a user ID, a login. High confidence, lower reach, because not every journey carries a shared ID. Probabilistic matching uses device, IP, and behavioral signals to extend past that boundary. Higher reach, lower precision. Both are legitimate; the tradeoff is precision versus coverage, and the right mix depends on how much of your journey is logged-in.
Neither of those is the answer that should worry you. The answer that should worry you is the third one.
A WORKED EXAMPLE — THE DEDUPE TELL
A vendor pitches itself as vendor-agnostic and neutral. Its dashboard reports 4.2× blended ROAS. You ask the dedupe question: when Meta and Google both claim the same $180 purchase, who wins? The answer: "We use the platform's reported conversion as the source of truth and apply our model on top."
That is the tell. If Meta claims the sale under a 1-day view-through window and Google claims it under a 30-day click window, and the vendor ingests both platform-reported conversions, it is now deduplicating two already-inflated numbers. Its "independent" 4.2× is a weighted average of Meta's and Google's self-graded homework.
Put a number on it. If 30 percent of the vendor's conversion volume is platform-reported view-through, a class of conversion the research shows is heavily overstated, then a 4.2× headline can be masking a true blended figure closer to 2.5 to 3× once you strip out conversions the ads did not cause. The audit outcome is not "reject." It is "this is a channel-mix rearranger, not a truth layer; validate it against a holdout before you trust the allocation."
The practical disqualifier is simpler than any of this. When a vendor will not tell you its dedupe method, it is usually because the answer is "we take the platform's word for it." A vendor that cannot articulate how it resolves a double-claimed conversion is a vendor that does not resolve it.
Question 4: Who pays whom
Every vendor says it is independent. Independence is a claim about the data supply chain and the money, not about the logo on the slide, so ask about the money directly.
Do you have a revenue-share, data-licensing, referral, or co-marketing relationship with any ad platform you report on? Do you resell or embed platform-reported conversions as ground truth? The laundering pattern is the combination: independent branding on the outside, platform-reported conversions treated as the source of truth on the inside. A vendor can be genuinely neutral and still ingest platform data as one signal to reconcile. The problem is the vendor that ingests it as ground truth and then charges you for a "neutral" read on top.
Expect a smart pushback here, because it is partly right: "everyone ingests platform data, you cannot avoid it." True. There is no platform-free attribution vendor, and pretending otherwise is its own kind of dishonesty. The distinction that survives the pushback is what the vendor does with the platform's number. A tool that collects your events server-side and treats each platform-reported conversion as one signal among several to reconcile is categorically different from a tool that treats the platform's conversion as ground truth and models on top of it. The audit's job is to locate the vendor on that spectrum, not to demand a purity that does not exist.
Question 5: Export rights
The last question is about the day you leave, which is the day the vendor has the least incentive to help you.
Ask whether you can export raw, event-level data, not aggregated reports, on day one and on the day you cancel, in a format you can query without them. The answer you do not want: "You can export aggregated reports as CSV; raw event data stays in our platform."
A WORKED EXAMPLE — THE EXPORT-HOSTAGE TELL
A vendor demos beautifully. You ask the export question. The answer: aggregated reports as CSV, raw event data stays in the platform. Fine, you think, you can pull a summary whenever you want.
Two years in, you have routed 40 million events through them. Your entire historical attribution baseline lives in their schema. Switching vendors now means starting your measurement history from zero, because the summaries you exported cannot be rebuilt into event-level truth. The export answer was a switching-cost number the whole time: no raw export equals a re-instrumentation project and a broken time series the day you leave.
Contrast a vendor that syncs raw event-level data to a warehouse you control, your own BigQuery or Snowflake, from day one. You own the substrate, the vendor owns the model, and you can leave with your data and your history intact. For anyone who intends to triangulate attribution against MMM and holdouts, no raw export is a walk-away.
Here is the one-week reconciliation query from Question 1, written against a warehouse you control. If the vendor cannot give you the raw events to run something like this, that is the answer to Question 5.
-- Reconcile vendor-observed conversions against your own backend count,
-- from raw event-level data synced to your warehouse.
WITH vendor_observed AS (
SELECT
DATE_TRUNC(event_ts, WEEK) AS week,
COUNT(DISTINCT conversion_id) AS vendor_conversions
FROM vendor_events
WHERE event_type = 'conversion'
GROUP BY 1
),
backend_truth AS (
SELECT
DATE_TRUNC(created_at, WEEK) AS week,
COUNT(DISTINCT order_id) AS actual_conversions
FROM orders
WHERE status = 'completed'
GROUP BY 1
)
SELECT
b.week,
b.actual_conversions,
v.vendor_conversions,
SAFE_DIVIDE(v.vendor_conversions, b.actual_conversions) AS coverage_ratio
FROM backend_truth b
LEFT JOIN vendor_observed v USING (week)
ORDER BY b.week DESC;
Aggregated CSVs are a hostage situation dressed as portability. You can see the summary but you cannot rebuild your history, run the query above, or leave without losing your baseline.
What the audit does not do
This audience deserves the limitations stated plainly, so read them before you run the script on a live vendor.
Passing all five questions does not make a vendor accurate. It makes the vendor auditable and clean-sourced. Those are different properties. The audit finds the least-biased allocator, and that is worth finding, but it is not truth.
No attribution vendor measures incrementality, full stop. MTA allocates credit for conversions that occurred; it cannot tell you which were caused by the ad. Only experiments and, at the macro level, MMM approach the causal question. There is no true number here, only better and worse lenses. Any sentence implying the audit surfaces "the real number" is wrong, and a measurement-literate buyer will dismiss the whole exercise if you let that sentence in.
The classic evidence is worth carrying into the room. In a set of 15 large-scale Facebook field experiments, Gordon, Zettelmeyer, Bhargava and Chapsky found that common observational and attribution methods overstate ad effectiveness by roughly a factor of three for purchase outcomes, and by much more for cheaper conversions like registrations and page views. In eBay's paid-search experiment, Blake, Nosko and Tadelis found that when eBay stopped bidding on its own brand keywords, essentially all of that traffic arrived anyway through organic and direct, which means the attribution report had been crediting brand ads for close to zero incremental sales. A perfectly attributed channel can be worth nothing. That is the gap between allocation and cause, and no vendor audit closes it.
ALLOCATION IS NOT CAUSE
The question: how far can an attribution number sit from what an experiment measures?
A few more honest edges. Deterministic dedupe is not free: high-precision matching lowers coverage, so you can dedupe cleanly and still be blind to the 30 to 40 percent of the journey with no shared identifier. Dedupe method is a question about how bias enters, not a box that, once ticked, delivers accuracy. Server-side collection reduces the platform-bias problem but does not eliminate it. And model editability, as noted, has a real downside if it ships without a defensible default. The audit is a filter for supply-chain cleanliness and auditability, not a certificate of correctness.
Run the five questions, score the seven categories against the scorecard at the top, and total to 100. The vendor that answers all five well is the lens your MMM and your holdouts get to argue with. The vendor that dodges even one is selling you a platform's homework with a nicer chart on top, and it will show up as a 1 in the category that matters most.
Key Takeaways
- ✓A vendor is only as independent as its data supply chain. Anything that ingests platform-reported conversions and re-serves them through a prettier model inherits the platform's bias by construction.
- ✓The coverage question only counts if you fix the denominator: coverage of what, and can they reconcile observed conversions against your own backend count live?
- ✓A model you cannot see or edit is a number you cannot reconcile against your MMM or your holdout tests, which is the whole reason to buy a third-party layer.
- ✓The disqualifier on dedupe is a vendor that won't name its method and silently defaults to platform-reported conversions as ground truth.
- ✓No attribution vendor measures incrementality. The audit finds the cleanest allocator; experiments and MMM supply the causal read it triangulates against.
If no attribution vendor measures incrementality, why buy one at all?▼
How do I get a real coverage number when every vendor claims about 100 percent?▼
Deterministic or probabilistic dedupe, which should I want?▼
What does a real conflict of interest look like? Most vendors say they are independent.▼
What export rights actually matter? I can always get a CSV.▼
Does high coverage mean the attribution is accurate?▼
How mature is your marketing measurement?
The free Measurement Maturity Assessment shows where you stand, where you're exposed, and what to fix first. 10 questions, 3 minutes.
Take the AssessmentReady to try server-side attribution?
Set up in 10 minutes. Free up to 30K records/month.