SKILL DETAIL
ad-attribution-gap
mbfinotti/advertising-skills/ad-attribution-gap
Quantify and explain the discrepancy between ad platform reporting, an analytics tool, and the source of truth (CRM or order system), classifying every unit of the gap as timing, definitional, or unexplained residual. Use whenever the user says the numbers don't match, that the platform reports more conversions or revenue than the CRM or order system, or mentions cross-platform reconciliation, double-counted conversions, an attribution discrepancy, or asks whether a reporting gap is normal - even if they never say 'attribution'. Covers B2B (CRM-anchored) and B2C/ecommerce (order-system-anchored). Do NOT use to fix broken or missing tracking - use mbfinotti/advertising-skills@ad-conversion-tracking instead.
Installation
npx skills add https://github.com/mbfinotti/advertising-skills --skill ad-attribution-gap
技能文件
SKILL.md
最近同步 · 2026年9月24日
evals/evals.json›
{
"skill_name": "ad-attribution-gap",
"evals": [
{
"id": 1,
"prompt": "I run growth at Maren & Oak, a DTC furniture brand on Shopify. For March, Meta reported 1,912 purchases and $412k in revenue, our analytics tool shows 1,104 purchases, and Shopify shows 1,351 paid orders and $296k. My CEO wants one number and asked me to get these three systems to agree. Can you reconcile them so they all show the same figure? One wrinkle: Meta's revenue includes shipping and tax, the Shopify report we use excludes them I think. We also processed about $18k of refunds in March.",
"expected_output": "A reconciliation setup that anchors on Shopify, normalizes all three systems to a common basis, classifies gap causes into timing/definitional/residual buckets, and explicitly refuses to make the three systems show one identical figure.",
"files": [],
"expectations": [
"Refuses to make the three systems agree on a single identical figure, and states that a perfect match would indicate a fabricated or hidden adjustment rather than rigor.",
"Anchors the conversion count and revenue on the Shopify order system, not on Meta or the analytics tool.",
"States that the ad platform and analytics explain where conversions came from, never how many there were.",
"Requires normalization to a common basis before comparing any numbers, naming at least date basis, timezone, and conversion definition as axes to align.",
"Identifies the platform's ad-interaction-date (click-date) stamping vs the order system's event-date stamping as a specific cause of misalignment.",
"Recommends one revenue basis of net revenue excluding tax and shipping, applied to every system, to resolve the Meta-includes-shipping-and-tax mismatch.",
"Treats the $18k of refunds as a definitional line (source of truth records reversals the platform and analytics rarely reverse), not as a tracking defect.",
"Carries view-through conversions as a definitional line pushing the platform's count higher, to be segmented out or quantified rather than treated as a bug.",
"Flags modeled or estimated conversions on the platform side and requires they never be compared to deterministic order counts without being labeled as modeled.",
"Classifies gap causes into exactly three buckets: timing, definitional, and unexplained residual.",
"States that an explicit unexplained residual must remain in the final report and must not be driven to zero.",
"Computes or sets up the gap per source pair (Meta vs orders, analytics vs orders) in absolute units and as a percentage of the anchor, rather than one blended three-way number.",
"Explains the analytics undercount (1,104 below 1,351 orders) through analytics-low mechanisms such as consent loss, ad blockers, or cross-device breaks, rather than suggesting the order count is inflated."
]
},
{
"id": 2,
"prompt": "Quick sanity check for my agency report at Bluewater Peak Supply. Last month Google Ads says 640 conversions, Meta says 510, TikTok says 210 - that's 1,360 total from paid - but our order system only logged 780 paid-attributed orders. So paid is over-reporting by 74%, right? Google alone claims 640 while the order data credits Google-tagged orders at about 330, which feels insane. Ad spend across everything was $96k and total revenue was $402k. How do I write up this over-count for the client?",
"expected_output": "A write-up that rejects the 1,360 sum, treats platform claims as overlapping against the 780 anchor, flags Google's roughly 1.9x ratio for double-counting investigation, and uses MER as a blended cross-check.",
"files": [],
"expectations": [
"Rejects summing 640 + 510 + 210 into a 1,360 paid total as invalid.",
"States that the true conversion count is 780 (the order system's number) with overlapping platform claims on it, not 1,360 conversions.",
"Refuses to report a '74% over-count' as the user computed it.",
"Explains cross-platform self-crediting (several platforms each claiming the same conversion) as a definitional cause to quantify, not a defect to fix.",
"Flags Google's 640 claimed vs roughly 330 anchor-credited (about 1.9x) as exceeding the roughly 1.5x ratio that suggests double-counting.",
"Labels the 1.5x ratio as a practitioner heuristic, not a standard.",
"Routes the suspected double-counting to the unexplained-residual/investigation bucket rather than accepting it as a documented definitional line.",
"Computes MER as total revenue divided by total ad spend from the source of truth (about 4.2 from $402k over $96k) as a blended cross-check.",
"Treats MER as a cross-check that sidesteps per-platform attribution, never as a target to hit, and does not present any cited MER figure as a benchmark to reach.",
"Requires normalizing date basis, attribution windows, and view-through/modeled inclusion per platform before quantifying each platform's gap.",
"Names the counting-rule difference (some platforms count every conversion per click, others count one) as a definitional candidate for part of the excess.",
"Computes each platform's gap as its own pair against the 780 anchor with sign, rather than one aggregate paid-vs-orders figure."
]
},
{
"id": 3,
"prompt": "I do marketing ops at Ferrostack, a B2B data-integration SaaS with roughly $40k ACV. I pulled all deals closed-won in Q2 from our CRM - 23 deals, $920k - and compared against LinkedIn Ads, which only claims credit for 4 of them. Our offline conversion uploads are keyed on the LinkedIn click ID we store on each lead, and the ops dashboard says only 38% of Q2 leads have a click ID on file. My CMO says this proves LinkedIn isn't working - last-touch shows it clearly - and wants me to rerun everything under a different attribution model to see if it looks better. Can you tell me how big our attribution gap really is, in percent?",
"expected_output": "A B2B reconciliation that re-cohorts by lead creation date, flags the 38% match rate as a capture defect, refuses the mid-analysis model switch, and reports absolute units alongside percentages.",
"files": [],
"expectations": [
"Rejects cohorting by Q2 close date and requires cohorting by lead creation date instead.",
"Explains that a close-date pull mixes cohorts the platforms saw months apart, because of lead-to-close lag.",
"Flags the 38% click-ID match rate as below the roughly 50% threshold that indicates a capture defect, not a normal loss.",
"States the normal practitioner-reported click-ID match-rate range as roughly 75-85%.",
"Sends the low match rate to the residual/investigation bucket with named candidate mechanisms such as cookies blocked, the ID stripped by a redirect, or the CRM not saving the field on every submission.",
"Classifies deals closing after the platform's import-acceptance window, or general lead-to-close lag, as timing with a platform-low direction, rather than as evidence LinkedIn is not working.",
"Refuses to switch attribution models mid-analysis and treats the model question as out of scope, redirecting it rather than picking a model.",
"Reports absolute units alongside every percentage, refusing the percent-only answer the user asked for, given the 23-deal volume.",
"Recommends longer comparison windows because percentage gaps swing wildly on small denominators.",
"Names offline conversion upload lag or processing latency as a timing cause of the platform undercount.",
"Notes that last-touch under-reports platform contribution on long sales cycles, and records the model difference as a definitional line rather than proof of failure.",
"Recommends attributing at account level rather than contact level, so the form-filler does not absorb credit owed to a multi-person buying group.",
"Assigns the click-ID capture fix an owner (marketing ops or RevOps) and hands the tracking fix off rather than treating it as part of the reconciliation itself."
]
},
{
"id": 4,
"prompt": "At Vanterra Skincare we check ad numbers every morning. Yesterday (Tuesday) our ad platform showed 41 purchases but the store backend had 67 paid orders - the platform is missing a third of our sales. My colleague says it's because the platform includes view-through conversions and modeled conversions, which inflate its numbers, so that explains it. The store reports in UTC-8 and the platform account is set to UTC. We also launched a new campaign Monday with a 7-day click window. What's going on and how do I fix the tracking?",
"expected_output": "An analysis that strikes the colleague's wrong-direction explanation, attributes the gap to single-day grain, timezone mismatch, and the still-open window, and declines to prescribe a tracking fix.",
"files": [],
"expectations": [
"Rejects the colleague's view-through/modeled explanation because those causes push the platform's count higher, which cannot explain a platform-low gap (direction check applied to strike the candidate cause).",
"Identifies the single-day comparison itself as manufacturing the worst-case discrepancy and recommends a 7-day-or-longer window or weekly grain.",
"Identifies the UTC vs UTC-8 timezone mismatch as shifting conversions across calendar days without any real discrepancy.",
"Identifies the still-open 7-day click window from Monday's launch as timing: the platform's number keeps growing after the pull.",
"Names ad-interaction-date stamping vs order-event-date stamping as a reason a single day cannot line up between the two systems.",
"Recommends letting recent conversions age or re-running after the window closes, rather than acting on the immature pull.",
"Declines to prescribe a tracking fix, stating that no defect is established from this comparison.",
"Anchors the order count on the store backend's 67 and treats the platform as explaining origin, not volume.",
"Explicitly classifies the candidate causes as timing or definitional rather than presenting an undifferentiated list of possibilities.",
"States that a defect requires a direction flip, a sudden jump, or a gap resisting known mechanisms, and that this gap is attributable so it does not qualify.",
"Predicts that on a short ecommerce cycle the timing share of the gap resolves within days and recommends re-running the comparison to confirm it.",
"Establishes or asks for each system's attribution window, counting rule, and conversion definition before putting a number on the gap."
]
},
{
"id": 5,
"prompt": "We finished digging into tracking at Loamworks, a B2C garden-subscription brand doing about $310k/month. Findings: (a) a duplicated purchase tag in our tag container inflating platform-reported purchases about 18%, call it $29k/month of over-reported revenue; (b) about 220 internal test orders per month polluting the order system, roughly $6k; (c) cross-device tracking losses our analytics vendor estimates at $55k/month of unattributed revenue, stable for the last four months; (d) UTM parameters stripped by our link shortener on email-to-site hops, about $9k/month misattributed to direct. Engineering is fully booked this quarter. Biggest number first means cross-device is our top priority, right? What order do we fix these in?",
"expected_output": "A fix ranking by revenue at stake per unit of fix effort that puts the duplicate tag first, keeps stable cross-device losses deferred and named, and shows the effort adjustment explicitly.",
"files": [],
"expectations": [
"Rejects 'biggest number first' and ranks defects by revenue at stake per unit of fix effort.",
"Places the duplicate purchase tag at or near the top: hour-scale, one owner, against a large over-count.",
"Ranks the test-order cleanup high despite its $6k being the smallest figure, because it is near-zero effort and owned outright by the analyst.",
"Does not place cross-device losses first despite $55k being the largest number.",
"States that cross-device identity work is quarter-scale or standing engineering effort and partly structural at any effort level.",
"Notes the cross-device gap has been stable for four months, and that a stable structural loss can wait while a growing one is what would promote it.",
"Given engineering is booked, names cross-device explicitly as deferred or struck from the ranked fix list, carried as a quantified known delta, rather than silently parked at the bottom of the ranking.",
"Identifies the stripped-UTM fix as spending a second team's hour (whoever owns the shortener or redirect chain), distinct from fixes the analyst owns alone.",
"Shows the revenue figure and the fix-effort adjustment for each defect so the reader can see why the biggest number is not first, rather than presenting a bare ordered list.",
"Instructs annotating the report history at each fix date, because the post-fix step-change is a discontinuity that would otherwise look like a defect next period.",
"Assigns an owner to each defect in the list.",
"Hands the duplicate-tag and dedup work off to a dedicated conversion-tracking process rather than detailing tag surgery inside the reconciliation."
]
},
{
"id": 6,
"prompt": "Endgame question. At Corvid Analytics (B2B, about $70k/month total ad spend) we've reconciled everything and two problems survive. First, Google and LinkedIn each claim credit for basically the same demo bookings - together they claim 118 of our 74 CRM-recorded demos each month. Second, about 9% of our monthly gap survives every explanation we throw at it, and it's been steady at that level. Our CFO will not approve turning ads off anywhere - no dark regions, no test cells, spend keeps running. Leadership wants certainty on which platform actually drives demos so we can move budget. What's the most rigorous next step - should we buy a media mix modeling platform?",
"expected_output": "A next-measurement recommendation that names the holdout as the only causal settlement, deletes it because spend cannot be withheld, challenges MMM fit, ranks the remaining options, and accepts the stable 9% residual.",
"files": [],
"expectations": [
"Identifies a holdout or geo test as the only causal answer that settles a two-platform ownership fight outright.",
"Given the CFO's refusal to withhold spend, deletes the holdout from the menu explicitly and states the causal option is off the menu, rather than parking it under 'later' or recommending it anyway.",
"Challenges the media-mix-modeling purchase: MMM fits a portfolio spanning channels no conversion record covers and requires a permanent owner for a standing pipeline and refresh cadence, which a two-platform account is unlikely to justify.",
"Recommends self-reported attribution ('how did you hear about us') as the default rung, running alongside the existing reconciliation.",
"States self-reported attribution's honest limits (coarse, self-report-biased) alongside its advantage of surviving consent loss and blocking entirely.",
"Presents server-side collection with shared event IDs as the rung that closes double-counting at the source, at the cost of engineering weeks across two teams.",
"Ranks the remaining options on more than one axis (such as value, effort, or efficiency) instead of a single blended list.",
"Does not sum the two platforms' claims: treats 118 claimed vs 74 CRM demos as overlapping claims on 74 demos.",
"Describes the mature end state as triangulation - platform attribution for bid optimization, experiments for causal ground truth, mix modeling for portfolio allocation - with no single method trusted alone.",
"States that platform reporting is structurally self-flattering, so the corrective signal must come from outside the platforms.",
"Names the obligation the recommended rung incurs (for example a privacy-notice line for a new form field, or consent and data-residency review for server-side collection).",
"Treats the steady 9% residual as documented and acceptable under the judgment test (stable, direction-consistent), not as a defect to chase to zero.",
"Assigns the chosen next step an owner and stops there, rather than designing the study or implementation inline."
]
},
{
"id": 7,
"prompt": "Marketing analyst at Pinwheel Labs, a PLG dev-tools company with about 40k monthly sessions. Total signups in analytics vs our billing system have tracked within a steady ~12% gap for over a year. But in the last two months the 'direct' channel share jumped from 31% to 44% of signups while paid-social and organic-social shares dropped - totals are still fine, just the mix moved. Separately, my manager wants to institutionalize an 'attribution correction factor': since analytics undercounts billing by 12%, we'll multiply analytics by 1.12 in every dashboard going forward and stop doing these tedious comparisons. Two questions: is our 12% gap normal, and can we ship the correction factor?",
"expected_output": "A verdict that the stable 12% total gap is normal, the sudden direct-share jump is a defect signal with named misattribution mechanisms, and the permanent 1.12 multiplier is rejected in favor of per-period re-derivation.",
"files": [],
"expectations": [
"Distinguishes the channel-mix shift (total intact) from a total-level gap and names misattribution-to-direct mechanisms: tracking parameters stripped by redirects, link shorteners, or in-app browsers, and/or AI-assistant referrals arriving with no referrer.",
"If citing the figure that roughly 70% of AI-assistant-driven traffic arrives with no referrer, labels it directional rather than audited.",
"Treats the sudden two-month direct-share jump as a defect signal under the judgment test (a sudden jump), warranting investigation, in contrast to the stable total gap.",
"Judges the steady 12% total gap normal using the stability, direction-consistency, and decomposability test rather than a bare percentage threshold.",
"When citing tolerance bands, labels their source status: only the analytics vendor's own documented figures (up to 10% pageview and 20% user/session discrepancy) are official, while other bands come from vendors selling tracking audits and are indicative only.",
"Rejects shipping the permanent 1.12 correction factor.",
"Explains the rejection: consent rates and media mix shift continuously, so the gap must be re-derived each period rather than frozen.",
"Points out that the last two months already demonstrate the factor's failure mode: the mix shifted, and a frozen multiplier would have masked exactly this change.",
"Recommends carrying known deltas forward as documented expectations re-derived each period, instead of stopping the comparisons.",
"Notes that because the total-level comparison is intact, conversion counting is sound and the problem is channel assignment, not the count itself."
]
},
{
"id": 8,
"prompt": "I'm the first data hire at Squall & Co, a B2C meal-kit brand. Board meeting in two weeks. Our CMO wants the deck to use Google Ads' reported ROAS of 4.1 because 'Google has the most complete data,' and wants us to standardize all reporting on Google Ads numbers as our source of truth going forward. I ran a gap analysis myself: the gross gap between Google and our order system is 1,410 conversions for the quarter. I could pin 940 of them on specific causes - window differences, view-through, refunds - but 470 I can't explain yet. I'm inclined to present it as fully reconciled to keep the deck clean. Advice?",
"expected_output": "A pushback that anchors board numbers on the order system, labels platform ROAS non-incremental, computes the 67% explained share against the 80% threshold, and prescribes iteration plus an honest residual statement over a clean fabrication.",
"files": [],
"expectations": [
"Rejects standardizing on Google Ads as the source of truth: platforms report claims on conversions, and the anchor must be the system that records the money (the order system).",
"States that board- or executive-facing numbers anchor on the finance-grade source of truth.",
"Requires that any platform ROAS shown in the deck be labeled non-incremental, with platform numbers reserved for in-platform bid optimization.",
"Computes the explained share as roughly 67% (940 of 1,410) and states it falls below the roughly 80% pass threshold.",
"Refuses to present the reconciliation as fully explained, stating that claiming a clean tie-out fabricates a number.",
"Prescribes iterating on the 470 unexplained conversions by walking the timing and definitional cause lists with the direction check (each candidate cause must push the gap the observed way).",
"Recommends re-running the comparison after attribution windows close to confirm suspected timing items actually resolved.",
"Notes that a residual of this size makes a real tracking defect the likely case (most audited accounts carry at least one), so known defect mechanisms should be checked before accepting the remainder.",
"States that if the share is still below the threshold at the deadline, the deck should carry the honest figure with named exclusions of what was ruled out, which beats a fabricated 100%.",
"Justifies the 80% threshold by reference to the officially published expected-discrepancy band, rather than presenting it as arbitrary.",
"Prescribes the reconciliation report shape: per-pair verdicts, a normalization basis, a variance table with bucket, direction check, and owner per cause line, a residual statement, and known deltas to carry forward.",
"Sequences the two weeks with hour- and day-scale checks only, deferring anything that needs a full test period past the board date."
]
}
],
"trigger_queries": [
{ "query": "Facebook says we got 1,200 purchases last month but Shopify only shows 890 - which one is right?", "should_trigger": true },
{ "query": "Why does Google Ads report 3x more conversions than our CRM?", "should_trigger": true },
{ "query": "our ga4 and meta numbers never match, is that normal", "should_trigger": true },
{ "query": "HubSpot shows 40 demos from paid but the ad platforms together claim 95, reconcile this for me", "should_trigger": true },
{ "query": "I need to explain to my CFO why the ad dashboard revenue doesn't equal our Stripe revenue", "should_trigger": true },
{ "query": "attribution discrepancy between google ads and our order system - how do I break it down?", "should_trigger": true },
{ "query": "cross-platform conversion reconciliation for last quarter", "should_trigger": true },
{ "query": "Are double-counted conversions why Meta plus Google claims add up to more sales than we actually had?", "should_trigger": true },
{ "query": "Is a 25% gap between analytics and backend orders something to worry about?", "should_trigger": true },
{ "query": "My boss wants one source of truth for conversions - the ad platform, GA, or Salesforce?", "should_trigger": true },
{ "query": "TikTok claims 500 purchases, our store shows 210 from TikTok. what gives", "should_trigger": true },
{ "query": "how much of the difference between platform-reported revenue and billing revenue is normal vs broken", "should_trigger": true },
{ "query": "Every tool gives me a different ROAS. Which number do I put in the board deck?", "should_trigger": true },
{ "query": "each of our ad channels takes credit for the same signups - untangle the overlap for me", "should_trigger": true },
{ "query": "our agency's reported conversions are way above what lands in the CRM, need to audit the difference", "should_trigger": true },
{ "query": "the numbers in the weekly marketing report don't add up across tools and finance is asking questions", "should_trigger": true },
{ "query": "why did analytics suddenly start disagreeing with our sales data this month", "should_trigger": true },
{ "query": "quantify how much of our platform-vs-CRM gap is timing vs definitions", "should_trigger": true },
{ "query": "view-through conversions are inflating Meta's numbers vs GA - how do I account for that in reporting?", "should_trigger": true },
{ "query": "we spent $80k, platforms claim $500k revenue, but bank deposits say $310k - help me explain the delta to leadership", "should_trigger": true },
{ "query": "LinkedIn ads says 60 leads, our marketing automation says 38, the CRM says 41. which do I trust and why do they differ?", "should_trigger": true },
{ "query": "email platform, analytics, and our marketplace dashboard all report different order counts - reconcile them", "should_trigger": true },
{ "query": "conversion numbers between our pixel and server events don't line up - how much mismatch is acceptable?", "should_trigger": true },
{ "query": "audit why paid channel revenue in our BI warehouse never ties to what the ad platforms say", "should_trigger": true },
{ "query": "my ecommerce client thinks their tracking is broken because the platform and the store disagree by 30%", "should_trigger": true },
{ "query": "explain the difference between what Google attributes to itself and what our last-click analytics report gives it", "should_trigger": true },
{ "query": "Our Q3 numbers: ads manager 2,340 purchases, backend 1,671. Write up what's going on for the exec team.", "should_trigger": true },
{ "query": "does anyone actually get their ad platform numbers to match their CRM? mine never do", "should_trigger": true },
{ "query": "how do I reconcile offline conversion uploads with what the CRM eventually records?", "should_trigger": true },
{ "query": "the marketing team and the finance team report different customer acquisition numbers and I'm stuck in the middle", "should_trigger": true },
{ "query": "the platform says our campaign drove 900 signups but the product database logged 640 - before I kill the campaign, sanity check?", "should_trigger": true },
{ "query": "need a framework for explaining ad reporting discrepancies to non-technical stakeholders", "should_trigger": true },
{ "query": "why is revenue in the ads dashboard 40% higher than actual invoiced revenue", "should_trigger": true },
{ "query": "GA shows fewer conversions than both the ad platform and the order system - where do the missing ones go?", "should_trigger": true },
{ "query": "B2B question: our ad platform never sees the deals that close months later. how should I compare its numbers to closed-won?", "should_trigger": true },
{ "query": "is it normal that after refunds our real revenue is way below what the ad platforms report?", "should_trigger": true },
{ "query": "walk me through separating real tracking problems from expected differences between reporting tools", "should_trigger": true },
{ "query": "meta and google both take credit for the same purchases in their dashboards - how big is the overlap really?", "should_trigger": true },
{ "query": "our conversion counts jumped in the platform but not in the database - did something break or is this normal?", "should_trigger": true },
{ "query": "monthly close: marketing-reported sales vs accounting-recorded sales differ again, need to document why", "should_trigger": true },
{ "query": "how do I set up a recurring reconciliation between ad platforms, analytics, and our order system?", "should_trigger": true },
{ "query": "client asks why their dashboard shows way more leads from paid campaigns than ever reach their inbox", "should_trigger": true },
{ "query": "the gap between platform conversions and CRM conversions doubled this month, what do I check first?", "should_trigger": true },
{ "query": "modeled conversions vs actual orders - how do I compare these honestly in one report?", "should_trigger": true },
{ "query": "I think our ad platform is double counting purchases, how do I prove it?", "should_trigger": true },
{ "query": "Our CPA doubled in the last three weeks and I don't know why - diagnose our ad account", "should_trigger": false },
{ "query": "Set up conversion tracking for our new Meta campaign before launch and verify the pixel fires once", "should_trigger": false },
{ "query": "My purchase event fires twice on the thank-you page - fix the deduplication between pixel and CAPI", "should_trigger": false },
{ "query": "Which attribution model should we use - first click, last click, or data-driven?", "should_trigger": false },
{ "query": "Is our CAC too high for a $49/month SaaS?", "should_trigger": false },
{ "query": "How should I split a $60k monthly ad budget across Google, Meta, and TikTok?", "should_trigger": false },
{ "query": "Design an A/B test to figure out which ad creative wins", "should_trigger": false },
{ "query": "Our ads get plenty of clicks but the landing page doesn't convert - audit it", "should_trigger": false },
{ "query": "When can I scale my winning campaign from $200/day to $1,000/day?", "should_trigger": false },
{ "query": "Build a retargeting sequence for cart abandoners", "should_trigger": false },
{ "query": "Are we overspending against this month's budget? 60% through the month at 75% of spend", "should_trigger": false },
{ "query": "Set a maximum allowable CAC and a ROAS floor for the company", "should_trigger": false },
{ "query": "Is my top video ad fatiguing or is it just seasonality?", "should_trigger": false },
{ "query": "Write 10 headline variants for our spring sale ads", "should_trigger": false },
{ "query": "Which negative keywords should I add from this search terms report?", "should_trigger": false },
{ "query": "Help me pick between Google Ads and LinkedIn for our B2B launch", "should_trigger": false },
{ "query": "Too many ad sets stuck in learning limited - should we consolidate campaigns?", "should_trigger": false },
{ "query": "Building a lookalike audience - which customer list should I upload as the seed?", "should_trigger": false },
{ "query": "Map the buying committee for our $80k ACV product and how to target each role", "should_trigger": false },
{ "query": "Score the hooks in these three video ad scripts", "should_trigger": false },
{ "query": "What ad formats work best for brand awareness on short-form video?", "should_trigger": false },
{ "query": "Draft a creative brief for a UGC campaign", "should_trigger": false },
{ "query": "I want to promote our CEO's LinkedIn post as a paid ad", "should_trigger": false },
{ "query": "Plan our audience targeting layers for a cold prospecting campaign", "should_trigger": false },
{ "query": "Should I switch from target CPA to maximize conversion value bidding?", "should_trigger": false },
{ "query": "Collect our competitors' running ads into a swipe file", "should_trigger": false },
{ "query": "Write a UGC script with 5 hook variants for our skincare brand", "should_trigger": false },
{ "query": "How do I become a senior media buyer? review my portfolio", "should_trigger": false },
{ "query": "Write a job description for a performance marketing manager", "should_trigger": false },
{ "query": "What PPC newsletters and podcasts should I follow to stay current?", "should_trigger": false },
{ "query": "Implement server-side tagging with a tag manager on our site", "should_trigger": false },
{ "query": "Set up GA4 ecommerce purchase events on our checkout", "should_trigger": false },
{ "query": "Build a marketing mix model in Python for our channel spend", "should_trigger": false },
{ "query": "Design a geo holdout test to measure Meta incrementality", "should_trigger": false },
{ "query": "Explain how multi-touch attribution and data-driven attribution work", "should_trigger": false },
{ "query": "Our GA4 sessions dropped 40% after the cookie banner change - fix consent mode", "should_trigger": false },
{ "query": "Reconcile our bank statement against the general ledger for month-end close", "should_trigger": false },
{ "query": "Migrate our conversion tracking from Universal Analytics to GA4", "should_trigger": false },
{ "query": "Why does my Salesforce report show different pipeline than my Salesforce dashboard? both are CRM views", "should_trigger": false },
{ "query": "Set up offline conversion uploads from HubSpot to Google Ads", "should_trigger": false },
{ "query": "Which UTM naming convention should we standardize on?", "should_trigger": false },
{ "query": "Audit our data warehouse dbt models for the marketing schema", "should_trigger": false },
{ "query": "My Meta ads were rejected for policy violations, help me get them approved", "should_trigger": false },
{ "query": "Forecast next quarter's revenue from the current pipeline", "should_trigger": false },
{ "query": "Calculate the incremental lift from our latest brand campaign", "should_trigger": false }
]
}
references/reconciliation-examples.md›
# Worked reconciliation examples
Two complete examples of the deliverable shape defined in SKILL.md. All figures are illustrative - realistic in shape and direction, not benchmarks.
Do not reuse these numbers as expectations for a real account. Derive every amount from the account's own data.
## Example 1 - B2C/ecommerce, one month
### 1. Headline
Anchor: order system, net revenue basis. Anchor count: **1,240 paid orders, $96,400 net revenue**.
Verdicts:
- Platform A vs orders: structural and explained.
- Analytics vs orders: structural and explained.
- Platform A + Platform B jointly claim 1,590 conversions against 1,240 orders: overlapping claims, not extra orders. Do not sum.
### 2. Normalization basis
- Grain: calendar month, single reporting timezone (the order system's).
- Date basis: order-event date. Platform A raw numbers (interaction-dated) re-pulled with a window wide enough to cover all interaction dates feeding this month's orders.
- Conversion definition: paid order (pending and cancelled excluded), dated definition v2 (in force since the 1st of the prior month).
- Counting rule: Platform A set to "every", order system counts orders - difference carried as a definitional line.
- Windows/models: platform windows as currently configured (read from settings, not assumed), analytics uses its cross-channel model - difference carried as definitional.
- Revenue basis: net revenue excluding tax and shipping, refunds deducted in-month.
- Lag maturity: last 5 days of the month flagged immature for Platform A's window. Estimated open-window share carried as timing.
### 3. Variance table
| Source pair | Metric | Amount | % of gross gap | Bucket | Cause | Direction check | Evidence | Owner | Status |
| ------------------------------------ | ---------------------------- | -------- | -------------- | ------------ | ------------------------------------------------------------------------------------------------------------------------- | --------------------------------------- | ---------------------------------------- | -------------- | ------------------ |
| Platform A (1,468) vs orders (1,240) | conversions | +96 | 42% | Definitional | View-through included by platform | platform high - passes | platform data segmented by click/view | Media buyer | Documented |
| same pair | conversions | +54 | 24% | Definitional | Modeled conversions included by platform | platform high - passes | platform's modeled-conversions breakdown | Media buyer | Documented |
| same pair | conversions | +31 | 14% | Definitional | "Every" counting: repeat orders on one click | platform high - passes | 31 multi-order clicks in export join | Analyst | Documented |
| same pair | conversions | +28 | 12% | Timing | Window still open on last 5 days | platform grows later - passes | prior-month lag curve | Analyst | Re-check next pull |
| same pair | conversions | +19 | 8% | **Residual** | Unexplained | - | dedup event-ID spot-check pending | Analytics eng. | **Investigate** |
| Analytics (1,082) vs orders (1,240) | conversions | −103 | 65% | Definitional | Consent + ad-blocker loss at expected baseline | analytics low - passes | consent-rate data | Analyst | Documented |
| same pair | conversions | −38 | 24% | Definitional | Cross-device journey breaks | analytics low - passes | new-user rate by browser | Analytics eng. | Documented |
| same pair | conversions | −17 | 11% | **Residual** | Channel-level only: "direct" share up 4 pts, total intact - suspect stripped parameters / AI-referral no-referrer traffic | mix shift, no total change - consistent | direct-share trend | Analyst | Monitor |
| Orders, revenue | gross $109,900 → net $96,400 | −$13,500 | n/a | Definitional | Tax + shipping + discounts + $2,900 refunds | gross high - passes | order-system finance data | Finance | Documented |
### 4. Residual statement
- Platform A vs orders: gross gap 228 conversions, explained 209 (92%), residual 19 (8%). Stable vs prior month (7%), same direction. **Passes** the ≥ 80% threshold and the judgment test, with the residual documented and the event-ID spot-check queued.
- Analytics vs orders: gross gap 158 conversions, explained 141 (89%), residual 17 (11%), confined to channel mix. **Passes**, with monitoring on direct share.
### 5. Defects and handoffs
None confirmed this period. If the event-ID spot-check finds unshared IDs between browser and server purchase events, open a defect (double-counting, platform high) and hand off to `mbfinotti/advertising-skills@ad-conversion-tracking`.
### 6. Known deltas to carry forward
Known deltas, re-derive each period rather than subtracting as a fixed correction:
- Platform A ≈ +14–19% vs orders, from view-through + modeled + counting rule.
- Analytics ≈ −11% vs orders, from consent and cross-device loss.
- Revenue: net = gross − ~12% (tax/shipping/discounts/refunds).
## Example 2 - B2B, one quarter cohort
### 1. Headline
Anchor: CRM. Cohort: **leads created in Q1** (not deals closed in Q1 - close-date pulls mix cohorts the platforms saw months apart). Anchor count: **412 leads, 47 closed-won, $588,000 closed revenue**.
Verdict: platform vs CRM - structural and explained after the offline-import match rate is accounted for. One capture defect found.
### 2. Normalization basis
- Cohort basis: lead creation date. Closed-won revenue followed to close regardless of close date.
- Conversion definition: CRM-accepted lead (spam and duplicates removed), dated definition.
- Bridge: offline conversion import keyed on the platform click ID stored on the lead, import lag ≈ weekly batch.
- Volume caveat: small denominators - absolute units reported alongside every percentage.
### 3. Variance table
| Source pair | Metric | Amount | % of gross gap | Bucket | Cause | Direction check | Evidence | Owner | Status |
| ---------------------------------- | ------- | ----------- | -------------- | --------------------- | --------------------------------------------------------------- | ---------------------------------- | ------------------------------------------------------------- | ------------- | ------------------------------ |
| Platform (341 leads) vs CRM (412) | leads | −44 | 62% | Definitional | Click-ID match rate 78% - unmatched leads invisible to platform | platform low - passes | import logs: 89 unmatched, 44 net of overlap with lines below | RevOps | Documented |
| same pair | leads | −18 | 25% | Timing | Weekly import batch lag at quarter edge | platform low, lands later - passes | import timestamps | RevOps | Re-check |
| same pair | leads | −9 | 13% | **Residual → Defect** | Click ID not saved on one landing-page form variant | platform low - passes | form audit: hidden field missing on variant C | Marketing ops | **Handoff** |
| Platform closed-value vs CRM $588k | revenue | −$96k | 100% | Timing | 8 deals closed after the platform's import-acceptance window | platform low - passes | close dates vs documented window | Analyst | Documented as structural floor |
| Contact vs account view | credit | 12 accounts | n/a | Definitional | Form-filler credited instead of buying group | n/a | account rollup | RevOps | Documented |
### 4. Residual statement
Platform vs CRM leads: gross gap 71, explained 62 (87%) as definitional plus timing, residual 9 (13%). The residual did not resist classification, though: the form audit attributed it to a missing hidden field, converting it to a **defect**, not an accepted residual. Post-fix, expect the match rate to recover toward the 75–85% practitioner range from its variant-C-weighted low.
### 5. Defects and handoffs
Ranked by revenue at stake per unit of fix effort, with the adjustment shown:
| # | Defect | Revenue at stake | Fix effort | Position after adjustment |
| --- | ---------------------------------------------------------------------- | ----------------------------------------------------------- | ---------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- |
| 1 | Missing click-ID capture on form variant C | est. −$27k pipeline/quarter at cohort close rate (−9 leads) | an hour: restore the hidden field on one form variant, fully reversible, one owner | first - smaller number than the import-window floor below, but the only one that is actually fixable this week |
| - | Closed revenue falling outside the platform's import-acceptance window | −$96k/quarter | not a defect: structural floor, no fix at any effort | excluded from the ranking, carried forward as a known delta |
Defect 1 owner: marketing ops. Hand off to `mbfinotti/advertising-skills@ad-conversion-tracking`. Annotate the history at fix date - the post-fix step-up is a discontinuity, not organic lift.
### 6. Known deltas to carry forward
Known deltas, re-derive both each quarter:
- Platform structurally under-reports closed revenue by whatever share of deals closes after its import-acceptance window (this cohort: $96k, 16%).
- Platform lead counts ≈ −20% vs CRM at current match rate.
## Negative example - what not to ship
> "Platform reported 1,468, orders show 1,240. After adjustments the numbers reconcile exactly to 1,240. Gap: 0%."
This is the bank-reconciliation instinct misapplied. An attribution reconciliation that ties to zero has hidden or invented at least one adjustment - there is always a residual, because consent loss, modeling, and identity gaps cannot be enumerated to the last unit. Ship the explained share and the residual, with the judgment-test verdict, instead of a tie-out.
SKILL.md›
---
name: ad-attribution-gap
description: "Quantify and explain the discrepancy between ad platform reporting, an analytics tool, and the source of truth (CRM or order system), classifying every unit of the gap as timing, definitional, or unexplained residual. Use whenever the user says the numbers don't match, that the platform reports more conversions or revenue than the CRM or order system, or mentions cross-platform reconciliation, double-counted conversions, an attribution discrepancy, or asks whether a reporting gap is normal - even if they never say 'attribution'. Covers B2B (CRM-anchored) and B2C/ecommerce (order-system-anchored). Do NOT use to fix broken or missing tracking - use mbfinotti/advertising-skills@ad-conversion-tracking instead."
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.1.10"
---
# Attribution Gap Reconciliation
You are a marketing measurement analyst. Deliver a reconciliation report: a number on the gap between what each reporting system claims, with every unit of that number attributed to a named cause. The gap between an ad platform, an analytics tool, and the system that records the money is structural and expected - a report where the numbers match perfectly is a report where someone fabricated a number.
Apply one test to every step: does it help put a number on the gap and attribute that number to a named cause? If not, skip it.
## Boundaries - hand off, don't absorb
- Fixing or installing tracking (tags, pixels, deduplication setup) is out of scope. When the analysis concludes "a tag is misfiring" or "dedup is broken", name the finding, quantify its share of the gap, assign an owner, and hand off to `mbfinotti/advertising-skills@ad-conversion-tracking`.
- Choosing an attribution model is out of scope. Treat model differences as a cause of discrepancy to measure and neutralize, never as a decision to make. If the user asks "which model should I use", redirect them to an attribution-model selection resource.
- Ad account performance diagnosis, landing page work, and CAC/ROAS benchmarking belong to other skills.
## The discipline, ported from financial reconciliation
Financial reconciliation rolls both sides to an adjusted balance before comparing, classifies every reconciling item, and only closes at a $0.00 difference. Port the first two moves and deliberately break the third:
1. Normalize both sides to a common basis before comparing a single number.
2. Classify every unit of variance into exactly one of three buckets: timing, definitional, or unexplained residual.
3. State the residual explicitly and do not drive it to zero. A bank reconciliation must tie to $0.00. An attribution reconciliation never does, because the systems genuinely count different things. The goal is a small, explained residual, not a zero one.
Lead with this judgment test - it beats any percentage threshold because it works at any account size:
- A gap is **normal** when it is _stable_ period over period, _in the expected direction_, and _decomposable into named causes_.
- A gap is a **defect** when it _flips direction_, _jumps suddenly_, or _cannot be attributed to a known mechanism_.
The test maps exactly onto the buckets: a defect is precisely a gap that resists classification into timing or definitional.
## Before starting - five questions
Ask these up front, batched. This is a tactical run, not a strategy interview.
1. Which sources are in play? Name the ad platform(s), the analytics tool, and the source of truth.
2. Which system holds the money - a CRM or an order/billing system? That system anchors the conversion count. If the user proposes anchoring on an ad platform, push back: platforms report claims on conversions, not money received.
3. What decision is the gap blocking? (Budget reallocation, a board number, trust in a channel.) This sets how deep to decompose.
4. What date range, at what grain? Recommend weekly or monthly - day-level comparison amplifies timezone and date-stamping artifacts into false discrepancies.
5. B2B, B2C/ecommerce, or blended? Blended businesses run the two funnels as two separate reconciliations (see the B2B vs B2C section).
Ask three more in the same batch. They cost one message and they set the ordering the run ends on - a defect list and a next-measurement choice can't be ranked for the user without them:
6. **What date must the answer land by?** A hard near-term date promotes the fast-acting options - self-reported attribution, an hour-long tag or form fix - and rules out anything needing a full test period.
7. **One-off win or compounding asset?** A one-off answer favors a holdout on the single channel in dispute. Only a compounding mandate justifies standing measurement work.
8. **What is the effort ceiling?** Analyst hours, engineering coordination, political capital to withhold spend, and how reversible the change has to be. A ceiling of "my own hours this week" removes every option that needs another team.
If you can read the user's exported reports (CSV, spreadsheet, warehouse extract), work from those. Otherwise, ask the user to paste each system's own values for:
- Total conversions and revenue for the agreed window.
- Date basis, timezone, currency.
- Attribution window and counting rule, as shown in the system's own settings.
## Step 1 - Normalize to a common basis
A comparison made before normalization is worthless, and normalization is the step practitioners skip. Most reported "discrepancies" shrink as soon as both reports cover the same conversion action, the same date range in the same timezone, and the same definitions. Align each axis below and record the chosen basis in the report:
1. **Date basis.** Ad platforms typically stamp a conversion on the ad-interaction date (click or impression) and backdate it. Analytics tools and the source of truth stamp the conversion event date instead. A Monday click with a Thursday purchase appears on Monday in one system and Thursday in the other, so a narrow pull can show a conversion in one and zero in the other. Re-date one side or widen the window until both bases are covered.
2. **Timezone.** Each system reports in its own configured timezone. A conversion near midnight lands on different calendar days with no real discrepancy. Weekly/monthly grain washes most of this out.
3. **Currency.** Confirm each system's currency and which side applies conversion. For revenue, the source of truth's rate on the transaction date is the anchor.
4. **Conversion definition.** Align which event counts, and each system's counting rule - some platforms count _every_ conversion per click, others count _one_. A single order can carry different units across dashboards.
5. **Attribution window and model.** Do not assume defaults - platform defaults change and differ per platform. If you can browse the web, look up each platform's current documented default window. Otherwise, ask the user to read it from the platform's settings. Match windows and models across systems where configurable. Where not configurable, record the difference as a definitional cause to quantify in Step 3.
6. **Click-through vs view-through.** Platforms include view-through conversions that click-based analytics never sees. Segment view-through out before comparing, or carry it as a definitional line.
7. **Modeled vs observed.** Platforms and analytics tools add modeled/estimated conversions where tracking is blocked. Never compare a modeled estimate to a deterministic order count without labeling which is which.
8. **Revenue basis.** Pick one basis: net revenue excluding tax and shipping is the recommended default. Apply it to every system, and record how refunds, cancellations, and discounts are treated on each side.
9. **Lag maturity.** Compare only lag-mature windows: if the attribution window is still open for part of the range, the platform number keeps growing after the pull, so flag the immature share as timing. Data also settles late on the analytics side, so let recent conversions age before treating a missing match as a gap - one practitioner method waits a full four days. Comparing too early manufactures a phantom discrepancy, and a single-day window manufactures the worst case: use 7 days or longer.
## Step 2 - Quantify the gap
1. Anchor the conversion count and revenue on the source of truth. Every other source explains _where_ conversions came from, never _how many_ there were.
2. For each pair - each ad platform vs the anchor, the analytics tool vs the anchor, and platform vs analytics - compute the gap in absolute units and as a percentage of the anchor, with its sign (which side is higher).
3. Never sum across ad platforms. If one platform claims 100 and another claims 80 while the order system recorded 120, that is 120 conversions with overlapping claims - not 180. The overlap between platform claims is itself a definitional line to quantify.
4. The gross gap per pair (absolute difference after normalization) is the quantity Step 3 must decompose.
5. Run the blended cross-check: MER = total revenue ÷ total ad spend, computed from the source of truth and actual billed spend. One attribution vendor cites around 5.0+ as a common figure, with the caveat that it varies widely by industry and stage. Treat MER as a cross-check that sidesteps per-platform attribution entirely, never as a target. If per-platform numbers improve while MER degrades, the platform numbers are the ones lying.
6. Ratio check: a platform-reported count above roughly 1.5x the anchor-recorded count suggests double-counting (a practitioner heuristic, not a standard) - send it to Bucket 3 for investigation rather than Bucket 2.
Steps 1–4 are identical for B2B and B2C. Only the anchor system differs.
## Step 3 - Classify every unit into exactly one bucket
Apply the direction rule before anything else: each cause pushes the gap in a known direction. A cause that would push the gap the opposite way cannot explain an observed gap of that sign - use this to strike candidate causes fast. In the lists below, "platform high" means the cause makes the ad platform report more than the anchor.
### Bucket 1 - Timing: real, resolves itself
The two systems will agree once processing catches up. No fix needed, but the comparison must be re-dated or re-run later. Note each item's expected resolution date.
- Ad-interaction-date vs event-date stamping (direction depends on which edge of the window the pull sits on).
- Attribution windows still open - conversions from recent clicks not yet landed (platform low, then catching up).
- B2B lead-to-close lag: the deal closes in the CRM weeks after the click. A platform whose window is shorter than the sales cycle never sees it (platform low vs closed revenue).
- Offline conversion uploads and processing latency - imports land hours or days after the event (platform low until processed).
- In-flight orders: pending/unfulfilled orders counted by one system, not yet by another.
### Bucket 2 - Definitional: real, explained, never closes
The systems are counting genuinely different things. Quantify each line and carry it forward as a documented, expected delta - do not "fix" it.
- Attribution window length differences (longer window → that system high).
- Attribution model differences - a platform credits 100% of its own click, a cross-channel analytics model splits credit across channels (platform high vs analytics).
- View-through conversions included on one side only (platform high).
- Modeled/estimated conversions included on one side only (that side high).
- Counting rule: "every" vs "one" conversion per click (the "every" side high).
- Cross-platform self-crediting overlap - several platforms each claiming the same conversion (sum of platforms high vs anchor, this is why platforms are never summed).
- Revenue gross vs net of tax, shipping, and discounts (gross side high).
- Refunds, cancellations, and chargebacks recorded by the source of truth but rarely reversed in platforms or analytics (platform/analytics high on revenue).
- Currency conversion applied at different rates or dates.
- Bot/invalid-traffic filtering rules that differ per system.
### Bucket 3 - Unexplained residual: the only bucket that signals a defect
Whatever remains after Buckets 1 and 2 are quantified. This is the only bucket that gets escalated, and the only one handed to `mbfinotti/advertising-skills@ad-conversion-tracking`. Known mechanisms that hide here:
- Broken, missing, or duplicate-firing tags (duplicates: platform high, often well past the ~1.5x ratio).
- Browser/server event deduplication failure (platform high).
- Consent, tracking-prevention, and ad-blocker loss beyond the expected baseline (analytics low).
- Tracking parameters stripped by redirects, link shorteners, or in-app browsers - traffic dumped into "direct" (analytics misattributes, channel-level gaps without a total-level gap).
- AI-assistant referrals: roughly 70% of AI-assistant-driven traffic arrives with no referrer and is misclassified as "direct" (a growth newsletter's figure - directional, not audited). Channel mix shifts toward "direct" with no total change.
- Cross-device and identity gaps - the journey breaks between devices or subdomains (analytics low, platform less affected).
- Test orders and internal traffic polluting the source of truth (anchor high - yes, the anchor can be wrong too).
Every line in every bucket gets: amount (units and revenue), direction check passed, evidence, and owner. If an amount can only be estimated, label it an estimate and state the basis.
## Step 4 - Judge the residual
1. Compute the residual per pair: gross gap minus quantified timing minus quantified definitional. Express it in units, revenue, and as a share of the gross gap.
2. Apply the judgment test: is the residual stable, direction-consistent, and small? Then document it and stop - do not chase it to zero.
3. Context for "small", stated honestly: the only officially published tolerance comes from a major analytics vendor's own documentation - pageview discrepancies up to 10% and user/session discrepancies up to 20% "can be expected and are not a cause for concern." Every other band below comes from vendors and agencies that sell tracking audits, so the direction of agreement is meaningful but the exact percentages are self-reported, not audited. Cite them as indicative triage heuristics, never as standards, and say so in the report:
| Pair | Cited normal | Cited investigate |
| ----------------------------- | ------------ | ---------------------------------------------------------------- |
| Ad platform vs analytics | 10–20% | materially above the band; 10–30% is also widely cited as normal |
| Analytics vs order backend | under ~25% | above ~35% |
| Server-side vs browser events | up to ~10% | ~30% alongside a match rate under 50% |
4. Calibrate the prior: agency audits report that most accounts carry at least one significant conversion-tracking defect. Finding a Bucket 3 defect is the common case, not the exception. The direction rule still has to place it, though, or it stays a residual rather than becoming a conclusion.
5. A residual that flips direction, jumps suddenly, or resists every known mechanism is a defect. Rank defects by revenue at stake **per unit of fix effort**, never by revenue alone. Show the adjustment: state each defect's revenue impact, then its fix effort, then where the ratio actually places it. Typical ordering across the common defect classes:
- efficiency: duplicate-firing tags > anchor pollution (test and internal orders) > missing click-ID capture > stripped tracking parameters > broken or missing tag > browser/server dedup failure > consent and tracking-prevention loss > cross-device identity gaps
- value: cross-device identity gaps > consent and tracking-prevention loss > browser/server dedup failure > duplicate-firing tags > broken or missing tag > missing click-ID capture > stripped tracking parameters > anchor pollution
The two orderings invert, which is the point of splitting them.
- Identity and consent loss carry the biggest revenue number and are the least fixable: a quarter of engineering or a standing job, partly structural at any effort. Consent-loss remediation also needs legal sign-off before anything ships, which pushes it further below its revenue rank.
- A missing hidden form field, a filter on test orders, or a duplicate tag is an hour of work and fully reversible.
The four hour-scale classes are not interchangeable despite costing the same hour:
- Duplicate tags lead: one owner fixes one tag container against a large over-count.
- Anchor pollution follows: near-zero effort, the analyst owns it outright.
- Click-ID capture and stripped parameters each spend a _second_ team's hour: marketing ops for a form field, whoever owns the redirect chain for the other.
What this efficiency order starves is the top of the value list. Cross-device identity gaps and consent loss lose every round on ratio and get deferred period after period while the hour-scale fixes cycle.
Promote either above the hour-scale classes when either condition holds:
- The gap it owns is _growing_ period over period rather than stable.
- The decision named in question 3 is a channel-level reallocation that the identity gap itself distorts.
A stable structural loss can wait. A growing one compounds against every future period.
Re-rank against this account before reporting: a defect class the team already fixes in-house moves up, and a hard deadline from the interview promotes every hour-scale class over every quarter-scale one. Delete rather than demote a class this account cannot act on at all. If no team will ever own server-side identity work here, strike it from the ranked fix list by name and carry it in the known-deltas section as an unowned, quantified delta.
A class parked at the bottom of the ranking reappears as scope next period. Assign an owner per defect and hand tracking defects to `mbfinotti/advertising-skills@ad-conversion-tracking`.
6. When the residual survives every mechanism, or when two platforms fight over the same conversions, the answer has to come from outside the reconciliation. Pick from the ranked next steps in the following section, name one, assign an owner, and stop there.
## When the reconciliation can't settle it - next measurement, ranked
Two situations reach past this skill:
- A Bucket 3 residual that survives every known mechanism.
- Two platforms both claiming the same conversions, where the reconciliation can only prove the overlap exists.
Name the next step and assign it an owner. Designing or running any of these is beyond this skill, and the two tracking rungs belong to `mbfinotti/advertising-skills@ad-conversion-tracking`.
Rank by what each buys per unit of effort, not by how definitive it sounds. The axes disagree, so read all four:
- efficiency: self-reported attribution > holdout or geo test > server-side collection with shared event IDs > consent-signal configuration and modeled conversions > media mix modeling
- value: holdout or geo test > media mix modeling > server-side collection with shared event IDs > self-reported attribution > consent-signal configuration and modeled conversions
- effort: media mix modeling > holdout or geo test == server-side collection with shared event IDs > consent-signal configuration and modeled conversions > self-reported attribution
- compliance cost: consent-signal configuration and modeled conversions > server-side collection with shared event IDs > self-reported attribution > holdout or geo test == media mix modeling
Both ties are real equalities, not deferred decisions. On effort, a holdout and a server-side build each cost a second team's commitment for weeks to a quarter: one paid in withheld spend, the other in engineering time. Neither is an hour nor a standing job.
On compliance cost, both sit at the floor of the axis: neither collects any new personal data. There is nothing to disclose, nothing to review, and reversing it re-collects nothing.
| Rung | You spend | You get | You owe |
| ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| Self-reported attribution | an hour of form work plus one reporting field, then a period before it reads | an independent channel signal that survives consent loss entirely - coarse and self-report-biased, but unblockable | a privacy-notice line for the new field |
| Holdout or geo test | a week to design, then a full test period (weeks in B2C, a quarter on B2B cycles), plus the political capital to withhold spend | the only causal answer, and the one that settles a two-platform ownership fight outright | nothing regulatory; fully reversible by turning the spend back on |
| Server-side collection with shared event IDs | engineering weeks across two teams | observed conversions recovered instead of estimated, and double-counting closed at the source | consent signal carried server-side, plus data-residency and contract review before the first event ships |
| Consent-signal configuration and modeled conversions | setup days, then vendor-owned output with no audit access | platform-side volume restored as an estimate - it recreates the modeled-vs-observed line from Step 1 rather than closing it | legal sign-off per jurisdiction; reversing a wrong configuration means re-collecting consent |
| Media mix modeling | a standing job: pipeline, modeling, and a refresh cadence someone owns | portfolio allocation across channels no conversion record covers at all, including offline and brand | nothing regulatory - it runs on aggregate spend and outcome data |
Default rung: self-reported attribution, running alongside the existing reconciliation.
What the efficiency order starves is the causal end of this menu. The holdout and media mix modeling buy evidence no reconciliation can produce, and both lose every efficiency round to an hour of form work: mix modeling ranks last on efficiency while ranking second on value.
Neither is ever promoted by the ratio, so each needs its own trigger, applied deliberately.
- Promote the holdout the moment two platforms claim the same conversions and the disputed budget justifies withheld spend. No cheaper rung settles an ownership fight.
- Promote mix modeling when the portfolio spans channels no conversion record covers at all, including offline and brand, and someone can own it permanently.
This ordering is a default, not a law: it shifts with context and with who executes it. Re-rank against what you already know about this account:
- A warehouse already modeling marketing data collapses mix modeling's effort.
- An existing server-side collection layer turns that rung into a config change.
- A strict consent jurisdiction makes consent-signal work mandatory regardless of its ratio.
The interview answers move it too: a hard near-term date demotes every rung needing a full test period, and a one-off mandate demotes both standing ones.
Delete a rung the user's constraints rule out instead of parking it at the bottom, and name the deletion in the report. A refusal to withhold spend from any region or audience deletes the holdout and geo test outright. Rank the remaining four and say the causal rung is off the menu, rather than leaving it under a "later" heading where it returns as scope next quarter.
- No engineering resource deletes server-side collection the same way.
- Nobody to own a refresh cadence deletes mix modeling.
- No consent-regulated traffic in the account deletes consent-signal configuration.
The mature end state is triangulation, not a winner:
- Platform attribution for daily bid optimization.
- Incrementality experiments for causal ground truth.
- Mix modeling for portfolio allocation.
- No single method trusted alone.
Platform reporting is structurally self-flattering ("Every platform's AI is optimized to make itself look good," as one growth newsletter puts it), which is why the corrective has to come from outside the platform.
## B2B vs B2C
The normalization axes, the three buckets, the direction rule, the never-sum rule, the report shape, and the treatment of platform self-crediting are identical for both. Do not hunt for a difference that is not there.
**B2B (anchor: CRM).**
- Cohort by _lead creation date_, not close date, or the comparison is meaningless - the dominant timing difference is lead-to-close lag, and a close-date pull mixes cohorts the platforms saw months apart.
- The bridge back to the platform is the offline conversion import, keyed on the ad platform's click ID captured in a hidden form field and stored on the lead. The import's own upload lag and match rate are causes of discrepancy.
- Practitioner-reported click-ID match rates run around 75–85%.
- Below ~50%, suspect a capture defect: cookies blocked, the ID stripped by a redirect, or the CRM not saving it on every submission. That goes to Bucket 3.
- Platforms also cap how long after the click an import is accepted. Check the current documented window.
- Browser tracking prevention can delete a cookie-stored click ID within days, which is why server-side capture on first click matters.
- Attribute at account level, not contact level, or whoever filled the form gets the credit for a multi-person buying group.
- Expect low volume: percentage gaps swing wildly on small denominators. Report absolute units alongside every percentage, and prefer longer windows.
- Last-touch under-reports platform contribution on long cycles - as B2B advertising practitioner AJ Wilcox puts it: "It ignores discovery. Prospects rarely convert after a single click." Record the model difference as definitional. Do not switch models mid-analysis.
**B2C/ecommerce (anchor: order/billing system).**
- The dominant definitional lines are refunds, cancellations, duplicate/test orders, and revenue booked gross vs net of tax, shipping, and discounts. Define one revenue basis (recommend net, excluding tax and shipping) and apply it everywhere.
- The order lifecycle (pending → paid → fulfilled → refunded/cancelled) creates reversals the ad platform never sees. A refund can even appear in one of the order system's own reports and not another, so audit the anchor too.
- The browser-vs-server bridge is a shared event ID, usually the order ID, used to deduplicate the two copies of each purchase event. Missing or mismatched event IDs are a Bucket 3 double-counting mechanism.
- Platform over-attribution is sharper than in B2B because view-through and modeled conversions carry more weight in the platform's count.
- Short cycles mean timing differences resolve in days - re-run the comparison after the window closes and most of Bucket 1 should vanish.
## The report
Deliver a reconciliation report with these sections (see [./references/reconciliation-examples.md](./references/reconciliation-examples.md) for a full worked example, one B2C and one B2B):
1. **Headline** - the anchor system, the anchor's number, and the one-line verdict per pair ("structural and explained" or "defect found, owner assigned").
2. **Normalization basis** - date basis, grain, timezone, currency, conversion definition, window/model settings, revenue basis. Date the definitions: a conversion definition changed without a date makes every trend comparison invalid.
3. **Variance table** - one row per cause, columns: source pair | metric | amount | % of gross gap | bucket | cause | direction check | evidence | owner | status.
4. **Residual statement** - per pair: gross gap, explained share, residual, and the judgment-test verdict.
5. **Defects and handoffs** - ranked by revenue at stake per unit of fix effort, showing the revenue figure and the effort adjustment as separate columns so the reader can see why the biggest number is not always first. Each with an owner and a next action.
6. **Known deltas to carry forward** - the definitional lines to re-apply next period, so the next reconciliation starts from documented expectations instead of alarm.
Anchor any board- or executive-facing number on the finance-grade source of truth. Platform-reported numbers are for in-platform bid optimization only, and any deck showing platform ROAS should label it non-incremental. Typical ownership, useful for the owner column:
| Role | Owns |
| ------------------------------------------ | -------------------------------------------- |
| Analyst (usually the reader of this skill) | Normalization and the reconciliation itself |
| Media buyer | Platform config, windows, tagging parameters |
| Marketing ops / RevOps | CRM fields, click-ID capture |
| Data/analytics engineering | Warehouse models, dedup keys |
| Finance | Net revenue - the figure of record |
## Failure modes
- Comparing before normalizing - the most common false alarm. A click-date total compared to an event-date total differs by definition, not by defect.
- Summing conversions across ad platforms, or summing platforms with different windows into one total.
- Driving the residual to zero. A perfect tie-out is evidence of fabrication, not rigor.
- Treating the 10–30% folklore band as a standard, or quoting any tolerance without labeling its source status.
- Comparing modeled estimates to deterministic counts as if both were counted transactions.
- Measuring an average gap once and subtracting it forever as a fixed correction - consent rates and media mix shift the gap continuously. Re-derive it each period.
- Assuming one system is "correct": two systems can both be right while measuring different things - and the anchor can be polluted (test orders, unreversed refunds).
- Cohorting B2B by close date instead of lead creation date.
- Reporting only percentages on low-volume B2B data.
- Fixing a tag without annotating the report history - the fix creates a discontinuity that will look like a defect next period.
## Objective and pass threshold
The natural score for this skill is the **explained share**: (quantified timing + quantified definitional) ÷ gross gap, per source pair.
- **Pass**: explained share ≥ 80%, and the residual passes the judgment test (stable, direction-consistent, no known-mechanism candidates left unexamined).
- Why 80%: the residual it tolerates (≤ 20% of the gap) sits inside the only officially published expectation band (up to 10–20% discrepancy is normal per the analytics vendor's own docs), so chasing the last fifth buys precision the underlying data cannot support.
- Iterate until met: for each unexplained unit, walk the Bucket 1 and Bucket 2 cause lists with the direction rule. Re-run after windows close to confirm suspected timing items actually resolved. Only then accept the remainder as residual, or escalate it as a defect.
- 100% is not achievable. If iteration stalls below 80% with no defect found, say so explicitly in the report with what was ruled out - an honest 70% with named exclusions beats a fabricated 100%.
## References
- `mbfinotti/advertising-skills@ad-conversion-tracking` - hand off confirmed tracking defects (broken tags, dedup failures) found in Bucket 3.
- [./references/reconciliation-examples.md](./references/reconciliation-examples.md) - full worked reconciliation examples (B2C and B2B) with realistic illustrative numbers.
**Optional integration note (vendor-specific, skip unless these are the user's tools).** The generalized mechanics above map to specific platforms and tools:
- Google Ads: ad-interaction-date stamping, "every"/"one" counting, offline conversion import with a 30-day click-ID window.
- GA4: event-date stamping, cross-channel model. The 10%/20% expected-discrepancy figure is from Google Analytics Help, support.google.com/analytics/answer/11986666.
- Meta Pixel plus Conversions API: deduplicated by a shared event ID.
- Click IDs as B2B join keys: `gclid`, `wbraid`/`gbraid` (Google), `msclkid` (Microsoft), `li_fat_id` (LinkedIn).
- Shopify/WooCommerce as B2C anchors: watch Shopify's gross-vs-net "Total Sales" formula and refunds-without-restock.
- Salesforce/HubSpot as B2B anchors.
Citations: MER's ~5.0+ figure is Northbeam's. The ~70% no-referrer AI-traffic figure is Demand Curve's (newsletter #331). The "optimized to make itself look good" line is Demand Curve's. The any-touch quote is AJ Wilcox's (B2Linked).