返回 Skills 目录
mbfinotti/advertising-skills已通过检查

SKILL DETAIL

ad-account-diagnostic

mbfinotti/advertising-skills/ad-account-diagnostic

Diagnose the root cause of an underperforming paid ad account - why the ads stopped working, why CPA went up, why ROAS dropped - instead of defaulting to 'increase the budget'. Weighs tracking, account structure, targeting, creative, bidding and budget, offer, and external forces, and returns a prioritised verdict with evidence and confidence. Covers B2B and B2C on search, paid social, video, and native. Use whenever the user mentions an ad account audit, a campaign structure review, wasted ad spend, rising cost per lead, or says their ads used to work - even if they never say 'diagnostic'. Diagnosis only, and it stops at the click: post-click page problems are mbfinotti/advertising-skills@paid-landing-page-audit and creative decay is mbfinotti/advertising-skills@ad-creative-fatigue.

安装量 · 180查看来源

Installation

npx skills add https://github.com/mbfinotti/advertising-skills --skill ad-account-diagnostic

技能文件

SKILL.md

最近同步 · 2026年9月24日

evals/evals.json›
{
  "skill_name": "ad-account-diagnostic",
  "evals": [
    {
      "id": 1,
      "prompt": "We run Velora Skin, a DTC skincare brand doing about $45K/month on Meta. Our ROAS fell from 2.9 to 1.6 over the last 12 days and the whole team agrees the creatives are burnt out - they are 9 weeks old and we have been meaning to refresh anyway. Can you confirm the fatigue read and help us decide which ads to kill first? Data from the last 12 days vs our 60-day baseline: CPM +3%, CTR -2%, reported CVR -38%. The drop shows up in every campaign - cold, retargeting, even brand - all starting the same Tuesday. Meta reported 348 purchases in the window; our Shopify order table shows 645 paid-attributed orders, and revenue is only about 5% below normal (the gap between Meta and Shopify normally runs around 8%). We also migrated checkout to a new subdomain that same week. Budget unchanged.",
      "expected_output": "A diagnosis that refuses the creative-fatigue read, reconciles platform vs backend first, records a tracking FAIL from the ~46% reconciliation gap tied to the checkout migration, screens every layer with four states, and hands the fix to a conversion-tracking workflow with a prediction and re-check date.",
      "files": [],
      "expectations": [
        "Compares Meta-reported purchases against the Shopify order figures (reconciliation) before diagnosing any other layer.",
        "Identifies the reported-vs-backend gap (roughly 46%, against the ~8% historical norm) as beyond the roughly-40% stop threshold and records a measurement/tracking FAIL.",
        "The verdict names measurement/tracking as the root cause, not creative fatigue.",
        "Declines to confirm the fatigue read and does not recommend killing or refreshing creatives as the fix.",
        "Cites the drop appearing uniformly across cold, retargeting, and brand campaigns on a single start date as a tracking signature, noting fatigue does not synchronise across unrelated audiences on one date.",
        "Connects the break to the checkout migration to a new subdomain in the same week.",
        "Uses the backend evidence (revenue only ~5% down versus a 38% reported CVR drop) to show demand is intact and the signal is broken.",
        "Uses the decomposition (CPM and CTR flat vs the 60-day baseline, CVR the only moved link) to rule out the auction and the ad before naming the cause.",
        "Presents a per-layer screen giving each layer one of four states (pass, FAIL, unknown, not applicable) rather than only naming the winning cause.",
        "Marks the creative layer pass with evidence (CTR flat against its own baseline).",
        "Avoids issuing verdicts on downstream layers through the broken signal, for example marking bidding unknown or deferring it until after the fix.",
        "Hands the tracking repair to a dedicated conversion-tracking fix workflow or owner instead of treating the diagnosis as the fix.",
        "States a prediction: reported CVR and ROAS recover toward baseline within roughly one lag-mature week once the purchase event fires again.",
        "Sets a re-check date one full comparison window (lag-mature) after the fix ships."
      ]
    },
    {
      "id": 2,
      "prompt": "I run growth at NorthPeak Analytics, a B2B data-quality SaaS. We spend about $30K/month on LinkedIn and Google. Our dashboard looks great - CPL is down 28% quarter over quarter to $62 - but sales keeps escalating that the leads are garbage. I pulled the CRM: lead-to-SQL rate went from 12% last quarter to 4% this quarter, worst in exactly the campaigns we shifted budget into. The campaigns optimize to our gated whitepaper form fill, which reconciles fine against the CRM (within a few percent). My plan is to tighten the audience targeting - the broad audiences are probably pulling in junk. Can you audit the account and help me tighten targeting?",
      "expected_output": "A diagnosis that refuses targeting tightening as the first move, names the raw form-fill optimization event as a structure-layer FAIL, judges the account on cost per SQL over B2B-length windows, and puts the CRM-outcome feedback loop at the top of the fix queue with owner, effort, compliance note, and prediction.",
      "files": [],
      "expectations": [
        "Does not adopt targeting tightening as the first fix, identifying the audience drift as a symptom of the optimization event rather than an independent cause.",
        "Names the optimization event (raw whitepaper form fill) as the root cause, recorded as a structure-layer FAIL.",
        "Explains the platform is doing exactly what it was asked: finding people who fill forms cheaply, in volume.",
        "Judges the account on cost per SQL (or cost per closed-won) rather than CPL, showing cost per SQL rose while CPL fell.",
        "Uses the CRM join (12% to 4% lead-to-SQL) as the carrying evidence, with the form-fill reconciliation noted as passing tracking.",
        "Uses the concentration of the SQL-rate collapse in the campaigns that received the shifted spend as localisation evidence.",
        "The top-ranked fix is feeding CRM outcomes (SQL or closed-won) back to the platforms and optimizing to a qualified event, ahead of any targeting change.",
        "Evaluates over B2B-length windows (roughly 90-day or 4-6 week minimum reads), not week-over-week.",
        "Flags that pushing CRM outcomes into ad platforms triggers a consent and data-processing review before the wiring ships.",
        "Presents the per-layer screen with four states; targeting is not marked FAIL as an independent cause.",
        "Predicts the post-fix shape: CPL rises while the lead-to-SQL rate and cost per SQL recover, over one to two mature windows.",
        "Assigns an owner and an order-of-magnitude effort to the top fix (roughly a week of marketing-ops/CRM wiring).",
        "Recommends no budget increase and defers any targeting redesign until the event fix has been judged at maturity."
      ]
    },
    {
      "id": 3,
      "prompt": "Harbor & Finch is a two-partner immigration law firm. We run Google Ads at about $7K/month and normally get 10-12 consultation requests a week. Last week we got 6. Nine days ago we raised the daily budget 50% because we wanted more volume. We have no access to the CRM data right now (the office manager owns it and is on leave). My partner wants to switch the bidding to manual CPC today, or alternatively double the budget to compensate for the drop. Which of the two should we do? We need a decision today.",
      "expected_output": "A refusal to pick either fix: an insufficient-evidence verdict citing the learning reset from the +50% edit, attribution-lag immaturity, and sub-noise conversion volume, with gate math for what clears it, an edit freeze, unknown layer states, low confidence due to unverified reconciliation, and a mature-window re-check date.",
      "files": [],
      "expectations": [
        "The verdict is insufficient evidence: no root-cause layer is named on this data.",
        "Refuses both proposed moves (the manual-CPC switch and the budget doubling) rather than picking one of them.",
        "Identifies the +50% budget change nine days ago as a significant edit that reset the learning phase.",
        "Explains that the most recent week under-reports because of attribution lag, so a trailing window looks like decline by construction.",
        "Notes that roughly 6 conversions is far below any read that clears the noise band, and below the ~30-conversion evaluation reference for automated bidding.",
        "States the gate math: roughly how many more days or conversions are needed at current volume before a verdict is possible.",
        "Recommends freezing further edits until a lag-mature window has accumulated.",
        "Issues no quantified performance prediction on this data, saying why.",
        "Sets a re-check date at a mature window, roughly two or more weeks after the edit.",
        "Records reconciliation as unverified due to the missing CRM access and caps stated confidence accordingly (low).",
        "Uses unknown states for the layers this data cannot power, instead of marking them clean.",
        "Recommends obtaining CRM or backend read access before the next diagnostic run.",
        "Explains that 20-50% day-to-day swings during learning are expected volatility, not evidence of a defect."
      ]
    },
    {
      "id": 4,
      "prompt": "I manage paid search for Brightside HVAC, a regional home-services company. Our search impression share fell from 71% to 54% over six weeks and my boss's read is that we are being outspent, so he asked me to model a 40% budget increase. The columns I pulled: Search lost IS (budget) = 4%, Search lost IS (rank) = 42%. CPA is stable at $58, and on 11 of the last 30 days we did not even spend the full daily budget. How much should we raise the budget - is 40% enough, or should we go higher?",
      "expected_output": "A refusal to size a budget raise: the lost-impression-share split shows rank, not budget, is the constraint; the unspent budget corroborates it; a 40% single-step edit would reset learning; the fix points at rank components via handoff, with the condition under which budget would have been the right answer stated.",
      "files": [],
      "expectations": [
        "Does not answer with a budget-increase amount; declines to size a raise on this evidence.",
        "Splits the impression-share loss into lost-to-budget versus lost-to-rank before making any recommendation.",
        "Identifies rank (42% lost-to-rank versus 4% lost-to-budget) as the binding constraint.",
        "States that increasing the budget will not recover impression share lost to rank.",
        "Uses the unspent daily budget (11 of the last 30 days) as corroborating evidence that budget is not the cap.",
        "Notes that a single-step budget increase beyond roughly 20% counts as a significant edit and resets learning.",
        "Points the fix at rank components (bid and quality/relevance), handed to a bidding/quality workflow or owner rather than executing a bid change itself.",
        "Reads the pattern through decomposition first (volume down while efficiency stable means delivery constrained) before naming the layer.",
        "Verifies, or explicitly conditions the bidding read on, measurement/tracking passing first.",
        "Presents the layer screen with the bidding/budget finding carrying the impression-share-split evidence.",
        "States a prediction tied to the rank fix (impression share recovers only if rank improves) with a re-check window.",
        "Names the condition under which budget would have been the right answer (lost-to-budget dominating with efficiency holding at the margin) and shows this data does not meet it."
      ]
    },
    {
      "id": 5,
      "prompt": "We just took over paid social for PetPal Box, a dog-toy subscription box, from another agency. Their handover doc says: 'CPMs are up industry-wide going into Q4, auction inflation, ride it out.' CPA is up 30% and my CEO wants me to confirm the market story in writing. Blended CPM is +34% vs the prior 60 days. When I break it down: our main prospecting ad set (a narrow stacked-interest audience) is +85% CPM with frequency up from 2.1 to 4.8 while unique reach has been flat for a month; the other four ad sets are +2 to +6% CPM. Can you confirm it's external so I can send the note?",
      "expected_output": "A refusal to confirm the market story: the breakdown shows the CPM rise concentrated in one saturated narrow ad set (frequency up, reach flat), which points internal to targeting; external requires uniform elevation and is a diagnosis of exclusion; the corrected note attributes only the small uniform residual to the market.",
      "files": [],
      "expectations": [
        "Declines to confirm the external/market story as written.",
        "Treats the inherited agency narrative as unverified until the data independently reproduces it, rather than as the hypothesis to confirm.",
        "Reads the breakdown rather than the blended +34% account-level average.",
        "Identifies the concentration (one ad set at +85% while siblings sit at +2-6%) as evidence of an internal, targeted cause.",
        "States the rule: external elevation shows up roughly uniformly across every campaign and audience, while concentrated elevation points internal.",
        "Reads frequency rising from 2.1 to 4.8 with flat unique reach as audience saturation.",
        "Names targeting (narrow-audience saturation) as the layer behind the concentrated CPM rise.",
        "Treats external as a diagnosis of exclusion that cannot be claimed while an internal layer FAILs.",
        "Compares against the account's own history rather than category or industry CPM claims.",
        "Runs, or explicitly conditions the verdict on, the tracking/reconciliation check before the targeting call.",
        "Hands the audience redesign to a targeting workflow or owner rather than designing the new audiences inside this diagnosis.",
        "States a prediction (that ad set's CPM and frequency normalise once the audience is broadened or rotated) and sets a re-check date.",
        "Provides the corrected note for the CEO: mostly internal saturation, with at most the small uniform +2-6% residual attributable to the market."
      ]
    },
    {
      "id": 6,
      "prompt": "Our analyst at Cobalt Gear (outdoor equipment ecommerce) built a performance report. He summed conversions from Meta (7-day click), Google Ads (30-day click), and GA4 (last-click) into one total column, and it shows total conversions down 22% this month vs last month. One thing worth mentioning: on the 12th of this month we changed Meta's attribution setting from 7-day click to 1-day click. Based on the report, which campaigns should we kill? I want to cut the bottom 20% this week.",
      "expected_output": "A refusal to kill campaigns off the summed total: summing across differing attribution windows and models is a category error, the mid-month settings change invalidates the comparison and mechanically deflates the number, sources must be reported side by side and reconciled against the order table, and the comparability requirements for a clean rebuild are specified.",
      "files": [],
      "expectations": [
        "Refuses to rank or kill campaigns based on the summed total.",
        "Identifies summing conversions across differing attribution windows and models as invalid - a category error, not a comparable total.",
        "Identifies the mid-month Meta setting change (7-day to 1-day click) as invalidating the month-over-month comparison on its own.",
        "Explains that part of the reported -22% is mechanical, produced by the shorter attribution window rather than by performance.",
        "Recommends reporting the three sources side by side until their definitions reconcile, instead of blending them.",
        "Notes each platform self-attributes, so platform claims cannot be summed with each other or with GA4.",
        "Recommends reconciling against the backend order table as the source of truth.",
        "Specifies the comparability requirements for the rebuilt comparison: equal window lengths, equal attribution-lag maturity, matched day-of-week composition.",
        "Notes the most recent days under-report due to attribution lag, so a month ending now reads artificially low.",
        "Issues no kill list; the verdict is insufficient evidence (or equivalent) pending a clean comparison.",
        "States exactly what would clear the gate: per-source side-by-side exports on consistent settings plus the backend join.",
        "Sets the re-check for when a clean, lag-mature comparison exists."
      ]
    },
    {
      "id": 7,
      "prompt": "I lead marketing (team of 2) at Fernway Travel, a tour operator. An audit of our ad account surfaced four things: (1) the purchase event fires twice on every booking - confirmed, the platform counts almost exactly double our booking system; (2) we run 14 campaigns and 11 of them get fewer than 10 conversions a month each; (3) our creatives are 5+ months old, and we have a creative studio on retainer with idle capacity; (4) two campaigns consistently underspend their budgets. Constraints: engineering has zero availability this quarter (confirmed by the CTO), we have no CRM - bookings live in a booking system with no export and IT says maybe next year - and the CEO wants visible action this week. Give me the quick wins first, in order.",
      "expected_output": "A ranked plan by outcome-per-effort, not cheapest-first: the double-fire tracking fix ranked first and escalated past the engineering freeze rather than demoted, the CRM-outcome rung deleted with its revival condition, consolidation of sub-volume campaigns handed off, creative promoted because of the retained studio, budget moves last, every finding tagged with severity, confidence, outcome, effort, and owner.",
      "files": [],
      "expectations": [
        "Rejects the cheapest-first quick-wins ordering and ranks by outcome bought per unit of effort.",
        "The double-firing tracking fix is ranked first despite engineering being unavailable.",
        "Handles the engineering constraint on the tracking fix as an escalation to whoever can lift it (CTO/CEO), not as a demotion or deletion, because it is the confirmed root-cause defect.",
        "States that every other fix's results will be measured through the doubled signal until the tracking fix ships, making downstream reads untrustworthy.",
        "Recommends consolidating the 11 sub-volume campaigns (below learning-volume gates), handed off as a consolidation plan rather than executed inside this diagnosis.",
        "Notes consolidation costs a learning window of volatility, setting that expectation against the CEO's this-week demand.",
        "Promotes the creative work above its default rank because the retained studio with idle capacity collapses its effort cost, and says that is the reason.",
        "Ranks budget and underspend moves last, with the reason: near-zero outcome while upstream layers fail, and significant edits reset learning.",
        "The offline/CRM-outcome feedback rung is deleted rather than parked at the bottom, naming the constraint (no CRM or export) and the revival condition (when an export exists).",
        "Tags each finding with severity, confidence, outcome bought, effort, and owner.",
        "Expresses effort as an order of magnitude (hours, a week, a quarter, a standing job), not currency or precise estimates.",
        "Provides a CEO-visible action for this week without letting the visibility demand reorder the queue.",
        "Hands each finding to its owner or follow-on workflow in rung order.",
        "States a prediction for the top fix, for example platform-reported conversions falling by roughly half toward booking-system truth once deduplication ships."
      ]
    },
    {
      "id": 8,
      "prompt": "Lumen Desks - we sell standing desks direct-to-consumer, about $80K/month across Google and Meta. ROAS slid from 3.4 to 2.1 over three weeks. My plan is to restructure the account (it has grown messy) and refresh the creatives, but sanity-check me before I start. The numbers: CPM flat (+1%), CTR flat (-1%), search impression share stable at 68%, but CVR is -41% starting March 3. On March 3 our in-house web team (they sit next to me) shipped a new product-page template and a 12% price increase. Platform-reported purchases are within 7% of our order system, same as always.",
      "expected_output": "A diagnosis that passes tracking on the stable 7% reconciliation, localises the failure past the click via the CVR-only decomposition, ties it to the March 3 release and price change as separate hypotheses, advises against the restructure and creative refresh, promotes offer-and-downstream to the top of the queue because the owners are in-house, and hands off at the click boundary.",
      "files": [],
      "expectations": [
        "Uses the ~7% reconciliation gap matching the historical norm to pass measurement/tracking before reading other layers.",
        "The decomposition isolates CVR as the only moved link, with CPM, CTR, and impression share flat.",
        "Concludes the failure sits past the click (offer and downstream), tied to the March 3 release and price change, not account structure and not creative.",
        "Advises against the planned restructure and creative refresh because no evidence implicates those layers.",
        "States that this diagnosis stops at the click and hands the page and offer work to a post-click/landing-page audit and the page owners.",
        "Promotes offer and downstream to the top of the fix queue, citing the promotion conditions: CVR collapsed while CPM, CTR, and impression share held, and the page owners are in-house.",
        "Separates the two March 3 candidates - the template change and the 12% price increase - as distinct hypotheses to isolate, not one blob.",
        "The layer screen marks structure, targeting, and creative pass with the flat-metric evidence attached.",
        "Recommends checking backend conversion on the paid path specifically, noting that unchanged CVR from other traffic sources is not a full rule-out.",
        "Recommends no budget or bid changes.",
        "Recommends testing one falsifiable hypothesis at a time (for example isolating the template from the price change) rather than stacking fixes.",
        "States a prediction (paid CVR recovers toward baseline if the page/offer fix ships) and a re-check one mature window later.",
        "Presents the verdict in a complete structured block covering window, volume/reconciliation, decomposition, localisation, layer screen, confidence, verdict, evidence, findings, prediction, handoff, and re-check."
      ]
    },
    {
      "id": 9,
      "prompt": "New CMO here at Juniper Wellness, a DTC supplements brand. Our Meta ROAS is 2.4; the industry benchmark report I bought says supplements average 3.5, and benchmark CTR is 1.6% versus our 0.9%. Clearly the account is underperforming badly. I need a full audit that identifies everything that's broken so I can present the fix list to the board next week. Extra data: our ROAS has been between 2.3 and 2.5 for 12 straight months, CTR around 0.9% the whole time, contribution margin is 55%, and blended MER is steady at 2.1. Nothing in the account changed recently.",
      "expected_output": "An audit that refuses the benchmark as baseline: the account is judged against its own 12-month history and a margin-derived break-even ROAS of roughly 1.8, concluding no defect is established; no fix list is manufactured for the board, blended MER against break-even replaces raw platform ROAS as the truth read, and the growth-versus-defect distinction is made explicit.",
      "files": [],
      "expectations": [
        "Refuses the industry benchmark as the baseline and judges the account against its own 12-month history.",
        "States why benchmarks mislead: different mix, geography, and definitions - directional context only.",
        "Derives a break-even ROAS from the 55% contribution margin (roughly 1/0.55, about 1.8) and judges the 2.4 ROAS against it as above break-even.",
        "Concludes no defect is established: stable metrics against the account's own history plus above-break-even economics do not support a broken-account verdict.",
        "Does not manufacture a fix list to close the benchmark gap - no creative, targeting, or bidding overhaul recommended off the benchmark deltas.",
        "Resists the audit-for-the-board pressure explicitly: findings require account evidence, not a count of red flags to fill a deck.",
        "Recommends blended MER judged against the margin-derived break-even as the truth read, over raw platform-reported ROAS.",
        "Notes platform-reported ROAS is claimed revenue, not evidence of incremental revenue.",
        "The layer screen shows pass states with evidence rather than invented FAILs.",
        "Frames the board answer as growth versus defect: wanting better than 2.4 is an improvement program, not a defect repair, and the diagnostic will not invent a defect to justify one.",
        "Recommends product-level contribution-margin reads as where a real opportunity review would start, since blended figures can hide margin-destroying products.",
        "Sets a monitoring baseline and re-check rather than a fix plan."
      ]
    },
    {
      "id": 10,
      "prompt": "Arcline Software, a B2B workflow SaaS. Our CPA rose 45% over the last six weeks with no single day where it broke - it just drifted up. Spend is flat. I can only export campaign-level daily data; the ad-set and ad breakdowns are blocked by our agency's portal. No CRM access either - the data team says maybe next month. Also worth mentioning: we swapped the cookie consent banner about five weeks ago. I don't need anything fancy, just go through each area and tell me yes or no: tracking fine? structure fine? targeting fine? creative fine? bids fine?",
      "expected_output": "A per-layer screen that refuses the yes/no format: four states with unknowns left unknown, confidence capped due to the unreachable backend, the consent-banner swap flagged as the first tracking suspect, the gradual no-break-date drift read as the structural-erosion pattern, the exact missing access named, concrete export steps and formulas supplied, and no fixes recommended yet.",
      "files": [],
      "expectations": [
        "Declines the yes/no format and answers per layer in four states (pass, FAIL, unknown, not applicable).",
        "Layers the available data cannot power are marked unknown and explicitly not treated as clean.",
        "States that with the backend unreachable every finding carries at most medium confidence, and says so in the verdict itself.",
        "Flags the consent-banner swap five weeks ago as a tracking-layer suspect that must be reconciled before other layers are read.",
        "Reads the gradual, no-break-date CPA drift as the structural-erosion/fragmentation pattern, distinct from a single-event cause.",
        "Names the exact missing access that would clear each unknown: a CRM or backend export, and ad-set plus ad-level breakdowns.",
        "States that localisation cannot complete at campaign-only granularity, and what the finer breakdown would unlock.",
        "Asks for, or explicitly requires, the change and edit log - what changed and exactly when - before any verdict.",
        "Provides concrete export steps and spreadsheet formulas for the user to run, including the delta formula (current - baseline) / baseline.",
        "Does not issue a definitive single-layer root-cause verdict on this data.",
        "Tags per-layer entries with severity and confidence as separate fields.",
        "Emits a compact state block (baselines, screen states, open unknowns, re-check date) the user can carry into the next session.",
        "Recommends no fixes - no budget, bid, creative, or structural changes - before the unknowns are resolved."
      ]
    }
  ],
  "trigger_queries": [
    { "query": "our google ads suddenly stopped converting last month, can you figure out why", "should_trigger": true },
    { "query": "audit my ad account", "should_trigger": true },
    { "query": "CPA has doubled since January and nothing obvious changed", "should_trigger": true },
    { "query": "why did our ROAS drop from 3.2 to 1.9", "should_trigger": true },
    { "query": "our facebook ads used to print money, now they barely break even", "should_trigger": true },
    { "query": "run a full paid media account audit for our ecommerce store", "should_trigger": true },
    { "query": "cost per lead keeps creeping up every month and I can't tell why", "should_trigger": true },
    { "query": "the agency wants more budget but performance keeps getting worse", "should_trigger": true },
    { "query": "something is wrong with our ad account, results fell off a cliff two weeks ago", "should_trigger": true },
    { "query": "we're wasting ad spend somewhere, help me find where", "should_trigger": true },
    { "query": "review our campaign structure, I think something is off", "should_trigger": true },
    { "query": "diagnose why our meta ads performance degraded", "should_trigger": true },
    { "query": "our ads stopped working and I don't know if it's the creative or the targeting", "should_trigger": true },
    { "query": "conversions dropped 40% overnight and platform support is useless", "should_trigger": true },
    { "query": "why is my cost per acquisition rising when we haven't touched anything", "should_trigger": true },
    { "query": "boss wants to know why leads got so expensive this quarter", "should_trigger": true },
    { "query": "we spend $60k a month on paid and results keep sliding, where do I even start", "should_trigger": true },
    { "query": "is our ad account broken or is it just the market", "should_trigger": true },
    { "query": "everyone says raise the budget, but I want to know what's actually wrong first", "should_trigger": true },
    { "query": "linkedin ads worked great in Q1, dead in Q3, help me work out what changed", "should_trigger": true },
    { "query": "can you do a health check on our ppc account", "should_trigger": true },
    { "query": "our new marketing hire says the whole account is set up wrong, can you verify", "should_trigger": true },
    { "query": "why would ROAS fall when we changed nothing", "should_trigger": true },
    { "query": "figure out whether our performance drop is seasonal or something we broke", "should_trigger": true },
    { "query": "our cost per demo went from $180 to $420, walk me through finding the cause", "should_trigger": true },
    { "query": "second opinion on our paid search account, the numbers look worse every week", "should_trigger": true },
    { "query": "should I be worried that CPM is up 30% across the whole account", "should_trigger": true },
    { "query": "our ads suddenly got expensive, tell me if it's us or the auction", "should_trigger": true },
    { "query": "the dashboard says our campaigns are fine but revenue from ads keeps falling", "should_trigger": true },
    { "query": "inherited an underperforming ads account at my new job, where do I start", "should_trigger": true },
    { "query": "conversion volume halved but spend stayed flat, what happened", "should_trigger": true },
    { "query": "our shopify store's paid traffic stopped converting, need to know the root cause", "should_trigger": true },
    { "query": "before we rebuild everything, can you diagnose what's actually failing in the account", "should_trigger": true },
    { "query": "the google rep keeps pitching fixes but the account has been declining for months, can you look at it properly", "should_trigger": true },
    { "query": "why do our ads perform worse every month even though we keep optimizing", "should_trigger": true },
    { "query": "our lead gen campaigns collapsed right after the site relaunch, is it related", "should_trigger": true },
    { "query": "CPL up 3x in six weeks, need answers before the board meeting", "should_trigger": true },
    { "query": "what's tanking our ad performance, creative fatigue or audience burnout?", "should_trigger": true },
    { "query": "I think the algorithm broke our campaigns, can you check", "should_trigger": true },
    { "query": "paid social results dropped and my CMO wants a root cause by friday", "should_trigger": true },
    { "query": "our account's efficiency has been eroding slowly for six months", "should_trigger": true },
    { "query": "did our ads stop working because of the iOS privacy changes or did we mess something up", "should_trigger": true },
    { "query": "help me figure out why the campaigns that used to be profitable aren't anymore", "should_trigger": true },
    { "query": "spent double this month for the same number of sales, why", "should_trigger": true },
    { "query": "what's wrong with my adwords account", "should_trigger": true },
    { "query": "our b2b ads generate tons of leads but sales says they're all junk", "should_trigger": true },
    { "query": "impressions are steady but conversions keep sliding, diagnose it for me", "should_trigger": true },
    { "query": "we doubled budget and got fewer conversions, make it make sense", "should_trigger": true },
    { "query": "why are we suddenly losing money on ad spend after two profitable years", "should_trigger": true },
    { "query": "my ecommerce ads went from 4x to 1.5x return, need a structured diagnosis", "should_trigger": true },
    { "query": "our campaigns never leave the learning phase and performance is all over the place, what's the underlying issue", "should_trigger": true },
    { "query": "prospecting campaigns are fine but retargeting collapsed, what would cause that", "should_trigger": true },
    { "query": "can you review why our ad results don't match last year despite bigger budgets", "should_trigger": true },
    { "query": "performance dropped right after we restructured the account, was it the restructure or something else", "should_trigger": true },
    { "query": "tell me if we actually need new creatives or if something else is broken", "should_trigger": true },
    { "query": "the account is bleeding money and the team keeps guessing at fixes, need a real diagnosis", "should_trigger": true },
    { "query": "why is our cost per click fine but our cost per sale terrible now", "should_trigger": true },
    { "query": "quarterly review time: figure out what degraded in our paid accounts", "should_trigger": true },
    { "query": "give me a root cause analysis of our declining ad performance", "should_trigger": true },
    { "query": "my ads worked until the consent banner update, now the results look awful, what's really going on", "should_trigger": true },
    { "query": "audit our wasted ad spend and tell me what's actually causing it", "should_trigger": true },
    { "query": "check that our conversion pixel fires correctly before we launch next week", "should_trigger": false },
    { "query": "meta says 320 purchases but shopify shows 210, explain the gap", "should_trigger": false },
    { "query": "which of our 30 campaigns should we merge so they exit learning faster", "should_trigger": false },
    { "query": "is this specific ad worn out? its CTR has been dropping for 3 weeks", "should_trigger": false },
    { "query": "build a negative keyword list from this search terms report", "should_trigger": false },
    { "query": "design the audience targeting plan for our new product launch", "should_trigger": false },
    { "query": "how should I split $50k a month between google, meta and tiktok", "should_trigger": false },
    { "query": "our winning campaign is ready to scale, how fast can I raise the budget", "should_trigger": false },
    { "query": "should we use target CPA or maximize conversions for this new campaign", "should_trigger": false },
    { "query": "are we on track to spend the monthly budget or are we under-pacing", "should_trigger": false },
    { "query": "audit the landing page our ads send traffic to", "should_trigger": false },
    { "query": "write 10 ad copy variants for our spring sale", "should_trigger": false },
    { "query": "compute our CAC from this spend data and tell me if it's healthy", "should_trigger": false },
    { "query": "set maximum CAC and minimum ROAS guardrails for the marketing org", "should_trigger": false },
    { "query": "which ad platforms should a b2b devtools startup even be on", "should_trigger": false },
    { "query": "score these five video ad hooks and tell me which deserve budget", "should_trigger": false },
    { "query": "design an a/b test plan for our new creative concepts", "should_trigger": false },
    { "query": "build a retargeting sequence from our funnel data", "should_trigger": false },
    { "query": "which customers should seed our lookalike audience", "should_trigger": false },
    { "query": "write ugc scripts for our skincare product", "should_trigger": false },
    { "query": "set up a competitor ad swipe file for the team", "should_trigger": false },
    { "query": "write a job description for a senior media buyer", "should_trigger": false },
    { "query": "how do I move from ppc specialist to head of growth", "should_trigger": false },
    { "query": "what newsletters and podcasts should a media buyer follow", "should_trigger": false },
    { "query": "write a creative brief for our video editor", "should_trigger": false },
    { "query": "map the buying committee for our enterprise abm ads", "should_trigger": false },
    { "query": "should we run our CEO's linkedin posts as paid ads", "should_trigger": false },
    { "query": "adapt our ad copy for AI chat assistant placements", "should_trigger": false },
    { "query": "which ad formats fit an app install objective", "should_trigger": false },
    { "query": "the platform numbers and GA4 never match, quantify how much is timing vs definitions", "should_trigger": false },
    { "query": "plan the migration so merging our ad sets doesn't reset learning", "should_trigger": false },
    { "query": "why did organic traffic drop 40% after the google update", "should_trigger": false },
    { "query": "audit our site's SEO, rankings fell hard this month", "should_trigger": false },
    { "query": "diagnose why our email open rates collapsed", "should_trigger": false },
    { "query": "our website conversion rate dropped, run a CRO audit", "should_trigger": false },
    { "query": "our churn spiked last quarter, find the root cause", "should_trigger": false },
    { "query": "sales pipeline dried up, diagnose our outbound motion", "should_trigger": false },
    { "query": "our app store conversion rate is falling, what's wrong", "should_trigger": false },
    { "query": "figure out why our SaaS trial-to-paid rate dropped", "should_trigger": false },
    { "query": "audit our google analytics setup for gaps", "should_trigger": false },
    { "query": "my google ads account got suspended, how do I appeal", "should_trigger": false },
    { "query": "help me structure a brand new ad account from scratch", "should_trigger": false },
    { "query": "what's the industry benchmark CTR for facebook ads in fashion", "should_trigger": false },
    { "query": "forecast next quarter's ad performance for the budget plan", "should_trigger": false },
    { "query": "build the media plan for our product launch", "should_trigger": false },
    { "query": "how do I set up enhanced conversions in google ads", "should_trigger": false },
    { "query": "write the monthly client report for our ppc account", "should_trigger": false },
    { "query": "reduce our AWS bill, cloud costs doubled", "should_trigger": false },
    { "query": "diagnose why our kubernetes pods keep crashing", "should_trigger": false },
    { "query": "review my resume for a performance marketing role", "should_trigger": false },
    { "query": "how many conversions does smart bidding need before I switch to tROAS", "should_trigger": false },
    { "query": "set the right frequency caps for each retargeting stage", "should_trigger": false },
    { "query": "pick which of these two hero images we should test first", "should_trigger": false },
    { "query": "translate our ad copy for the german market", "should_trigger": false },
    { "query": "what UTM naming convention should we use across campaigns", "should_trigger": false },
    { "query": "design a geo holdout to measure incrementality of our brand search campaigns", "should_trigger": false },
    { "query": "negotiate better rates with our ad agency", "should_trigger": false },
    { "query": "our competitor launched aggressive ads, build a counter campaign", "should_trigger": false },
    { "query": "how much budget do we need to hit 500 leads a month", "should_trigger": false },
    { "query": "clean up the duplicate conversion actions in google ads for me", "should_trigger": false },
    { "query": "explain why meta's attribution changed after iOS 14", "should_trigger": false }
  ]
}
references/examples.md›
# Worked Examples

Three diagnoses in the verdict-block shape, each with the tempting wrong read it replaces. Numbers are illustrative account data, not benchmarks. Findings are listed in the efficiency order of SKILL.md's Prioritisation section, each tagged `[severity, confidence, outcome bought, effort, owner]`.

## Table of Contents

- [Example 1 - B2C ecommerce: looks like creative fatigue, is a tracking break](#example-1---b2c-ecommerce-looks-like-creative-fatigue-is-a-tracking-break)
- [Example 2 - B2B SaaS: cheap leads, empty pipeline - the wrong conversion event](#example-2---b2b-saas-cheap-leads-empty-pipeline---the-wrong-conversion-event)
- [Example 3 - the false positive: a "collapse" that is attribution lag plus a learning reset](#example-3---the-false-positive-a-collapse-that-is-attribution-lag-plus-a-learning-reset)

## Example 1 - B2C ecommerce: looks like creative fatigue, is a tracking break

Situation: DTC store, paid social, ~$60K/month. Reported ROAS fell from 3.1 to 1.8 in ten days. The team's instinct: "the creatives are burnt out, we need new ads" - the ads _are_ eight weeks old, which makes the story feel right.

Decomposition says otherwise: CPM flat vs the 60-day baseline, CTR flat, AOV flat - the entire drop is in reported CVR, and it fell across every campaign, every audience, cold and retargeting alike, starting the same Tuesday. Fatigue does not synchronise across unrelated audiences on one date. The site changelog shows a checkout-platform migration deployed that Tuesday; the order table shows revenue down only ~4%.

```
ROOT-CAUSE VERDICT - dtc-apparel, 2026-08-12
platform(s)    : paid social | model: B2C
window         : Aug 1-10 vs baseline Jun 28-Jul 27 (lag maturity matched: yes)
volume         : $19.4K, 214 reported conversions | reconciliation gap: 41% vs order table (baseline norm: 9%)

decomposition  : CVR -42% reported; CPM +2%, CTR -1%, AOV +1% - single failing link
localisation   : uniform across all campaigns and audiences, common start date Aug 5

layer screen
  measurement/tracking : FAIL - purchase event missing on new checkout domain; gap jumped 9%→41% on deploy date  [critical, high]
  structure            : pass - no changes, volumes above learning gates                                          [-, high]
  targeting            : pass - frequency and reach trends unchanged                                              [-, high]
  creative             : pass - CTR flat vs each creative's own baseline                                          [-, high]
  bidding/budget       : unknown - delivery now optimising on starved signal; re-screen after fix                 [medium, low]
  offer & downstream   : pass - backend CVR ~flat; revenue -4% vs -42% reported                                   [-, high]
  external             : pass - CPM flat; category calm                                                           [-, medium]

confidence     : high - backend reconciliation is direct evidence; single break date; uniform pattern
verdict        : measurement/tracking - conversion event lost in checkout migration
evidence       : reconciliation ratio broke on deploy date; drop uniform across unrelated audiences; upstream metrics flat
findings       : 1) restore + dedupe purchase event
                    [critical, high, restores every downstream number and the delivery signal, ~a day of dev, web dev]
                 2) re-screen bidding after 7 lag-mature days
                    [medium, low, unquantified, an hour once the window matures, media buyer]
prediction     : reported CVR recovers to ~baseline within one lag-mature week of the event firing; ROAS follows
handoff        : mbfinotti/advertising-skills@ad-conversion-tracking (fix), then re-run this diagnostic
re-check       : 2026-08-26
```

**The wrong move**: shipping new creative. It would have cost two weeks of production, reset delivery on fresh ads mid-breakage, and "failed" - because the measured CVR was broken, the new ads would report just as badly, burning the creative budget _and_ the team's trust in creative testing. Never diagnose downstream layers through a failed reconciliation gate.

## Example 2 - B2B SaaS: cheap leads, empty pipeline - the wrong conversion event

Situation: B2B SaaS, search + paid social, ~$40K/month. Dashboard looks great: CPL down 35% quarter over quarter.

Sales says the leads are junk; pipeline is flat. The tempting read: "targeting got worse, tighten the audiences."

The account optimises to raw form fills. Decomposition shows CTR and CVR-to-form _improved_ - the platform is doing exactly what it was asked: finding people who fill forms cheaply.

CRM join shows lead→SQL rate fell from 14% to 5% in the same quarter, concentrated in the campaigns that shifted spend toward the cheapest-CPL audiences. This is the Happy Cog failure mode - "leads look great in dashboard, sales say trash" - and per Swydo, cost per closed-won, not CPL, is the KPI that tells the truth in B2B.

```
ROOT-CAUSE VERDICT - b2b-saas, 2026-08-12
platform(s)    : search + paid social | model: B2B
window         : May-Jul vs baseline Feb-Apr (lag maturity matched: yes - 90-day windows per long cycle)
volume         : $118K, 1,240 leads, 74 SQLs | reconciliation gap: 6% on form fills (backend = CRM)

decomposition  : CPL -35%, but cost per SQL +61%; failing link is post-conversion quality, not the funnel to form
localisation   : concentrated in campaigns optimising to form-fill with broadest audiences

layer screen
  measurement/tracking : pass - form event reconciles at 6%; but no offline/CRM outcome feeds back to platforms   [-, high]
  structure            : FAIL - optimization event is raw form fill; platform rewarded for junk volume            [high, high]
  targeting            : pass-with-note - drift is the *symptom* of the event choice, not an independent cause    [medium, medium]
  creative             : pass - stable engagement, no decay pattern                                               [-, medium]
  bidding/budget       : pass - targets met; the targets measure the wrong thing                                  [-, high]
  offer & downstream   : pass - demo-page CVR stable for the SQLs that do arrive                                  [-, medium]
  external             : n/a - no cost-side anomaly to explain                                                    [-, -]

confidence     : high - CRM join is direct evidence; pattern tracks spend shift; volume clears the gate on 90-day windows
verdict        : structure - optimising to a conversion event the business does not value
evidence       : cost per SQL up while CPL down; SQL-rate collapse concentrated where the cheap-lead spend went
findings       : 1) feed CRM outcomes (SQL/closed-won) back to platforms; optimise to a qualified event
                    [high, high, recovers most of the SQL-cost delta and keeps paying, ~a week of CRM wiring,
                     marketing ops]  - Koda: the offline feedback loop "consistently improves lead quality
                     more than any targeting adjustment"
                 2) judge campaigns on cost per SQL at 4-6 week maturity, not CPL
                    [medium, high, judgement that tracks revenue instead of form volume, near-zero, media buyer]
prediction     : CPL rises, lead→SQL rate recovers toward ~14%, cost per SQL falls within 2 windows of the event switch
handoff        : mbfinotti/advertising-skills@ad-conversion-tracking (offline import wiring); targeting redesign only
                 if drift persists after the event fix - mbfinotti/advertising-skills@ad-audience-targeting
re-check       : 2026-10-15 (one 4-6 week B2B window, lag-mature)
```

**The wrong move**: tightening targeting first. The platform would keep hunting cheap form fills inside the narrower audience, CPL would rise, quality would stay junk - and the "fix" would look like it made things worse, inviting the next reflex: more budget.

## Example 3 - the false positive: a "collapse" that is attribution lag plus a learning reset

Situation: lead-gen account, ~$9K/month, low volume (~55 conversions/month). Monday panic: "conversions fell off a cliff last week - the account is broken, should we double the budget to compensate?"

The screen: the trailing 7 days always under-report (conversions attribute over a multi-week window - recent days are immature by construction); the account's own history shows every trailing week "down" ~30% before maturing flat. And the budget was raised 40% eight days ago - a significant edit widely treated as resetting learning, inside which 20-50% day-to-day swings are normal (Niblin). Conversion volume in the panic window: 9 - far below any read that survives the noise band, and below Google's ≥30-conversions evaluation reference for Target CPA.

```
ROOT-CAUSE VERDICT - leadgen-local, 2026-08-12
platform(s)    : search | model: B2B
window         : Aug 4-10 vs baseline Jul 1-28 (lag maturity matched: NO - comparison window immature)
volume         : $2.1K, 9 conversions in window | reconciliation gap: unverified (no CRM access this run)

decomposition  : reported CVR -33% - but inside the account's own historical immature-week band
localisation   : not meaningful at this volume

layer screen
  measurement/tracking : unknown - no backend access; ratio history unavailable          [medium, low]
  structure            : pass - unchanged                                                [-, medium]
  targeting            : pass - unchanged                                                [-, medium]
  creative             : unknown - 9 conversions cannot power a creative read            [-, low]
  bidding/budget       : unknown - +40% budget edit 8 days ago; learning likely reset    [medium, medium]
  offer & downstream   : pass - no site changes logged                                   [-, medium]
  external             : pass - CPM flat                                                 [-, medium]

confidence     : low - immature window, sub-noise volume, unverified reconciliation
verdict        : insufficient evidence - observed "collapse" fully explainable by attribution lag + learning reset
evidence       : every historical trailing week shows the same immature dip; edit date precedes the volatility
findings       : 1) wait: judge only a lag-mature window ≥14 days post-edit
                    [high, high, avoids paying for variance, near-zero - freeze and wait, media buyer]
                 2) get CRM read access before the next diagnostic run
                    [medium, -, unblocks reconciliation on every later run, an hour of access admin, ops]
gate math      : at ~2 conversions/day, ≈3 more weeks are needed for the noise band to shrink below a 30% delta
prediction     : none issued - a prediction on this data would be noise laundered as analysis
handoff        : none - no fix is recommended, because no defect is established
re-check       : 2026-08-31, mature window, edits frozen until then
```

**The wrong move**: doubling the budget "to compensate". Another significant edit would reset learning again, extend the volatile window, run CPAs 20-50% hotter through it (Grow With Sakib), and - because the trailing week always looks bad - the dashboard would "confirm" the account is broken, justifying the next panic edit. The right move costs nothing: freeze, wait for maturity, then diagnose. The most expensive failure mode in low-volume accounts is not a defect - it is treating variance as a defect and paying for the fix.
references/layer-evidence.md›
# Layer Evidence

One section per diagnostic layer, in the fixed order. Each gives the evidence that confirms the layer as the root cause, the evidence that rules it out, and the single check that settles the call.

Every comparison is against the account's own history, never an industry benchmark. All named thresholds are attributed practitioner reference points, not universal constants.

## 1. Measurement / tracking

- **Confirms**: conversions drop across every campaign and channel simultaneously on the same date (a real market does not turn everything off at once - a tag change does); reported-vs-backend gap above roughly 40% (OnlyDeb's reference point), or a reconciliation ratio that swings week to week instead of holding stable; duplicate or inactive conversion actions; conversion events firing on button click instead of confirmed outcome; a site release, consent-banner change, or tag-manager publish dated at the break.
- **Rules out**: the drop is isolated to one campaign, ad set, or segment while siblings hold (Clixtell - an isolated drop points at targeting, creative, or the page, not tracking); the backend confirms the decline at the same magnitude; the reconciliation ratio is unchanged from the account's historical norm.
- **Settling check**: trace one real test event end-to-end - action → browser/server request → platform receipt → deduplication → report → CRM record - and reconcile a lag-mature 7+ day window against backend truth by both event date and processing date. Ratio stability, not exact equality, is the pass criterion. A recent-days-only gap is attribution lag, not breakage.
- **Handoff**: `mbfinotti/advertising-skills@ad-conversion-tracking` for the fix; `mbfinotti/advertising-skills@ad-attribution-gap` when the finding is cross-platform disagreement rather than a broken pipe.

## 2. Structure

- **Confirms**: CPA drifting up steadily over months with no single break date (Precisionly); many campaigns or ad sets each starving below the conversion volume automated bidding needs (reference points: roughly 15-30 conversions/month per campaign, OnlyDeb; Google recommends ≥30 for Target CPA evaluation); overlapping ad sets bidding on the same users; efficient campaigns capped by budget while inefficient ones spend freely; the account optimising to an event the business does not value; hyper-segmentation fragmenting the learning signal (Perpetual Traffic ep. 804 reports removing state-level geo splits cut CAC 20-25% within a week).
- **Rules out**: a prior period of good performance under the identical structure with a more recent break date (Foxwell) - structure did not change, so something else did; per-unit volume comfortably above the learning thresholds.
- **Settling check**: map spend, conversion volume, and optimization event per campaign/ad set. If most units sit below the volume gates, or two units serve the same audience with separate budgets, structure FAILs.
- **Handoff**: diagnosis only - the merge/consolidation plan belongs to `mbfinotti/advertising-skills@ad-campaign-consolidation`.

## 3. Targeting

- **Confirms**: audience overlap warnings; narrow cold audiences saturating as spend competes with itself (AdStellar); frequency climbing while reach flattens; CPM rising in specific ad sets while the account's other audiences hold; healthy CTR but traffic that never converts anywhere (wrong intent, not wrong ad).
- **Rules out**: broad, healthy-sized audience with normal frequency but poor CVR - that points past the ad to the page or offer (Pigeon Digital); decline uniform across unrelated audiences (points external or to tracking).
- **Settling check**: per-ad-set frequency and reach trend vs the account's own history, plus overlap inspection. Saturation shows exposure concentrating; mis-targeting shows engagement without downstream quality.
- **Handoff**: `mbfinotti/advertising-skills@ad-audience-targeting` for redesign; `mbfinotti/advertising-skills@ad-negative-keywords` when search terms show the mismatch.

## 4. Creative

- **Confirms**: CTR falling while CPM holds flat (Metamktgagency - the auction is unchanged; the ad is losing the click); relevance/engagement diagnostics below the account's norm; decline concentrated in the oldest creatives while newer ones hold; frequency above roughly 3 with declining CTR (AdStellar's reference point).
- **Rules out**: stable CTR with rising CPM (auction/external, not creative - Metamktgagency); decline equally present in a fresh creative launched into the same window; CVR collapse with healthy CTR (downstream of the click).
- **Settling check**: per-creative CTR trend against each creative's own baseline. This skill only _names_ the layer - the full differential (baseline, confounder screen, decay measurement) runs in `mbfinotti/advertising-skills@ad-creative-fatigue`, which hands back here if the confounders point elsewhere.

## 5. Bidding / budget

- **Confirms**: bid strategy mismatched to conversion volume (a target-based strategy fed too few conversions never exits learning); a recent significant edit dating the volatility (budget/bid/target changes reset learning - CPAs run 20-50% higher during learning, Grow With Sakib); lost impression share concentrated in lost-to-budget with strong efficiency (genuinely capped); or lost-to-rank while budget goes unspent (bids/quality, not money - Workshop Digital); targets moved repeatedly without waiting out the recalibration window.
- **Rules out**: impression share stable vs history; no significant edits in the window; delivery smooth and budgets pacing normally.
- **Settling check**: the lost-IS split (budget vs rank) plus the edit log against the volatility dates. Trustworthy Digital's decision reference: lost-to-budget above 50% argues the constraint is money; lost-to-rank above 50% argues it is rank - and above roughly 60-80% impression share, diminishing returns make more budget the wrong buy either way.
- **Handoff**: `mbfinotti/advertising-skills@ad-bidding-strategy`, `mbfinotti/advertising-skills@ad-budget-pacing`, `mbfinotti/advertising-skills@ad-spend-allocation`, `mbfinotti/advertising-skills@paid-media-scaling`. This skill makes no budget or bid recommendation itself.

## 6. Offer & downstream

- **Confirms**: CPM and CTR healthy but CVR down (Pigeon Digital - "points past the ad, onto the page and the offer"); the drop dating to a price change, promo end, page release, or checkout change; the problem reproducing on a direct walk through the funnel; healthy CVR but poor revenue outcome (AOV/mix shift, not the account at all).
- **Rules out**: CVR stable while upstream metrics moved; backend conversion rate from other traffic sources unchanged is _not_ a full rule-out (paid traffic can hit a different page or geo) - check the paid path specifically.
- **Settling check**: CVR by landing page and by date against the site's change log. If the failing link is post-click, this skill stops: flag it, with the evidence, and hand everything past the click - message match, friction, the fix list - to `mbfinotti/advertising-skills@paid-landing-page-audit`.

## 7. External

- **Confirms**: the cost metric elevated evenly across every campaign and audience (AdStellar - "if CPM is elevated across every campaign evenly, the issue is likely external"); timing aligned with known seasonality or market events; The HQ Digital's test - the account's CPM rose during a period when every advertiser in the category was bidding, versus competitors flat while the account's costs rose (internal); competitor entry visible in auction-insight-style reports.
- **Rules out**: elevation concentrated in one branch of the account; competitors' pressure flat while the account's costs rose; any unresolved FAIL in layers 1-6 - external is a diagnosis of exclusion and cannot be claimed over an unchecked internal layer.
- **Settling check**: the uniformity test from the breakdown-and-compare step, corroborated by at least one external signal (seasonality calendar, auction insights, category evidence). Verdict `external` carries the account's realistic floor for the period - the finding is "wait, or re-set targets to the market", never "spend through it" without incrementality evidence.
references/metric-decomposition.md›
# Metric Decomposition

How to localise the failing link before naming a cause. The outcome metrics (CPA, ROAS) are outputs, not levers - decompose them first, then read direction.

## The identity

Pigeon Digital's decomposition: ROAS is an output of four levers.

```
impressions × CTR              = clicks
clicks      × CVR              = conversions
conversions × AOV              = revenue
spend        = impressions × CPM / 1000
=> ROAS  ≈ (CTR × CVR × AOV) / (CPM / 1000)      - impressions cancel
   CPA   = spend / conversions = (CPM / 1000) / (CTR × CVR)
```

Three levers push ROAS up (CTR, CVR, AOV), one pushes it down (CPM). A ROAS or CPA move _must_ be expressible as a move in at least one of the four - find which one(s) actually moved vs the account's own history before any layer talk. Judge against economics, not vanity: break-even ROAS ≈ 1 / gross-margin rate (AdDogs), so a "drop" that stays above break-even is a different conversation from one below it.

## Sequential elimination down the chain

Walk impression → click → conversion → revenue and stop at the first broken link (Pigeon Digital's chaining logic):

1. **CPM moved, CTR held** → the auction changed: external pressure, seasonality, saturation, or relevance decay. Run the uniformity test (below).
2. **CTR fell, CPM held** → the ad is losing the click: creative layer (Metamktgagency). Distinguish all-clicks CTR from link/outbound CTR - engagement can mask falling intent clicks.
3. **CVR fell, CTR and CPM held** → "points past the ad, onto the page and the offer" (Pigeon Digital) - offer & downstream layer - _unless_ tracking broke (a tag failure looks exactly like a CVR collapse; the reconciliation gate must already have passed) or targeting shifted the traffic mix to lower-intent clickers.
4. **AOV/value fell, all else held** → mix shift, promo, pricing - usually not an ad-account problem at all.

Two links moving together is information, not noise:

- CPM up _and_ CTR down is the classic creative-decay signature.
- Everything down simultaneously on one date is the classic tracking signature.

## Direction table

| Observation (vs own history)             | First reading                   | Layer to check first                     |
| ---------------------------------------- | ------------------------------- | ---------------------------------------- |
| CPM up uniformly, everywhere             | Auction/seasonality             | External                                 |
| CPM up in one ad set only                | Audience saturating or narrowed | Targeting                                |
| CPM up + CTR down together               | Relevance decay                 | Creative                                 |
| CTR down, CPM flat                       | Ad losing the click             | Creative                                 |
| CTR fine, CVR down                       | Post-click or signal            | Offer/downstream - after tracking passed |
| CVR down across all channels at once     | Broken conversion signal        | Measurement/tracking                     |
| CPA up slowly over months, no break date | Erosion, fragmentation          | Structure                                |
| CPA up sharply from a known date         | Whatever changed that date      | Edit log first                           |
| Volume down, efficiency stable           | Delivery constrained            | Bidding/budget (lost-IS split)           |
| Revenue down, conversions stable         | AOV/mix shift                   | Offer - often not the account            |

## Lost impression share: budget vs rank

Search impression share = impressions won / impressions eligible. Its loss splits into two numbers with opposite fixes (Adalysis; Workshop Digital):

- **Lost IS (budget)** - delivery stopped because the budget capped out. More budget genuinely buys more of the same delivery _only if_ efficiency at the margin holds.
- **Lost IS (rank)** - the ad lost the auction on rank (bid × quality). "Simply increasing your budget won't guarantee a 100% search impression share if your ad rank is insufficient" (Workshop Digital) - budget does nothing here.

Trustworthy Digital's reference points:

- Lost-to-budget above 50%: the constraint is money, and often the right move is _reducing_ bids to buy more clicks at the same spend.
- Lost-to-rank above 50%: the constraint is rank.
- 60-80% impression share: where diminishing returns typically begin for mid-market non-brand terms.

All three are practitioner reference points to test against the account, not laws.

## Frequency

Frequency is a lagging, confirming signal - an average over a window, at ad-set level, not the marginal effect of the next impression. AdStellar's reference point:

- Frequency above roughly 3 with declining CTR supports a creative/saturation read.
- Frequency climbing while unique reach flattens month over month is the earlier saturation tell.

Never build a verdict on a frequency number alone, and never compare cold-audience frequency against retargeting frequency - tolerated exposure differs by an order.

## Learning-phase mechanics

Automated delivery recalibrates after significant edits (targeting, placement, optimization event, creative, bid strategy, meaningful bid/budget changes - and budget moves beyond roughly 20% in one step are widely treated as significant, per practitioner interpretation relayed by Modern Marketing Institute and Grow With Sakib). During recalibration, performance is volatile by design: ROAS swings of 20-50% day to day are normal (Niblin), and CPAs run 20-50% above post-learning levels (Grow With Sakib).

Reference exit points:

- Roughly 50 optimization events in 7 days on Meta (Meta Business Help Center, as relayed by AdStellar).
- Up to three weeks or 1-2 conversion cycles for Google Smart Bidding (Google Ads Help).

Diagnostic consequences:

- Date every edit before reading any trend: a "collapse" that starts at an edit date is a reset, not a root cause.
- Data from inside a learning window fails the Evidence Gate.
- A unit that never accumulates the exit volume is a _structure_ finding (consolidate the signal), not a bidding finding.

## Breakdown-and-compare

The localisation move that separates internal from external: pull the failing metric by campaign, then ad set, then ad, and compare each against its own historical baseline (AdStellar; The HQ Digital).

- **Uniform elevation** across every branch → the cause is above the account: auction, season, market. The HQ Digital's discriminator: costs rising while the whole category was bidding is external; costs rising while competitors stayed flat is internal.
- **Concentrated** in one branch → the cause lives in that branch; drill one level down and repeat until the smallest failing unit is found.

Comparability rules, non-negotiable:

- Equal window lengths.
- Equal attribution-lag maturity (recent days under-report - a trailing window always looks like decline).
- Matched day-of-week composition.
- Known outages excluded.
- Never sum or compare conversion counts across differing attribution windows, counting methods, or models - report side by side until definitions reconcile.

Breaking any of these manufactures a phantom root cause.
SKILL.md›
---
name: ad-account-diagnostic
description: "Diagnose the root cause of an underperforming paid ad account - why the ads stopped working, why CPA went up, why ROAS dropped - instead of defaulting to 'increase the budget'. Weighs tracking, account structure, targeting, creative, bidding and budget, offer, and external forces, and returns a prioritised verdict with evidence and confidence. Covers B2B and B2C on search, paid social, video, and native. Use whenever the user mentions an ad account audit, a campaign structure review, wasted ad spend, rising cost per lead, or says their ads used to work - even if they never say 'diagnostic'. Diagnosis only, and it stops at the click: post-click page problems are mbfinotti/advertising-skills@paid-landing-page-audit and creative decay is mbfinotti/advertising-skills@ad-creative-fatigue."
license: MIT
metadata:
  author: Maya-Beth Finotti
  version: "1.4.5"
---

# Account Diagnostic

Find the root cause of an underperforming ad account before anyone touches it. The whole discipline is refusing the reflex fix: an account that "stopped working" has at least seven candidate failure layers, and the most common responses - raise the budget, swap the creative, blame the algorithm - each treat one unproven hypothesis as a verdict. This skill works the layers in a fixed order, measurement first, because every downstream number is read off the conversion signal: as Adspirer puts it, the checklist is "ordered the way an experienced PPC manager actually works it: measurement first (because every other number is wrong if tracking is broken)".

It decomposes the metric chain to localise the failing link before naming a cause, compares the account against its own history rather than industry benchmarks, and refuses to issue any verdict the data cannot support. The output is a diagnosis with confidence, evidence, and a prioritised handoff plan - never an executed change.

This skill stops at the account boundary; everything past the click hands off elsewhere, and this skill never becomes a budget skill - it exists precisely because "increase the budget" is the wrong default.

- **Past the click** (message match, page friction, the page fix list): `mbfinotti/advertising-skills@paid-landing-page-audit`. This skill only flags "the evidence points downstream of the click" and hands off.
- **Fragmentation verdict** (merge plan): `mbfinotti/advertising-skills@ad-campaign-consolidation`.
- **Creative-layer verdict** (decay analysis): `mbfinotti/advertising-skills@ad-creative-fatigue`, which hands back here when the cause is not creative - the two are reciprocal.
- **Tracking verdict** (fix checklist): `mbfinotti/advertising-skills@ad-conversion-tracking`.
- **Cross-platform discrepancy quantification**: `mbfinotti/advertising-skills@ad-attribution-gap`.
- **Search-term mining**: `mbfinotti/advertising-skills@ad-negative-keywords`.
- **Targeting design**: `mbfinotti/advertising-skills@ad-audience-targeting`.
- **Budget and bidding decisions**: `mbfinotti/advertising-skills@ad-budget-pacing`, `mbfinotti/advertising-skills@ad-spend-allocation`, `mbfinotti/advertising-skills@paid-media-scaling`, and `mbfinotti/advertising-skills@ad-bidding-strategy`.

## Interview

Ask before opening any data. One question per message; offer multiple-choice answers where possible; skip anything already supplied or visible in the data.

- Which platform(s) does the account run on? (search / paid social / video / native / several)
- B2B or B2C/ecommerce?
- Monthly spend, and roughly how many conversions per month on the event the account optimises to? (Decides whether the Evidence Gate can clear at all.)
- What window is under suspicion - when did performance degrade, and against which prior period is it being judged?
- What changed, and exactly when? (Budget, bids, creative, targeting, optimization event, landing page, site release, consent banner, price, promo calendar - dates matter more than the list.)
- What can you export, and at what granularity? (Per-campaign per-day is the working minimum; per-ad-set and per-ad breakdowns unlock the localisation step.)
- What is the target CPA/ROAS - and is it derived from unit economics or inherited from a dashboard?
- Is CRM or backend revenue data reachable for reconciliation? (Yes, directly / yes, via someone / no.)
- Which attribution window and counting settings does reporting use, and did they change in the window?
- Who will implement the fixes, and what is their effort ceiling - hours available, whose hands (developer, marketing ops, creative studio, page owner), and how reversible a change is allowed to be? (The handoff plan is addressed to them; see Prioritisation for how a ceiling reshapes the queue.)
- By what date does the recovery have to show up in reporting? (A hard deadline promotes fast-acting fixes - tracking repair, unblocking a genuinely budget-capped campaign - and demotes offline-outcome wiring and creative programmes, whose payoff lands a window or two later.)
- One-off recovery, or a compounding asset? (A compounding mandate promotes the offline-outcome feedback loop and structural consolidation above the quick wins; see Prioritisation.)

## Workflow

1. Run the Interview; collect every answer before touching data.
2. **Reconcile before interpreting.** Compare platform-reported conversions against CRM/backend truth over a lag-mature window, against OnlyDeb's practitioner reference points:
   - Gap above roughly 40%: something is broken - stop, record a tracking FAIL, and hand off.
   - Gap within roughly 10%: proceed to the other layers.

   These are attributed starting points, not laws - the account's own historical ratio, and whether that ratio is _stable_ week to week, is the real criterion. If backend data is unreachable, every later finding carries at most medium confidence, and the verdict must say so.

3. **Decompose the failing metric** with [references/metric-decomposition.md](references/metric-decomposition.md). Chain the funnel - impressions → clicks → conversions → revenue - and express the outcome as CPM × CTR × CVR × AOV (Pigeon Digital) to find which single link actually moved. Name no cause before the failing link is localised.
4. **Break down and compare.** Pull the failing metric by campaign, then ad set, then ad, and compare each against the account's own history - never industry benchmarks. AdStellar's rule for CPM generalises: elevated evenly across everything points external; concentrated in one branch points to a targeted internal cause. This step is what separates internal from external causes.
5. **Walk the layers in fixed order** (below) using [references/layer-evidence.md](references/layer-evidence.md). For each layer record one of four states - `pass` (evidence rules it out), `FAIL` (evidence confirms it), `unknown` (evidence missing), `not applicable` - with its own severity and confidence. Never collapse to a binary: `unknown` is not `pass`, and treating an unchecked layer as clean is the most common way audits go wrong.
6. **Apply the Evidence Gate** (below). If it fails, the verdict is `insufficient evidence`: state exactly what additional days, spend, or access would clear it, and stop. No verdict on underpowered data.
7. **Issue the Root-Cause Verdict** (block below; worked versions in [references/examples.md](references/examples.md)). A compound verdict - two unrelated confirmed causes - is legitimate; rank them.
8. **Rank the findings by efficiency** (Prioritisation section) and write the handoff: which sibling skill or owner takes each finding, in rung order.
9. **Log the prediction.** Every verdict must state which metric should move, in which direction, by roughly how much, once the recommended fix ships - this is the raw material for the Measuring section.
10. Set a re-check date one full comparison window (at equal attribution-lag maturity) after the fix ships.
11. If your harness has persistent memory, memorise the account's baselines, the per-layer screen, the verdict, fixes applied, predictions, and re-check dates, so the next run starts from history instead of re-deriving it. If it does not, emit a short state block the user can paste into the next session.

If your harness can read the exports or run the calculations, compute every step directly; otherwise emit the exact export steps and spreadsheet formulas (columns, ratio, delta `(current - baseline) / baseline`) for the user to run and report back.

## The Layer Order

Fixed order, because each layer's evidence is only readable if the layers before it hold:

- **Tracking** corrupts every number.
- **Structure** corrupts the data pooling and learning that targeting and creative reads depend on.
- **Bidding** reads assume the auction inputs above it are sane.
- **Offer & downstream** sits past the click.
- **External** causes are a diagnosis of exclusion, claimed only when internal evidence rules the others out - the uniformity test in step 4 can fast-path there, but never skip the tracking check to get to it.

Baker's 8-layer framework states the why: auditing creative on an account whose pixel feeds 56% wrong data "produces conclusions that look right but are mathematically meaningless" - and across B2B SaaS accounts spending $3K-$250K/month he reports 80% of underperformance tracing to pixel health and creative diversity, not bidding or budget. This is a sequence, not a menu: the layers are not alternatives to pick between, so they carry no efficiency ranking - every one gets screened, in this order. The ranking lives one step later, over the _fixes_ the screen produces (see Prioritisation).

1. **Measurement / tracking** - is the conversion signal real, deduplicated, and consented?
2. **Structure** - fragmentation, overlapping campaigns self-competing, budget traps, wrong optimization event.
3. **Targeting** - audience overlap, saturation, too narrow/broad, wrong intent.
4. **Creative** - resonance decay; the differential itself runs in `mbfinotti/advertising-skills@ad-creative-fatigue`.
5. **Bidding / budget** - strategy-volume mismatch, learning state, impression share lost to budget vs rank.
6. **Offer & downstream** - price, promo, page, checkout; flag and hand to `mbfinotti/advertising-skills@paid-landing-page-audit`.
7. **External** - auction inflation, seasonality, competitor entry; confirmed by uniform elevation plus market evidence.

Per-layer confirming evidence, ruling-out evidence, and the check that settles each: [references/layer-evidence.md](references/layer-evidence.md).

## Evidence Gate

Refuse a verdict the data cannot carry. All four checks must pass before any layer FAIL becomes a verdict:

- **Conversion volume.** Enough conversions in both windows that the observed delta exceeds noise (band ≈ `p ± 2 × sqrt(p × (1-p) / n)` on the relevant rate). Attributed reference points, not laws: Google recommends evaluating automated bidding over periods with at least 30 conversions for Target CPA and 50 for Target ROAS (Google Ads Help), and OnlyDeb notes campaigns under roughly 15-30 conversions/month lack signal for the algorithm at all - low-volume accounts (most B2B) gate on leading metrics instead and say so in the confidence line.
- **Learning-phase state.** No verdict from data collected while delivery is recalibrating. Reference points: Meta ad sets typically exit learning after roughly 50 optimization events in 7 days (Meta Business Help Center, as relayed by AdStellar), and Google Smart Bidding can take up to three weeks or 1-2 conversion cycles (Google Ads Help). Niblin: ROAS swinging 20-50% day-to-day is normal during learning - judge only after exit or 7+ days.
- **Window comparability.** Baseline and comparison windows must have equal duration, equal attribution-lag maturity (recent days always under-report; a trailing window always "looks like" decline), matched day-of-week composition, and no known outages. Jyll Saskin Gales audits all active campaigns together over a 90-day window because campaigns "interact with each other" - never judge one campaign in isolation over a cherry-picked week.
- **Attribution consistency.** Never sum or compare conversions across differing attribution windows, counting methods, or models - a 7-day-click figure plus a 30-day-click figure is not a total, it is a category error. Report side by side until definitions reconcile; a settings change mid-window invalidates the comparison entirely.

When the gate fails: verdict `insufficient evidence`, plus the arithmetic of what clears it - the additional days or spend needed for `n` to shrink the noise band below the observed delta at current traffic, or the specific access (CRM export, ad-level breakdown) that is missing.

## Root-Cause Verdict

Deliver one block per diagnosis:

```
ROOT-CAUSE VERDICT - <account>, <date>
platform(s)    : <...> | model: B2B | B2C
window         : <comparison window> vs baseline <baseline window> (lag maturity matched: yes/no)
volume         : <spend>, <conversions> conversions in window | reconciliation gap: <x% vs backend | unverified>

decomposition  : <which link moved: CPM / CTR / CVR / AOV - direction and size vs own history>
localisation   : <uniform across account | concentrated in: campaign/ad set/ad>

layer screen
  measurement/tracking : pass | FAIL | unknown | n/a - <evidence>   [severity, confidence]
  structure            : pass | FAIL | unknown | n/a - <evidence>   [severity, confidence]
  targeting            : pass | FAIL | unknown | n/a - <evidence>   [severity, confidence]
  creative             : pass | FAIL | unknown | n/a - <evidence>   [severity, confidence]
  bidding/budget       : pass | FAIL | unknown | n/a - <evidence>   [severity, confidence]
  offer & downstream   : pass | FAIL | unknown | n/a - <evidence>   [severity, confidence]
  external             : pass | FAIL | unknown | n/a - <evidence>   [severity, confidence]

confidence     : high | medium | low - <basis: gate math, reconciliation state, unknowns count>
verdict        : <root cause layer(s), or insufficient evidence - with the one-sentence causal story>
evidence       : <the two or three observations that carry the verdict>
findings       : <ranked by efficiency - each with severity, confidence, outcome bought, effort, owner>
prediction     : <metric expected to move, direction, rough size, by when>
handoff        : <sibling skill or owner per finding, in sequence>
re-check       : <date - one full window after fix, lag-mature>
```

## Prioritisation

Rank findings before handing off - a flat list of twenty observations is a pitch deck, not a diagnosis. Rank by efficiency: the outcome each fix buys per unit of effort it costs. Never by cheapest, and never by the layer screen's own row order - the screen is a diagnostic sequence, the ranking is a work queue.

Default fix-class order, highest value/effort ratio first:

1. **Tracking repair** - buys back the readability of every other number in the account and puts delivery on a real signal again; costs hours to a day of one developer. Always ships first: every later fix is measured through it.
2. **Optimization-event / offline-outcome fix** - buys a platform that chases qualified outcomes instead of cheap volume (Koda: the offline feedback loop "consistently improves lead quality more than any targeting adjustment"); costs about a week of marketing-ops and CRM wiring, then keeps paying without further work.
3. **Structural consolidation** - buys starved units the conversion volume they need to exit learning; costs a week of rebuild plus one learning window of volatility, and un-merging later is expensive.
4. **Targeting redesign** - buys reach on audiences that are not yet saturated; costs about a week, and reverses cleanly.
5. **Offer & downstream fixes** - the outcome can beat everything above it, but the work sits outside the account: costs coordination with whoever owns the page, the price or the checkout, so their calendar sets the pace instead of the auditor's.
6. **Creative replacement** - buys the click back where decay is confirmed; costs a production cycle and never ends - a standing job, not a fix.
7. **Budget and bid moves** - buy nothing until every rung above passes, and each significant edit spends a learning window (see "Why 'Increase the Budget' Is the Wrong Default").

The axes disagree, which is exactly why the reflex fix keeps winning arguments:

- efficiency: tracking > event fix > consolidation > targeting > offer & downstream > creative > budget moves
- value: offer & downstream == tracking > event fix > consolidation > creative > targeting > budget moves
- effort: offer & downstream > creative > event fix == consolidation > targeting > tracking > budget moves
- compliance cost: event fix > tracking > every other rung (none) - pushing CRM outcomes into an ad platform triggers a data-processing and consent-basis review and cannot be un-sent; re-firing a conversion event touches the consent banner and the platform's data-use terms.

Budget moves sit last on effort and last on efficiency at once: near-zero effort buying near-zero outcome is not a cheap win, it is a rounding error with a learning reset attached.

**What this order starves: offer & downstream.** It ties tracking for the highest value on the page and carries the highest effort, so the ratio pushes it to rung 5 - and the default-rung rule then reaches it only after four other rungs pass or ship. An account can run this diagnostic for a year and never once touch the price, the offer or the page. Promote it to the top on either condition: the decomposition localises the failure past the click (CVR collapsed while CPM, CTR and impression share held), or the page, price and checkout owners sit on the user's own team - its effort was coordination cost, not work, and in-house ownership deletes that cost.

Default rung: start at the highest rung whose layer the screen marked `FAIL`, and never below rung 1 while tracking is `FAIL` or `unknown`. Move one rung down only once the rung above is `pass` or already shipped.

The ordering is a default, not a law - it shifts with the account and with who executes it. Re-rank it against what you already know here:

- An in-house developer makes rung 1 a same-day job.
- A creative studio already on retainer promotes rung 6 above its default place.
- A business that tolerates no learning-window volatility demotes rungs 3 and 7.

The Interview's deadline, one-off-versus-compounding and effort-ceiling answers move rungs the same way.

A constraint the user actually stated does something different from re-ranking: it removes the rung. Delete a ruled-out fix class from the ranked findings rather than parking it at the bottom, and name it as deleted in the handoff with the constraint that killed it and what would revive it - "offline-outcome loop: deleted, no CRM access anywhere in scope; revive when an export exists".

A fix left sitting last on a list nobody will reach is indistinguishable from one nobody has considered, and it comes back next quarter as a fresh idea. One exception: the fix the verdict names as the root cause is never deleted. It escalates to whoever can lift the constraint - a tracking FAIL under an effort ceiling with no developer time is an escalation, not a demotion and not a deletion.

Tag every finding with:

- **Severity**: what it costs if unfixed.
- **Confidence**: how sure the evidence is, kept separate from severity.
- **Outcome bought**: the fix's recoverable value (Conner Crowe's discipline of attaching one to each fix); mark "unquantified" when the data cannot say.
- **Effort**, as an order of magnitude: near-zero, an hour, a week, a quarter, a standing job.
- **Owner**: from the Interview's implementer answer.

Score with ICE (Impact, Confidence, Ease) by default. Reach for RICE (Reach × Impact × Confidence / Effort - Sean McBride at Intercom) only when the ranking must survive someone who did not run the diagnostic: `efficiency: ICE > RICE`, `value: RICE > ICE only when the plan has to be defended`, since ICE costs minutes and reproduces most of RICE's ordering while RICE costs an extra reach estimate and buys defensibility, not accuracy. Hand the plan off in rung order, each finding with its owner attached.

## Why "Increase the Budget" Is the Wrong Default

The most common prescription, and the one this skill exists to refuse. Four mechanisms, each sourced:

- **Learning resets.** Budget increases beyond roughly 20% in one edit are widely treated as significant edits that restart the learning phase on Meta (practitioner interpretation relayed by Modern Marketing Institute and Grow With Sakib); Grow With Sakib estimates each reset costs 5-15% ROAS the following week, with CPAs running 20-50% higher during learning. More money into a broken account buys a worse version of the same problem.
- **Budget vs rank.** Lost impression share splits into lost-to-budget and lost-to-rank (Adalysis; Workshop Digital) - and "simply increasing your budget won't guarantee a 100% search impression share if your ad rank is insufficient" (Workshop Digital). Diagnose the split first; if the loss is to rank, budget does nothing.
- **Diminishing returns.** Trustworthy Digital observes diminishing returns above roughly 60-80% impression share, where the next increment costs more than the qualified revenue it returns. Near the top of impression share, more budget buys the worst impressions in the auction.
- **Correlation is not incrementality.** Platform-reported conversions are claimed, not caused: "Platform ROAS, Google Analytics, and last-click all report claimed revenue. None of them show that a single dollar was incremental" (Stella). Lifesight reports a grocery chain's geo holdout on non-brand paid search finding a 0% sales lift. Scaling a channel because its dashboard looks good scales the claim, not the revenue. Ben Heath says it plainly: "Big budgets don't guarantee results. Great creative and a strong offer are a much better place to focus."

A budget recommendation is only ever this skill's output as a _handoff_ - to `mbfinotti/advertising-skills@ad-spend-allocation` or `mbfinotti/advertising-skills@paid-media-scaling` - after the diagnosis shows delivery is genuinely budget-capped and everything upstream passes.

## B2B vs B2C

The method is identical in both: same layer order, same decomposition, same gate, same verdict shape. What diverges is the data regime and which numbers tell the truth.

| Dimension           | B2B                                                                                                                                                                                                                                                                                                   | B2C/ecommerce                                                                                                                                                                                             |
| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Evaluation window   | Conversion volume is low and the sales cycle long (GrowthSpree cites an 84+ day average, with SQL/pipeline outcomes landing 30-90 days after the tracked conversion); evaluation windows stretch to 4-6 weeks minimum before data is reliable (Koda), and the Evidence Gate leans on leading metrics. | Volume supports the full statistical gate and daily reads.                                                                                                                                                |
| Dominant root cause | Optimising to the wrong conversion event: cheap form fills the platform can find in volume, while "leads look great in dashboard, sales say trash" (Happy Cog).                                                                                                                                       | Platform ROAS is "systemically inflated post-iOS14" (Eightx).                                                                                                                                             |
| Fix class           | The offline feedback loop: feeding CRM outcomes (MQL/SQL/closed-won) back to the platform, which Koda reports "consistently improves lead quality more than any targeting adjustment".                                                                                                                | Read blended measures instead of raw platform ROAS: MER (total revenue / total marketing spend; Superscale) judged against a margin-derived break-even (roughly 1 / contribution-margin %, per Meerkats). |
| Truth-telling KPI   | Cost per SQL and cost per closed-won - "the only KPI that tells the truth in B2B" (Swydo) - never CPL.                                                                                                                                                                                                | Contribution margin by product, since blended ROAS "hides the products destroying your margins" (Jordan Glickman).                                                                                        |
| Source of truth     | CRM.                                                                                                                                                                                                                                                                                                  | The order table.                                                                                                                                                                                          |

In both models, the reconciliation step in the Workflow is the same act, only the source of truth differs.

## Auditor Failure Modes and Biases

The account is not the only thing under diagnosis - so is the auditor. Screen the auditor's own read against this table before issuing a verdict.

| Trap                                        | Why it burns                                                                                                                                                                                              | Fix                                                                                      |
| ------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| Anchoring on the previous owner's narrative | The first story heard becomes the hypothesis everything gets fitted to (inStreamly)                                                                                                                       | Ignore the inherited narrative until the data independently reproduces it                |
| Recency bias                                | Overweighting the last few days; daily noise mimics collapse                                                                                                                                              | Zoom out to weekly/monthly views before reading any intraday move (Go4Trades; Niblin)    |
| Confirmation bias                           | Data gets interpreted to support the strategy already in place (inStreamly)                                                                                                                               | Write the disconfirming evidence for the leading hypothesis before concluding            |
| Goodhart / metric fixation                  | Optimising the score, not the objective - Google's Optimization Score reaches 100% just by dismissing recommendations (Store Growers)                                                                     | Treat platform scores as means, never as the health verdict (Goodhart, 1984)             |
| Over-trusting platform-reported numbers     | "Meta claimed one number, Google claimed another, and the actual backend revenue was a third number entirely" (HYROS); every self-attributing platform reports as if it were the only channel (Enalitica) | Reconcile against backend truth first; never sum platform claims                         |
| The pitch-audit incentive                   | A free audit sold before a proposal is structurally biased toward "everything is wrong" (Conner Crowe)                                                                                                    | Findings need evidence and an outcome-per-effort estimate each, not a count of red flags |
| Fixing the layer where the symptom appears  | The visible metric is downstream of the cause; a CVR drop is rarely a CVR problem                                                                                                                         | Trace backward through the chain; fix at the source layer                                |
| One hypothesis, tested by fixing it         | Stacked unverified fixes destroy the baseline and the attribution of recovery                                                                                                                             | One falsifiable hypothesis at a time; predict the metric move before shipping the fix    |
| Benchmark shopping                          | Industry averages have different mix, geography, and definitions                                                                                                                                          | The account's own history is the baseline; benchmarks are directional context only       |
| Best-practice cargo-culting                 | Checklists applied without account context - there are "many different ways to structure a successful account" (Foxwell)                                                                                  | Every finding must cite this account's evidence, not a generic rule                      |

## Measuring Whether This Worked

This skill's KPI is diagnostic accuracy, tracked on a rolling log of every verdict and its step-9 prediction:

- **Prediction hit rate**: share of verdicts where the predicted metric moved in the predicted direction (at meaningful size, judged at the lag-mature re-check) after the recommended fix - and only that fix - shipped.
- **Overturn rate**: share of verdicts later overturned - a second diagnosis, an incrementality test, or the fix's failure showing the named root cause was wrong.

Starting floor - this skill's own working target, not a researched constant: iterate on the method until at least 70% of high-confidence verdicts hit their prediction and fewer than 15% of all verdicts are overturned, then tighten from the account's own log.

- A low hit rate with a clean gate usually means the layer evidence is being read too loosely.
- A high overturn rate concentrated in one layer means that layer's ruling-out checks are too weak.

Verdicts issued despite an `unverified` reconciliation line should be tracked separately - if they overturn more often, that is the argument for insisting on backend access next time.

## Reference

- Read [references/layer-evidence.md](references/layer-evidence.md) when walking the layer screen - per layer: confirming evidence, ruling-out evidence, and the check that settles it.
- Read [references/metric-decomposition.md](references/metric-decomposition.md) when localising the failing link - the CPM × CTR × CVR × AOV identity, the funnel chain, what each metric direction means, lost impression share, frequency, and learning-phase mechanics.
- Read [references/examples.md](references/examples.md) when writing the verdict block - a B2C case where the cause is tracking, a B2B case where the cause is the wrong conversion event, and a convincing false positive the Evidence Gate catches, plus the wrong move alongside the right one.