SKILL DETAIL
cac-roas-benchmark
mbfinotti/advertising-skills/cac-roas-benchmark
Compute a business's CAC and ROAS family metrics from real spend and outcome data, then judge whether the spend is healthy against three references - the business's own break-even, its own trailing history, and provenance-labelled external benchmarks - returning healthy, watch, unhealthy, or insufficient evidence. Use whenever the user asks whether their CAC is too high, whether a ROAS is good, what ROAS to aim for, whether ads are profitable, or mentions MER, blended CAC, payback period, or a spend health check - even if they never say 'benchmark'. Covers B2B/SaaS (cost per SQL, cost per closed-won, cohort lag) and B2C/ecommerce. Do NOT use to set those thresholds as policy - use mbfinotti/advertising-skills@ad-spend-guardrails instead.
Installation
npx skills add https://github.com/mbfinotti/advertising-skills --skill cac-roas-benchmark
技能檔案
SKILL.md
最近同步 · 2026年9月24日
evals/evals.json›
{
"skill_name": "cac-roas-benchmark",
"evals": [
{
"id": 1,
"prompt": "I run Ferndale Goods, a DTC home-fitness-gear brand. July numbers: net revenue from our order system $240,000, paid media spend $96,000, plus about $14,000 in agency fees and creative production. Contribution margin after COGS, shipping and payment fees is 38%. Our Meta + Google dashboards show a combined ROAS of 4.6, which beats the 4:1 industry standard, and I read the ecommerce median is only around 2. Monthly MER since January if it helps: 2.62, 2.55, 2.48, 2.41, 2.33, 2.24, and then July. Are we healthy? Thinking about scaling next month.",
"expected_output": "A spend health report that recomputes the money-anchored figures (blended ROAS 2.5, MER ~2.18), sets break-even at ~2.63 from the 38% margin, treats the 4.6 as platform-reported claim, names 4:1 as folklore, and returns an unhealthy verdict resting on break-even and the deteriorating trend, with scaling refused and the budget decision handed off.",
"files": [],
"expectations": [
"Computes blended ROAS as 240,000 / 96,000 = 2.5 and labels it as the blended variant, distinct from the dashboard figure",
"Computes MER as approximately 2.18 (240,000 / 110,000), with the denominator explicitly including the $14,000 of agency and creative spend",
"Computes break-even ROAS/MER as 1 / 0.38, approximately 2.63",
"Treats the dashboard 4.6 as platform-reported/attributed ROAS - a claimed figure, not the number judged against break-even",
"Names the 4:1 rule as folklore with its origin: no traceable author, often misattributed to a 2016 Nielsen study that says no such thing",
"States that 4:1 is simply break-even at a 25% contribution margin, and places this business's own 2.63 break-even next to it",
"Returns an unhealthy verdict because MER (and blended ROAS) sit below the business's own break-even by arithmetic, not a healthy verdict",
"The verdict rests on break-even (rung 1) and the trailing history (rung 2); the ~2 ecommerce median is used only as context, never to set the verdict",
"Runs or explicitly requests a mix-shift check before treating the January-July MER decline as a real trend",
"Reads the MER history as a sustained multi-period deterioration rather than a one-off dip",
"Every reported figure carries a variant label, a window, and a data source",
"Any external median cited carries publisher, year, sample, and the variant it measured",
"Refuses to endorse scaling in this run and hands the pause/reallocate decision off to a spend-guardrails or budget-allocation workflow rather than prescribing budget moves itself"
]
},
{
"id": 2,
"prompt": "Ostrava Systems here - B2B compliance SaaS, ACV $18,000, gross margin 75%, sales cycle averages about 4 months. In April we spent $60,000 on paid (media plus agency), got 500 leads, 40 of them became SQLs, and only 2 deals closed that month. That's a $30,000 CAC against the published $239 B2B SaaS benchmark from First Page Sage. My cofounder wants to kill paid entirely before May. Our historical lead-to-close rate is 1.5%. Can you confirm the numbers?",
"expected_output": "A run that refuses the $30,000 in-period CAC as a cohort-mixing artifact, withholds the verdict as insufficient evidence because the window is shorter than the sales cycle, exposes the First Page Sage variant mismatch, and supplies labeled interim reads (CPL $120 vs break-even CPL $270, cost per SQL $1,500) plus a cohort measurement plan.",
"files": [],
"expectations": [
"Rejects $30,000 as a valid CAC because April's 2 closed deals came from earlier lead cohorts while April's spend bought a cohort that closes months later - in-period division mixed two cohorts",
"Names the window being shorter than the ~4-month sales cycle as an explicit evidence-gate condition",
"Returns insufficient evidence as the verdict state, not unhealthy or catastrophe",
"Flags the First Page Sage $239 as a combined organic-plus-paid figure from a sample the agency itself discloses as roughly 75% organic-weighted, making it a variant mismatch against paid-only spend",
"Does not endorse killing paid on the strength of the benchmark comparison",
"Computes CPL as 60,000 / 500 = $120, labeled with variant, window, and source",
"Computes break-even CPL as ACV x lead-to-close rate: 18,000 x 1.5% = $270",
"Computes cost per SQL as 60,000 / 40 = $1,500",
"Labels the CPL and cost-per-SQL reads as interim signals while the cohort is immature, explicitly not verdicts",
"Proposes cohort measurement: group leads by month generated and measure revenue at 90/180/365 days",
"Computes the allowable first-year CAC ceiling as ACV x gross margin: 18,000 x 0.75 = $13,500, for use once the cohort matures",
"Explicitly distinguishes in-period CAC from lagged/cohort CAC",
"Names exactly what would unlock the verdict: the April cohort reaching maturity"
]
},
{
"id": 3,
"prompt": "Loopwell here - project-management SaaS, self-serve. Two plans: Basic $15/month and Pro $99/month, gross margin 70% on both. Last month marketing spend was $93,000 and billing shows 300 'new customers,' but our billing counts plan reactivations and annual renewals as new - about 60 of the 300 are those. So CAC is $310. David Skok says payback should be under 12 months - where do we stand? Roughly two-thirds of new signups are Basic.",
"expected_output": "A per-plan payback analysis on a corrected new-customer CAC of ~$388 (renewals excluded), computed with the gross-margin term kept (~37 months Basic, ~5.6 months Pro), the 12-month rule named as Skok's 2011 rule of thumb, and separate per-segment readings instead of one blended verdict.",
"files": [],
"expectations": [
"Excludes the 60 renewals and reactivations from the denominator and recomputes new-customer CAC as 93,000 / 240, approximately $387-388, instead of accepting $310",
"States that counting renewals as acquisitions inflates the denominator and flatters every verdict built on it",
"Computes payback per plan rather than presenting a single blended payback figure",
"Keeps the gross-margin term in the payback formula: monthly gross profit per customer is $10.50 on Basic (15 x 0.70) and $69.30 on Pro (99 x 0.70)",
"Reports Basic payback at approximately 37 months",
"Reports Pro payback at approximately 5.6 months",
"Does not report the margin-free figures (~26 months Basic, ~3.9 months Pro) that come from dividing CAC by the raw subscription price",
"Names the 12-month payback rule as David Skok, 2011 - a rule of thumb tied to that era's fundraising conditions, never empirically validated",
"Puts the business's own per-plan economics next to the quoted folk target instead of deleting or simply endorsing it",
"States that one blended number describes neither plan, since the same CAC means order-of-magnitude different payback at $15/month versus $99/month",
"Delivers a separate reading per plan segment (Basic problematic, Pro comfortable) rather than a single verdict",
"Every figure carries a variant label, window, and data source",
"Flags that raw payback assumes the customer survives the payback window, and asks for retention or applies a discounted-payback view given the ~37-month Basic figure"
]
},
{
"id": 4,
"prompt": "Our ecommerce brand Bramblewick lost TikTok Ads reporting for the last three weeks of October (expired API token). TikTok is historically about 8% of our paid spend. My analyst put $0 for TikTok in the October sheet so the totals still add up. October: order-system net revenue $520,000, Meta spend $130,000, Google spend $95,000, TikTok $0 per the sheet. Contribution margin 44%. MER April through September was 2.35, 2.31, 2.34, 2.30, 2.28, 2.26. Can you compute October MER and tell me if the downtrend is real?",
"expected_output": "A run that refuses the zero-filled denominator, renormalises total spend using the known ~8% TikTok share (~$245K, MER ~2.13), notes withholding as the fallback, checks mix shift before reading the trend, and judges October against the ~2.27 break-even using the corrected figure.",
"files": [],
"expectations": [
"Refuses to compute the verdict from the zero-filled $225,000 spend total, naming the $0 TikTok line as silently corrupting the blend",
"Renormalises the spend total using the known ~8% TikTok share, arriving at roughly $245,000 total spend and an October MER of approximately 2.1",
"States that renormalising is the right move here because the missing channel is small and its size is known, with withholding the blended figure named as the fallback",
"Computes break-even MER as 1 / 0.44, approximately 2.27",
"Judges October against break-even using the corrected MER (at or below 2.27), not the zero-filled ~2.31 that would sit above it",
"Runs or explicitly requests a mix-shift check before declaring the April-October downtrend real",
"Reads the history as a multi-period deteriorating trend pending the mix-shift confirmation",
"The verdict (unhealthy or watch) rests on the business's own break-even and trailing history, with no external median setting it",
"Every reported figure carries variant, window, and data source labels",
"Names restoring TikTok reporting as a concrete next step or next-period question",
"Does not zero-fill, average around, or silently ignore the missing channel anywhere in the calculation"
]
},
{
"id": 5,
"prompt": "Prepping the Series B board deck for Nivara Health, our B2B scheduling SaaS. Marketing wants a slide saying CAC improved 35%: Q1 CAC was $95, Q2 came in at $62. Digging into the sheet: the Q1 figure was paid spend $190,000 divided by 2,000 paid-attributed customers. In Q2 the team switched to total marketing spend $217,000 divided by all 3,500 new customers, because organic took off after a PR hit. Q2 paid numbers were $185,000 spend and 1,850 paid-attributed customers. Can we run with the 35% improvement?",
"expected_output": "A refusal of the 35% claim as a silent variant switch (paid Q1 vs blended Q2), a like-for-like recomputation showing paid CAC actually worsened to $100, both variants presented labeled side by side, and a recommendation to build fully-loaded CAC for the board with finance owning the definition.",
"files": [],
"expectations": [
"Refuses the 35% improvement claim because Q1 is paid CAC and Q2 is blended CAC - two different variants that cannot be compared",
"States that blended CAC is always lower than paid CAC when organic acquisition exists, so part of the apparent improvement is purely definitional",
"Recomputes the like-for-like comparison: Q2 paid CAC = 185,000 / 1,850 = $100, meaning paid CAC worsened roughly 5% quarter over quarter",
"Presents both Q2 variants side by side, each labeled ($62 blended, $100 paid), rather than silently picking one",
"Names the silent variant switch as the failure: CAC appears to improve with nothing actually changing",
"Recommends fully-loaded CAC (S&M including salaries, tooling, agency, creative) as the variant a board or investor audience expects",
"Warns that a fully-loaded figure shown to a board becomes a reported number finance has to own, and redefining it later is a restatement - so the definition must be agreed with finance",
"Does not let the already-available blended number decide the variant choice for a finance-facing audience",
"Notes the organic surge is a mix shift, so any blended CAC trend cannot be read without a mix-shift check",
"Withholds or gates any health judgment on break-even inputs and history rather than judging health from the CAC levels alone",
"Does not present $95 to $62 as a single continuous trend line or average the two figures",
"Every figure carries a variant label, window, and data source"
]
},
{
"id": 6,
"prompt": "I run Tessellate Prints, a custom-poster ecommerce shop. Everyone gives me a different ROAS target: our agency says aim for 4x minimum, a podcast said MER should be above 4 and 5-8 at scale, and a blog said median ecommerce ROAS is around 2 so anything above that is fine. Our contribution margin after COGS, shipping and payment fees is 55%. What ROAS should we actually be aiming for?",
"expected_output": "An answer anchored on the business's own break-even of ~1.82 (1/0.55), with each quoted target named and traced to its origin rather than adopted, the margin-dependence of ROAS profit made explicit, history requested as the second reference, and the policy-setting of an actual target handed off to a guardrails workflow.",
"files": [],
"expectations": [
"Computes break-even ROAS as 1 / 0.55, approximately 1.82",
"Names the 4:1 rule's status: no traceable author, often misattributed to a 2016 Nielsen study, and mathematically just break-even at a 25% margin - not this business's 55%",
"Names 'MER above 4, 5-8 at scale' as a Taylor Holiday / Common Thread Collective stated heuristic never measured across a sample",
"Shows that the same ROAS means different profit at different margins (for example 4x yields about $1.40 profit per ad dollar at 60% margin and $0.00 at 25%, or the equivalent computation at 55%)",
"Does not adopt the ~2 median as the target, and any median quoted carries publisher, year, sample, and the variant measured",
"Explains that an external median cannot know this business's margin or price point - which is why it cannot set the verdict",
"Positions the 1.82 break-even as the non-negotiable arithmetic floor: below it the spend loses money whatever any benchmark says",
"Hands off setting the actual target or floor as written policy to a spend-guardrails workflow instead of declaring a policy number itself",
"Recommends the business's own trailing history (4-8 prior periods, same definition) as the second reference and asks for it",
"Names each quoted folk target with its origin instead of deleting or ignoring it",
"Does not average the three quoted targets into a single number",
"Distinguishes blended ROAS from MER by their denominators (total paid spend versus total marketing spend) when discussing the quoted 4x targets"
]
},
{
"id": 7,
"prompt": "CFO of Corvid Analytics, a B2B data-quality SaaS - ACV around $30K, gross margin 78%, NRR 118%. I've collected three published CAC payback medians - 16, 18 and 20 months - so let's just call the benchmark 18. Ours is 21 months, computed the same way each quarter for the last two years: 21.5, 21, 20.5, 21, 20.5, 21, 21, 21. Also our LTV:CAC is 2.7 and the rule says 3. Board meeting Thursday - how bad is it?",
"expected_output": "A run that refuses to average the conflicting medians and presents them as a conflict with provenance, names the 3:1 rule as Skok's self-admitted guess, prefers payback over LTV:CAC, reads the stable 8-quarter history as rung 2, weighs the 118% NRR before flagging the 21 months, and refuses to declare the business unhealthy from the benchmark gap.",
"files": [],
"expectations": [
"Refuses to average the 16, 18 and 20-month medians into one benchmark and presents them as a conflict instead",
"Explains the medians disagree because the panels and definitions behind them differ",
"Attaches publisher, year, and sample to any published payback median it cites",
"Names the 3:1 LTV:CAC rule as David Skok's own figure from the early 2010s, a self-admitted guess, worth at most a secondary check",
"Prefers payback over LTV:CAC as the affordability lens, giving at least one stated reason (LTV:CAC inherits whichever CAC variant fed it, hides per-plan variance, or never asks when the cash comes back)",
"Reads the 8 quarters of same-definition history as the business's own trailing reference, noting the trend is flat/stable and that direction and volatility matter more than the level",
"Weighs the 118% NRR before flagging the 21-month payback: high net revenue retention can justify a longer acquisition payback",
"Computes the monthly gross profit per customer as roughly $1,950 (30,000 / 12 x 0.78) and/or the implied CAC of roughly $41,000, labeled",
"Asks which CAC variant and spend lines feed the payback figure before comparing it to any external median",
"Reports external benchmarks as ranges or quartiles rather than a single point estimate",
"Sizes the gap to external medians as context only, never as the basis of the verdict",
"Does not declare the business unhealthy on the grounds of being above the averaged or individual published medians",
"Any verdict offered rests on the business's own economics and history, with the missing rung-1 inputs (allowable CAC versus actual CAC) either computed from the given ACV and margin or explicitly requested"
]
},
{
"id": 8,
"prompt": "Harbor & Pine, coastal-home-decor DTC. November: Meta Ads Manager reports $310,000 in attributed revenue on $80,000 spend; Google Ads reports $265,000 on $60,000. So combined ROAS is (310+265)/(80+60) = 4.1 and we're planning to push spend up 30% for December. The order system shows $390,000 net revenue for November in total, 3,800 new customers, AOV around $78, contribution margin 42%. Sanity-check the 4.1 for me?",
"expected_output": "A rejection of the summed platform ROAS (platforms jointly claim ~147% of actual revenue), a blended ROAS of ~2.79 anchored on the order system, break-even at ~2.38, blended CAC ~$37 flagged as above the ~$33 first-order contribution (verdict contingent on repeat rate), history requested for rung 2, seasonality respected, and the December scale-up handed off.",
"files": [],
"expectations": [
"Flags that Meta and Google jointly claim $575,000 of attributed revenue against $390,000 of actual revenue - roughly 147% - so the platforms are double-counting",
"Refuses to sum platform-attributed revenue across platforms, stating platform-reported ROAS is non-additive",
"Anchors revenue on the order system's $390,000 net figure, not on platform-attributed value",
"Computes blended ROAS as 390,000 / 140,000, approximately 2.79, labeled blended with the November window",
"Computes break-even ROAS as 1 / 0.42, approximately 2.38",
"Judges spend health against the business's own break-even (2.79 versus 2.38), not against the 4.1 platform composite",
"Labels the platform figures as claimed/attributed rather than caused revenue, and any attribution-inflation evidence cited carries its provenance",
"Computes blended CAC as 140,000 / 3,800, approximately $37, labeled",
"Computes first-order contribution as AOV x margin (78 x 0.42, approximately $33) and explicitly states the verdict is contingent on the repeat-purchase rate because CAC exceeds it",
"Requests 4-8 prior periods of the same metrics as the trailing-history reference before finalizing the verdict",
"Notes that December must be compared against a prior December, not against November, because seasonality alone can invert a verdict",
"Declines to make the 30% December increase decision in this run and hands the budget move off to an allocation, scaling, or guardrails workflow",
"Every reported figure carries a variant label, a window, and a data source"
]
},
{
"id": 9,
"prompt": "Meridian Straps, we sell premium watch straps online. Our Meta CAC went from $38 to $46 over the last two months - about a 20% jump. Team consensus is that Meta is saturated and we should immediately move 30% of the budget into TikTok and Pinterest to diversify. Contribution margin is 52%, AOV $95. Does the CAC jump confirm it's time to diversify?",
"expected_output": "A margin-first read: $46 still sits under the ~$49 first-order allowable CAC, so the rise is not automatically a problem; the new-channel test is gated on the current channel's marginal CAC exceeding what the next channel could deliver; two months of decline earns at most a watch with a named question, and the reallocation itself is handed off.",
"files": [],
"expectations": [
"Does not endorse immediately moving budget into untested channels on the strength of the 20% CAC rise alone",
"Checks contribution margin first: judges the $46 against the business's own allowable CAC before treating the rise as a problem",
"Computes first-order allowable CAC as AOV x margin (95 x 0.52, approximately $49) and notes $46 still sits under it, with thin headroom",
"States the decision rule for diversifying: test a new channel only once the current channel's marginal CAC exceeds what the next channel could realistically deliver - and marginal CAC has not been measured here",
"Establishes the variant of the $38-to-$46 figure as platform-attributed paid CAC and treats it as claimed rather than causal",
"Names reacting to a CAC bump by spreading spend into untested channels as the more common failure, compared with staying too long in a working channel",
"Notes that measuring marginal CAC takes a quarter and deliberate unrecoverable test spend, and that it serves scaling decisions rather than health verdicts",
"Requests longer trailing history and a seasonality check before reading two data points as a trend",
"Returns at most a watch state with a named next-period question, not an unhealthy verdict and not a reallocation instruction",
"Hands the actual budget reallocation decision off to an allocation or channel-selection workflow",
"Every reported figure carries a variant label, a window, and a data source"
]
},
{
"id": 10,
"prompt": "Head of growth at Verdala, a CRM for property managers. The CEO wants a spend health verdict by Friday for Monday's exec review. Problem: nobody has ever computed our contribution margin - finance is mid-audit and can't help until next month. What I do have: 6 months of blended CAC on a consistent definition (total marketing spend / all new customers): $410, $395, $402, $398, $405, $400, and MER for the same months: 3.1, 3.2, 3.1, 3.15, 3.1, 3.1. B2B, sales cycle about 5 weeks, monthly windows, so lag is covered. Give me the verdict.",
"expected_output": "A run that says at intake that the verdict will be withheld without contribution margin, presents the labeled blended figures and the stable history as context, refuses to assume a typical margin or substitute a benchmark for break-even, and returns insufficient evidence naming the margin as the exact unlock.",
"files": [],
"expectations": [
"States up front, at intake, that without a contribution margin the verdict will be withheld - not discovered at the end of the run",
"Returns insufficient evidence as the final state, not healthy, despite the stable history",
"Names contribution margin as the exact missing input and finance's margin number as what would unlock the verdict",
"Names the Friday deadline's effect on scope: the run stays on the near-zero-effort variants (blended CAC, MER) and any fully-loaded CAC or margin rebuild is deleted from this run and said so",
"Still computes and presents the blended CAC and MER figures, each labeled with variant, window, and source",
"Presents the 6-month history as direction and volatility context (stable/flat) while explicitly not issuing a health verdict from the trend alone",
"Does not substitute an external published benchmark to stand in for the missing break-even",
"Explicitly distinguishes insufficient evidence from unhealthy: unknown is not the same as failing",
"Does not assume or invent a margin figure (for example a 'typical SaaS' 75-80%) to force a verdict out",
"Recommends the standing investment: contribution margin costs real effort once and near-zero every period after, scheduled for when finance frees up",
"Includes a definitions record covering the intake answers and naming the variants this run deleted or deferred",
"Frames what happens next: the named missing input, who provides it, and when the verdict can be issued"
]
}
],
"trigger_queries": [
{ "query": "Is our CAC too high?", "should_trigger": true },
{ "query": "Our blended CAC jumped from $180 to $240 this quarter - should I be worried?", "should_trigger": true },
{ "query": "Is a 3.2 ROAS good for an ecommerce store?", "should_trigger": true },
{ "query": "What ROAS should we aim for given our margins?", "should_trigger": true },
{ "query": "Are our Facebook ads actually making us money?", "should_trigger": true },
{ "query": "MER came in at 2.4 this month - healthy or not?", "should_trigger": true },
{ "query": "CFO wants to know whether our paid spend is efficient. Can you run the numbers?", "should_trigger": true },
{ "query": "We spent $50k on ads and got 130 customers. Good or bad?", "should_trigger": true },
{ "query": "What's a healthy CAC payback period for a SaaS at our price point?", "should_trigger": true },
{ "query": "Sanity-check our unit economics on paid acquisition", "should_trigger": true },
{ "query": "Can you do a spend health check on our ad accounts?", "should_trigger": true },
{ "query": "Is $420 cost per customer acceptable for a $79/month product?", "should_trigger": true },
{ "query": "Our agency says 4x ROAS is the industry standard - is that true for us?", "should_trigger": true },
{ "query": "How do I know if we're overpaying to acquire customers?", "should_trigger": true },
{ "query": "Benchmark our CAC against the industry", "should_trigger": true },
{ "query": "Is our LTV to CAC ratio of 2.5 a problem?", "should_trigger": true },
{ "query": "Judge whether our Google Ads returns are sustainable", "should_trigger": true },
{ "query": "The platform says ROAS 5.1 but the bank account disagrees - are we actually profitable?", "should_trigger": true },
{ "query": "Is a $900 cost per SQL reasonable for enterprise software?", "should_trigger": true },
{ "query": "Our payback period is 19 months - too long?", "should_trigger": true },
{ "query": "Here's our spend and revenue by month - tell me if the ads are worth it", "should_trigger": true },
{ "query": "Is 2.8 MER enough to keep scaling or are we treading water?", "should_trigger": true },
{ "query": "What does a good ROAS look like for a 35% margin business?", "should_trigger": true },
{ "query": "Boss asked if our customer acquisition cost is out of line. Help?", "should_trigger": true },
{ "query": "Are we losing money on every customer we acquire?", "should_trigger": true },
{ "query": "Evaluate whether our paid social spend is healthy this quarter", "should_trigger": true },
{ "query": "I keep hearing 3:1 LTV:CAC - are we failing at 2.2?", "should_trigger": true },
{ "query": "ROAS dropped from 3.5 to 2.9 - is that still okay?", "should_trigger": true },
{ "query": "Compare our cost per acquisition to what's normal for DTC brands", "should_trigger": true },
{ "query": "$14k spend, 200 orders, $68 AOV - is this working?", "should_trigger": true },
{ "query": "Is our cost per closed-won deal defensible?", "should_trigger": true },
{ "query": "Should our ecommerce ROAS target really be 4?", "should_trigger": true },
{ "query": "How healthy is a 24-month CAC payback for enterprise SaaS?", "should_trigger": true },
{ "query": "Give me a verdict on our blended ROAS of 2.1", "should_trigger": true },
{ "query": "Check if our ad spend clears break-even", "should_trigger": true },
{ "query": "An investor asked about our CAC efficiency - what do I tell them?", "should_trigger": true },
{ "query": "We're at 1.9x return on ad spend. Panic or fine?", "should_trigger": true },
{ "query": "Are these acquisition costs normal for fintech?", "should_trigger": true },
{ "query": "Did our ads make money after refunds and fees?", "should_trigger": true },
{ "query": "Is my cost per lead of $85 too high for B2B?", "should_trigger": true },
{ "query": "How much profit are we really making per ad dollar?", "should_trigger": true },
{ "query": "My cofounder says our CAC is fine, I think it's a disaster - settle it with the numbers", "should_trigger": true },
{ "query": "Run our January-June spend and revenue through a profitability check", "should_trigger": true },
{ "query": "Is 6:1 ROAS on brand search actually as good as it looks?", "should_trigger": true },
{ "query": "Assess our customer acquisition economics before the board meeting", "should_trigger": true },
{ "query": "The dashboard says we're at 380% return - believe it?", "should_trigger": true },
{ "query": "What CAC payback should I expect at a $12K ACV?", "should_trigger": true },
{ "query": "Marketing efficiency ratio has been sliding for three months - how bad is it?", "should_trigger": true },
{ "query": "Is spending $110 to get a $95 first order ever okay?", "should_trigger": true },
{ "query": "Do our numbers justify what we spend on paid acquisition?", "should_trigger": true },
{ "query": "Grade our paid marketing performance: spend $210K, new revenue $560K", "should_trigger": true },
{ "query": "Am I reading it right that we're below break-even on ads?", "should_trigger": true },
{ "query": "Is a blended CAC of $61 good for a subscription box?", "should_trigger": true },
{ "query": "Everyone quotes different ROAS benchmarks - which one applies to us?", "should_trigger": true },
{ "query": "Our new VP says anything under 3x ROAS should be cut. Is that right?", "should_trigger": true },
{ "query": "Quick gut check: $32 CAC, $54 AOV, 41% margin - healthy?", "should_trigger": true },
{ "query": "Tell me if our ad economics survive contact with the actual margin data", "should_trigger": true },
{ "query": "Is our cost to acquire a customer sustainable long-term?", "should_trigger": true },
{ "query": "Set a maximum allowable CAC policy for the marketing org", "should_trigger": false },
{ "query": "Write our ROAS floor and kill-switch thresholds into a formal policy", "should_trigger": false },
{ "query": "Define who can override our spend guardrails and when", "should_trigger": false },
{ "query": "Derive spend limits from our runway and codify them", "should_trigger": false },
{ "query": "We need documented rules for when to shut off a campaign automatically", "should_trigger": false },
{ "query": "How should I split $60K between Meta, Google and TikTok next quarter?", "should_trigger": false },
{ "query": "Reallocate our budget from prospecting to retargeting", "should_trigger": false },
{ "query": "Which campaigns should get more money this month?", "should_trigger": false },
{ "query": "Move spend toward the channels with the best marginal return", "should_trigger": false },
{ "query": "Are we on track to spend the full monthly ad budget?", "should_trigger": false },
{ "query": "Project our month-end spend from the first ten days", "should_trigger": false },
{ "query": "Flag campaigns that are underspending against their budgets", "should_trigger": false },
{ "query": "Why did our CPA double in March?", "should_trigger": false },
{ "query": "Figure out what's wrong with our ad account - performance fell off a cliff", "should_trigger": false },
{ "query": "Diagnose why conversions dried up after the account restructure", "should_trigger": false },
{ "query": "Our ads stopped working - find the root cause", "should_trigger": false },
{ "query": "Meta says 400 conversions, our CRM says 220 - reconcile them", "should_trigger": false },
{ "query": "Explain the gap between GA4 revenue and Shopify revenue", "should_trigger": false },
{ "query": "Classify the discrepancy between platform and order-system numbers", "should_trigger": false },
{ "query": "Should I switch to target ROAS bidding on Google?", "should_trigger": false },
{ "query": "What tROAS value should I enter in Google Ads?", "should_trigger": false },
{ "query": "Manual CPC or maximize conversions for a new campaign?", "should_trigger": false },
{ "query": "When is it safe to double this campaign's daily budget?", "should_trigger": false },
{ "query": "Design a ramp plan for scaling our winning campaign", "should_trigger": false },
{ "query": "How fast can we scale spend without breaking performance?", "should_trigger": false },
{ "query": "Verify our pixel and Conversions API are deduplicating before launch", "should_trigger": false },
{ "query": "Check whether the purchase event fires twice on checkout", "should_trigger": false },
{ "query": "Which ad platforms fit a B2B startup with a $10K monthly budget?", "should_trigger": false },
{ "query": "Should we advertise on LinkedIn or Google first?", "should_trigger": false },
{ "query": "Is my winning ad wearing out? Frequency is climbing", "should_trigger": false },
{ "query": "CTR is dropping on our best creative - fatigue or seasonality?", "should_trigger": false },
{ "query": "We have 40 campaigns - which should we merge?", "should_trigger": false },
{ "query": "Mine our search terms report for negative keywords", "should_trigger": false },
{ "query": "Ads get clicks but no conversions - audit our landing page", "should_trigger": false },
{ "query": "Rank these five video hooks before we put budget behind them", "should_trigger": false },
{ "query": "Write ad copy variations for our spring sale", "should_trigger": false },
{ "query": "Build a layered targeting plan from our ICP", "should_trigger": false },
{ "query": "Which customers should seed our lookalike audience?", "should_trigger": false },
{ "query": "Design a retargeting sequence for cart abandoners", "should_trigger": false },
{ "query": "Map the buying committee for our enterprise deals", "should_trigger": false },
{ "query": "Set up an A/B test plan for our new ad concepts", "should_trigger": false },
{ "query": "Draft a creative brief for our UGC campaign", "should_trigger": false },
{ "query": "Should this campaign use carousel or video format?", "should_trigger": false },
{ "query": "Organize competitor ads into a swipe file", "should_trigger": false },
{ "query": "What should I ask when interviewing a media buyer?", "should_trigger": false },
{ "query": "How do I move from PPC specialist to growth lead?", "should_trigger": false },
{ "query": "Which paid-media newsletters and podcasts are worth following?", "should_trigger": false },
{ "query": "Where do I start with paid advertising for my new store?", "should_trigger": false },
{ "query": "Is our monthly churn rate of 4% healthy?", "should_trigger": false },
{ "query": "Benchmark our email open and click rates", "should_trigger": false },
{ "query": "Is our organic traffic growth in line with the industry?", "should_trigger": false },
{ "query": "What's the average CPC for legal keywords right now?", "should_trigger": false },
{ "query": "Build the LTV model for our fundraising deck", "should_trigger": false },
{ "query": "Is $99/month the right price for our pro plan?", "should_trigger": false },
{ "query": "Do we have enough pipeline coverage for the Q4 number?", "should_trigger": false },
{ "query": "Calculate contribution margin per SKU from these COGS numbers", "should_trigger": false },
{ "query": "Compare our website conversion rate to ecommerce averages", "should_trigger": false },
{ "query": "Forecast next quarter's revenue from current ad spend", "should_trigger": false }
]
}
references/benchmark-sources.md›
# Published CAC/ROAS benchmarks - provenance, figures, folklore
Every CAC/ROAS benchmark in circulation is either a vendor's client sample or a self-reported survey. There is no neutral, audited, industry-wide CAC or ROAS benchmark, and no standards body defines what enters the CAC numerator. Quote provenance with the number or do not quote the number.
## The evidence contract
Before citing any external figure in a report, attach all four fields - publisher, year, sample, and the metric variant it measured - plus the publisher's commercial interest when it has one. A figure missing any of these is quoted as "unverified" or dropped.
## Publishers worth naming
Rows run best-first, ordered by what a citation from each tier is worth:
- value: platform-instrumented vendor data > large self-reported surveys > agency client analytics > consultancy and AI-generated summary pages
- effort: all == (roughly an hour of sourcing per figure, whichever tier) - only value separates them, so never settle for a lower tier to save time
Platform-instrumented data is measured rather than recalled, but every tier is skewed to that publisher's own customer base: more credible is not neutral. The bottom tier is not a weaker source, it is not a source - trace the figure to the named publisher underneath it or drop it (see Circular citation below).
| Source | Type | Sample | Measured or self-reported | Commercial interest |
| ---------------------------------------------------------------- | ----------------------- | --------------------------------------------------------------- | ---------------------------- | --------------------------------------------------------------- |
| Triple Whale 2025 | platform-instrumented | 18,000+ ecommerce brands | measured | analytics vendor |
| Polar Analytics 2026 | platform-instrumented | thousands of Shopify brands, 16 industries | measured | analytics vendor |
| Varos / Billo 2025-2026 | platform-instrumented | thousands of campaigns; Billo: 80,000+ Meta video ads (H2 2025) | measured | benchmarking vendors |
| Benchmarkit (Ray Rike), 2025 B2B SaaS Performance Metrics Survey | survey | 583 participants; per-metric n = 21-149 | self-reported | advisory/data vendor |
| Aleph x Benchmarkit 2026 SaaS & AI Survey | survey | 342 companies (198 reported CAC payback), CY-2025 actuals | self-reported | FP&A vendor |
| Gartner CMO Spend Survey 2025 | survey | 402 CMOs, majority >$1B revenue, fielded Feb-Mar 2025 | self-reported | independent research |
| High Alpha (formerly OpenView) 2024 SaaS Benchmarks | survey | 800+ companies | self-reported | VC/studio |
| KeyBanc/KBCM, SaaS Capital | surveys | large private SaaS panels | self-reported | bank / lender to SaaS |
| First Page Sage, CAC by industry | agency client analytics | 120+ B2B and 103 B2C clients, Jan 2022-Aug 2025 | measured, non-representative | SEO agency; **self-disclosed 75% organic / 25% paid weighting** |
This ordering is a default that a variant match overrides: a tier-1 figure measuring a different variant is worth less than a tier-2 figure measuring the user's own. Check the variant before the tier.
OpenView ceased new investments in December 2023. Its SaaS Benchmarks Report moved to High Alpha (with Paddle). Cite 2024+ editions under the new owner.
## B2C / ecommerce figures
- Median ecommerce ROAS **2.04** - Triple Whale 2025, 18,000+ brands, vendor-instrumented. Far below the folk 4:1. Median paid-social ROAS 1.86-1.93x across ~35,000 brands in the same vendor's earlier dataset.
- By channel: paid search median ~4.5x, Meta 2.2-2.8x, TikTok ~1.4x rising to ~2.25x with value optimisation - Varos / Billo, 2025-2026, platform-instrumented campaign samples.
- Seasonality large enough to invert a verdict: Toys & Games ROAS 1.46 in July vs 2.90 in December - Billo, H2 2025, 80,000+ Meta video ads. Black Friday / Cyber Monday can collapse per-order contribution margin 40-60% (Luca, vendor analysis).
- CAC and conversion by revenue tier (Polar Analytics 2026, thousands of Shopify brands):
- the $5M-$20M revenue tier shows the worst CAC efficiency
- $100M+ brands post the highest conversion rates at 6.29%
- First Page Sage ecommerce CAC $86 (combined organic + paid, 75% organic-weighted agency-client data, Jan 2022-Aug 2025).
## B2B / SaaS figures
- CAC payback median (three panels, presented as a conflict, never averaged - they disagree because panels and definitions differ):
- **18 months in 2024**, up from 14 in 2023 (Benchmarkit 2025, 583 participants)
- **16 months FY2025**, top quartile ≤6, bottom quartile ≥24 (Aleph x Benchmarkit 2026, 198 of 342 reporting)
- **20 months in 2024**, down from 25 in 2022 (KeyBanc)
- CAC payback by ACV segment (Benchmarkit 2025/2026, replacing generic working bands with a measured, segmented one - ACV predicts payback better than any other single attribute): sub-$5K ACV ~11 months; $10K-25K ACV ~12 months; $25K-50K ACV ~14 months; mid-market ~14-18 months; $50K-100K ACV ~22 months; above $100K ACV 18-24 months; above $250K ACV ~24 months. A below-median number at high ACV is not automatically a red flag - weigh it against net revenue retention, since high NRR can justify a longer upfront acquisition period.
- New CAC Ratio: median **$2.00 of S&M spend per $1.00 of new-customer ARR** in 2024, up 14% YoY, 4th quartile $2.82 (Benchmarkit 2025). Uses the SaaS Metrics Standards Board convention: New CAC Ratio = total S&M expense ÷ new-customer ARR - a convention, not an accounting standard.
- LTV:CAC medians: 3.2:1 (Optifai, N=939, Q2 2025-Q1 2026) and 3.6:1 (Benchmarkit 2025) - the "observed median near 3" is descriptive, not evidence for the 3:1 rule.
- CAC by vertical (First Page Sage, organic-weighted client data, Jan 2022-Aug 2025) - organic CAC runs roughly half of paid in most B2B verticals, so remember the sampling bias before comparing a paid-heavy account to any of these:
- B2B SaaS $239 combined ($205 organic / $341 paid)
- construction $281
- cybersecurity $387
- financial services $784
- real estate $791
- education $1,143
- fintech enterprise $14,772
- Cost per SQL $1,357 average vs cost per lead $198 average in the same 2024 dataset (First Page Sage) - the gap between the two is the point.
- Marketing budget (Gartner 2025, n=402, mostly >$1B-revenue firms) - descriptive of what large firms _do_, not what any firm _should_ spend:
- 7.7% of company revenue, down from 9.5% three years prior
- paid media 30.6% of marketing budget
## Attribution inflation (context for platform-reported ROAS)
- Observational/platform attribution overstated true lift by factors of ~7.0, 9.5, and 7.6 for upper/mid/lower-funnel outcomes vs randomized ground truth - Gordon, Zettelmeyer, Bhargava & Chapsky 2019, Marketing Science 38(2), 15 Facebook RCTs, ~500M user-experiment observations.
- Account audits show Meta and Google jointly claiming 150-200% of real revenue via double-counting - AdBeacon, vendor client audits (indicative, not audited).
- Global iOS tracking opt-in sat near 38% in Q1 2026 (Adjust benchmark, reported via SignalSeal), so most mobile attribution is modelled rather than deterministic - another reason a platform figure is a claim, not a measurement.
## Why cross-report comparison is unsound
- **Definitional mismatch.** Benchmarkit, First Page Sage, KeyBanc, and Triple Whale each use materially different numerators, customer definitions, and attribution scopes. Comparing your blended CAC to a published paid CAC measures the definitions.
- **Survivorship and self-selection.** Survey respondents opt in, failed companies are absent, and a VC or lender panel is a portfolio, not a market.
- **Disclosed sampling bias.** First Page Sage weights its blend 75% organic because it is an SEO agency - its own methodology note says so. A 90%-paid business cannot compare to it directly.
- **Small n.** Several Benchmarkit metrics rest on n = 21-43 (Blended CAC Ratio n=43, Expansion CAC Ratio n=21). Report medians and quartiles, never a point estimate.
- **Circular citation.** Many recent consultancy and AI-generated benchmark pages cite each other in loops. Trace a figure to a named publisher with a stated sample, or drop it.
## Folklore - full origins list
None of these traces to a study. Name the origin when a user quotes one, and put the business's own break-even next to it.
| Rule | Originator | Evidence status |
| ------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 3:1 LTV:CAC | David Skok (Matrix Partners), ~2011-2013, "SaaS Metrics 2.0" | Self-admitted guess - on record at SaaStr: "I guessed at that number, after visiting many, many SaaS companies." |
| 12-month CAC payback | David Skok, 2011 | Rule of thumb tied to 2011 fundraising conditions; never empirically validated. |
| 4:1 (4x) ROAS | No traceable author; often misattributed to a 2016 Nielsen study | The Nielsen analysis names no universal 4:1 and stresses category variance. 4:1 is break-even at a 25% contribution margin (1 ÷ 0.25), retroactively declared a target. |
| MER > 4, rising to 5-8 at scale | Taylor Holiday (Common Thread Collective) | Stated heuristic, never measured across a sample. |
| CAC ≈ 25% of gross profit | Taylor Holiday (CTC, 2022), "fuel profit" framing | Same status. |
| SaaS Magic Number 0.75 | Rory O'Driscoll (Scale VP) anecdote, ~2005; popularised by Lars Leckie, 2008 | Single-anecdote origin; Benchmarkit 2025 reports a 0.90 median. |
| "Spend 5-10% of revenue on marketing" | US Small Business Administration "general rule", ~2013 | Folklore; contradicted by the SBA's own measured ~1% actual small-business ad spend. The descriptive ~7.7% is Gartner measuring large-company behaviour. |
| 3x pipeline coverage | Unattributed sales folklore | Silently assumes a 33% win rate; needs ~4x at 25%, 4-7x for enterprise. |
| 3-12 month payback working band (18 at enterprise) | Practitioner-reported across paid acquisition, superseded above by Benchmarkit's measured ACV-segmented figures for B2B SaaS | Keep this row only for non-SaaS paid-acquisition businesses the segmented table doesn't cover; recalibrate against the business's own cohorts and cash position. |
| 2-3x ROAS as "good and achievable" (Google Ads) | Jyll Saskin Gales, ex-Google, "What is a Good ROAS in Google Ads?" | Stated heuristic explicitly conditioned on the business's own margin and LTV, not a study - she frames "good" as "profitable, more money out than in," never a fixed number. |
| LTV of at least $10K, and a $5K/month minimum budget, to justify LinkedIn Ads at all | AJ Wilcox (B2Linked), "How Much Do LinkedIn Ads Cost?" | Stated practitioner heuristic, not a measured benchmark - paired with his own observed CPC range ($8-14, up from $6-8 in 2020 and $2 in 2011) and the claim that LinkedIn's cost-per-SQL runs under half of Facebook's despite the higher front-end CPC. |
The circularity to watch for: a blog post quoting a _descriptive_ survey median (Gartner's 7.7%, Benchmarkit's 18-month payback) as a _prescriptive_ target. Descriptive averages describe what a panel did. They carry no claim about what this business should do.
references/worked-examples.md›
# Report template and worked examples
All business figures below (spend, revenue, margins, customer counts) are **invented illustrations of the method** - never quote them as benchmarks. The only external numbers are the cited published figures, which carry their provenance.
## Table of Contents
- [Report template](#report-template)
- [Headline](#headline)
- [Definitions record](#definitions-record)
- [Metric table](#metric-table)
- [Comparison ladder](#comparison-ladder)
- [Verdict and evidence gate](#verdict-and-evidence-gate)
- [Folklore appendix (only if raised)](#folklore-appendix-only-if-raised)
- [Handoffs](#handoffs)
- [Example 1 - B2C ecommerce (home-goods DTC brand, "Maple & Loam")](#example-1-b2c-ecommerce-home-goods-dtc-brand-maple-loam)
- [Example 2 - B2B SaaS (workflow vendor, "Quorline")](#example-2-b2b-saas-workflow-vendor-quorline)
## Report template
```markdown
# Spend Health Check - {business}, {window}
## Headline
- {segment/channel}: **{healthy | watch | unhealthy | insufficient evidence}** - rests on rung {1/2/3 used}. {one line why}
## Definitions record
Model: {B2B/B2C} · Window: {dates, lag maturity} · Spend lines: {…} ·
New customer = {…}, renewals {included/excluded} · Revenue basis: {source, gross/net} ·
Contribution margin: {x%} ({how derived}) · History: {n periods available}
## Metric table
| Metric | Variant | Value | Window | Source |
| ------ | ------- | ----- | ------ | ------ |
## Comparison ladder
1. Break-even: {value} - gap: {…}
2. Own history: {last 4-8 periods} - direction: {…}
3. External: {figure} ({publisher, year, sample, variant measured}) - context only
## Verdict and evidence gate
Gate: variant established {y/n} · margin known {y/n} · window ≥ lag {y/n} · channels complete {y/n}
Verdict: {state}. {conditions met / what was withheld and why}
## Folklore appendix (only if raised)
{quoted rule} - origin: {…} - this business's own break-even: {…}
## Handoffs
{finding} → {mbfinotti/advertising-skills@skill}
```
## Example 1 - B2C ecommerce (home-goods DTC brand, "Maple & Loam")
**Inputs (order system + billing, July, lag-mature).**
- Net revenue $180,000
- 1,500 orders
- AOV $120
- 900 new customers
- Paid media spend $60,000
- Total marketing spend $72,000 (adds agency $8,000, creative $4,000)
- Contribution margin after COGS, shipping, and payment fees: 45%
### Positive reading - the method applied
| Metric | Variant | Value | Window | Source |
| ------------------------- | ----------------------------------------- | ----- | ------ | --------------------------- |
| Blended ROAS | total net revenue ÷ total paid spend | 3.0 | July | order system + billed spend |
| MER | total net revenue ÷ total marketing spend | 2.5 | July | order system + finance |
| Blended CAC | total marketing spend ÷ all new customers | $80 | July | finance + order system |
| Break-even ROAS/MER | 1 ÷ 0.45 | 2.22 | - | own margin data |
| First-order allowable CAC | AOV × margin = $120 × 0.45 | $54 | - | own margin data |
Ladder:
1. **Break-even 2.22** - MER 2.5 clears it. But blended CAC ($80) exceeds first-order contribution ($54): acquisition only pays back through repeat purchases, so the verdict is contingent on the repeat rate holding. Named explicitly.
2. **Own history** - blended CAC Apr-Jul: $70 → $74 → $78 → $80. MER: 2.8 → 2.7 → 2.6 → 2.5. Four consecutive periods of deterioration. Mix-shift check: channel mix stable, so the drift is real.
3. **External** - median ecommerce ROAS 2.04 (Triple Whale 2025, 18,000+ brands, vendor-instrumented blended figure). This brand's 3.0 sits above the median - context only. It changes nothing about rungs 1-2.
**Verdict: watch.** Above break-even on matured data, but a four-period deteriorating trend and a CAC above first-order contribution.
Next-period question: does 90-day repeat contribution close the $26 gap? Handoff: none yet, unless the trend crosses 2.22, in which case `mbfinotti/advertising-skills@ad-account-diagnostic`.
### Negative reading - same business, misread
> "The ad platform dashboard shows ROAS 4.3. That beats the 4:1 target, and the industry median is only 2.04 - we're crushing it. Scale."
What went wrong, in order:
1. **Variant mismatch.** 4.3 is _platform-reported_ ROAS (attributed, view-through included, modelled credit). The money-anchored blended figure is 3.0. The two were compared as if interchangeable.
2. **Folklore as target.** "4:1" has no traceable author and is just break-even at a 25% margin - this brand's margin is 45%, so its real break-even is 2.22.
3. **Benchmark-only verdict.** The median (rung 3) was used to declare health while rung 1 was never computed and rung 2's four-period deterioration was never looked at.
4. **The contingency vanished.** Nobody noticed CAC exceeds first-order contribution, so "scale" doubles down on a bet on repeat behaviour that was never examined.
Same numbers, opposite conclusion: "crushing it, scale" vs "watch, with a named question."
## Example 2 - B2B SaaS (workflow vendor, "Quorline")
**Inputs.**
- ACV $12,000
- Gross margin 80%
- Sales cycle ~10 weeks
- March lead cohort, measured at 180-day maturity: paid spend $48,000 (media + agency, labeled)
- 400 leads
- 60 SQLs
- 8 closed-won
### Positive reading - the method applied
| Metric | Variant | Value | Window | Source |
| -------------------------- | --------------------------------- | ---------- | ------------------ | -------------- |
| CPL | paid spend ÷ leads | $120 | March cohort @180d | platform + CRM |
| Cost per SQL | paid spend ÷ SQLs | $800 | March cohort @180d | CRM |
| Paid new-customer CAC | paid spend ÷ closed-won | $6,000 | March cohort @180d | CRM + billing |
| Allowable CAC (first-year) | ACV × gross margin | $9,600 | - | own economics |
| Break-even CPL | ACV × lead-to-close (2%) | $240 | - | own economics |
| Payback | CAC ÷ (monthly gross profit $800) | 7.5 months | - | derived |
Ladder:
1. **Break-even** - CAC $6,000 sits under the $9,600 first-year contribution ceiling. CPL $120 sits under the $240 break-even CPL. Above water by arithmetic.
2. **Own history** - prior cohorts at the same 180-day maturity: $6,800 → $6,400 → $6,100 → $6,000. Improving.
3. **External** - published SaaS CAC-payback medians conflict: 18 months (Benchmarkit 2025, 583 participants, self-reported), 16 months (Aleph x Benchmarkit 2026, 198 of 342 reporting), 20 months (KeyBanc 2024, private SaaS panel). Presented as a conflict, not averaged. Quorline's 7.5 months sits inside all three panels' healthy half - context only.
**Verdict: healthy** - rungs 1 and 2 both clear, cohort-mature data, every figure labeled.
### Negative reading - same business, misread
> "March: we spent $48,000 and closed 3 deals. That's a $16,000 CAC. First Page Sage says B2B SaaS CAC is $239. This is a catastrophe - kill paid."
What went wrong, in order:
1. **Window shorter than the sales cycle.** The 3 deals closed _in_ March came from January's leads. March's spend bought a cohort that closes in June. In-period division mixed two cohorts and manufactured a $16,000 figure that describes neither.
2. **Variant mismatch on the benchmark.** First Page Sage's $239 (client analytics, Jan 2022-Aug 2025) is a *combined organic + paid* figure from a sample the agency itself discloses as 75% organic-weighted - compared here against a paid-only CAC. Its own paid-channel figure ($341) is still a different sample, different cost scope, different customer definition.
3. **Benchmark-only verdict, again.** Rung 1 (allowable CAC $9,600) and rung 2 (improving cohorts) were never consulted. The entire verdict rests on a mismatched rung 3.
4. **The evidence gate was skipped.** An immature window is an explicit withhold condition - the honest March-time output was "insufficient evidence until the cohort matures; interim read: cost per SQL $800 vs break-even CPL $240 × SQL rate," not a verdict at all.
Same numbers, opposite conclusion: "catastrophe, kill paid" vs "healthy" - the difference is entirely in the comparison, not the arithmetic.
SKILL.md›
---
name: cac-roas-benchmark
description: "Compute a business's CAC and ROAS family metrics from real spend and outcome data, then judge whether the spend is healthy against three references - the business's own break-even, its own trailing history, and provenance-labelled external benchmarks - returning healthy, watch, unhealthy, or insufficient evidence. Use whenever the user asks whether their CAC is too high, whether a ROAS is good, what ROAS to aim for, whether ads are profitable, or mentions MER, blended CAC, payback period, or a spend health check - even if they never say 'benchmark'. Covers B2B/SaaS (cost per SQL, cost per closed-won, cohort lag) and B2C/ecommerce. Do NOT use to set those thresholds as policy - use mbfinotti/advertising-skills@ad-spend-guardrails instead."
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.4.1"
---
# CAC & ROAS Spend Health Check
You are a marketing unit-economics analyst. Your deliverable is a spend health verdict: the business's CAC and ROAS family metrics, each labeled with its variant, window, and data source, judged against three references in strict priority order.
The calculation is trivial - the defensibility lives entirely in what each number is compared against. A verdict that rests only on an external published benchmark is the failure this skill exists to prevent.
## Boundaries - hand off, don't absorb
- Setting the target or guardrail that defines "healthy" as policy (max acceptable CAC, ROAS floor, kill-switch thresholds) is out of scope - that is `mbfinotti/advertising-skills@ad-spend-guardrails`. This skill judges current spend against arithmetic and history, not against a policy it invents.
- Recommending how to move budget between campaigns or channels is `mbfinotti/advertising-skills@ad-spend-allocation`. A verdict is not a reallocation plan.
- Tracking daily or weekly spend against a budget is `mbfinotti/advertising-skills@ad-budget-pacing`.
- Diagnosing _why_ an account underperforms is `mbfinotti/advertising-skills@ad-account-diagnostic`. Reconciling disagreeing reporting systems is `mbfinotti/advertising-skills@ad-attribution-gap` - when its reconciled numbers exist, use them as this skill's inputs instead of raw platform exports.
## Before starting - the intake
Ask these up front, batched. This is a tactical run on real numbers, not a strategy interview.
1. B2B, B2C/ecommerce, or blended? A blended business runs as two separate checks.
2. What evaluation window - and does it cover the conversion lag? (B2B: 4-6 weeks minimum, longer than the sales cycle for cohort reads.)
3. Which spend lines are in the CAC numerator: media only, media plus agency fees/tooling/creative production, or fully loaded with salaries?
4. What counts as a "new customer" in the denominator - and are renewals, repeat buyers, and reactivations excluded?
5. What is the revenue basis: platform-attributed, or from the order/billing system? Gross, or net of refunds and cancellations?
6. What is the contribution margin (or gross margin plus variable costs) per order or customer? Without it, break-even cannot be computed and the verdict may have to be withheld.
7. Are 4-8 prior periods of the same metric, on the same definition, available?
Then three scoping questions, because Step 1's metrics differ by orders of magnitude in effort and in how long their payoff lasts - that ordering cannot be picked for the user:
8. By when must the verdict land - a date, not "soon"?
9. A one-off read, or a standing check that will be repeated every period?
10. What is the effort ceiling: hours available, and whether finance or data can be pulled in at all?
Re-rank Step 1's menus against the answers, and say out loud which answer moved which metric:
- A hard near-term date keeps the run on the near-zero-effort variants (blended CAC, blended ROAS, MER) and deletes fully-loaded CAC, marginal CAC and any margin rebuild from this run - a narrower verdict, on time. Say which variants you deleted and why. A variant left ranked last reappears halfway through as a week of work nobody scheduled.
- A standing check promotes the slow ones: contribution margin, a consistent-definition history, and cohort grouping each cost a week once and near-zero every period after.
- Finance data that cannot produce a variant deletes it outright rather than demoting it: no salary-allocation rule means no fully-loaded CAC, and no contribution margin means no rung 1 either. Name both deletions at intake, and say then that the verdict will be withheld - never at the end of the run.
If you can read the user's exports (CSV, spreadsheet, warehouse extract), work from those. Otherwise ask for the totals per period. Record every answer in the report's definitions section - CAC is not a GAAP or IFRS term, no standards body defines what enters the numerator, so the definition agreed here is the only definition that exists.
## Step 1 - Establish the metric variant
"CAC" and "ROAS" each name a family of metrics, not one metric. Most bad verdicts trace to this step being skipped: two people quoting "our CAC is $240" routinely mean different numbers, and a figure compared to a benchmark measuring a different variant measures the definitions, not the business.
Both menus below are ordered by what each variant returns per unit of work to produce it - not by which is cheapest. Compute in that order and stop once the verdict is settled.
| CAC variant | Formula | Decision it serves | What it costs to produce |
| ---------------- | ------------------------------------------------------------- | ---------------------------------------- | ----------------------------------------------------------------------------------- |
| Blended CAC | total marketing spend ÷ all new customers (organic included) | whole-engine economics; feeds break-even | near-zero - both totals are already reported |
| New-customer CAC | any variant with renewals and reactivations excluded | subscription and repeat-purchase | an hour where billing already flags customer status; a week where it doesn't |
| Paid CAC | paid media spend ÷ paid-attributed new customers | channel efficiency | an hour to split the attribution - and worth no more than that attribution is worth |
| Fully-loaded CAC | S&M incl. salaries, tooling, agency, creative ÷ new customers | finance, boards, investors | a week: finance's ledger, a salary-allocation rule, and agreement on what counts |
| Marginal CAC | cost of the next increment of spend | scaling decisions, not health | a quarter, and deliberate test spend that cannot be recovered |
- efficiency: blended > new-customer > paid > fully-loaded > marginal
- value (for a health verdict): blended == new-customer > fully-loaded > paid > marginal
- effort: marginal > fully-loaded > paid > new-customer == blended
- compliance cost: fully-loaded > every other variant == none - once it reaches a board deck or an investor update it is a reported figure finance has to own, and redefining it later is a restatement, not an edit.
Marginal CAC's own decision rule (Demand Curve): a rising CAC is not automatically a problem to fix - check contribution margin first, and only test a new channel once the current channel's marginal CAC exceeds what the next channel could realistically deliver. Reacting to a CAC bump (their own example: a 20% rise) by immediately spreading spend into an untested channel is the more common failure than staying too long in a channel already working.
The ties are conditional, and the condition is what to check - not a way of leaving two variants undecided.
- **Value tie (blended == new-customer):** holds only where the business has no renewals, repeat purchases, or reactivations to miscount. The moment it has any, new-customer CAC sits strictly above blended, because it is the same read with a denominator error removed.
- **Effort tie (new-customer == blended):** holds only where billing already flags customer status, which makes both of them totals you already hold. Where it does not, new-customer costs a week and drops below paid on the effort line.
- **Compliance-cost tie (every variant == none, except fully-loaded):** every variant but fully-loaded ties at literally zero, because none of them ever leaves the marketing team's own report.
New-customer CAC is a correction applied to whichever variant you compute, not a rival to them - and it is not optional for a subscription or repeat-purchase business, where counting renewals as acquisitions inflates the denominator and flatters every verdict built on it.
**What this order starves:** fully-loaded CAC, and cohort CAC with it. Both are high on value, and both cost a week, so a ratio ranks them below the blended figure already sitting on the dashboard every single time.
- **Fully-loaded CAC** is the only variant finance and a board will recognise. Promote it whenever the verdict is going anywhere near a board deck, an investor update, or a fundraise.
- **Cohort CAC** is the only honest read of a lagging B2B business. Promote it whenever the sales cycle is longer than the window - Step 4's evidence gate withholds the verdict without it anyway.
Never let "the blended number is already there" decide a run where the audience is finance.
- Blended is always lower than paid when organic acquisition exists.
- Fully-loaded is always higher than media-only.
A worked example in circulation (Eightx): $84K spend, 2,000 new customers of which 1,200 from paid = **$42 blended and $70 paid CAC for the same month**. Silently switching variants makes CAC "improve" with nothing changing.
| ROAS family | Formula | What it costs to produce | Caveat |
| ------------------------------- | ------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ |
| Blended ROAS | total revenue ÷ total paid spend | near-zero - two totals the business already has | hides per-channel and per-SKU variance |
| MER | total revenue ÷ total _marketing_ spend | an hour the first time, to assemble agency, tooling and creative into the denominator | denominator is the only difference from blended ROAS |
| Contribution-margin ROAS / POAS | contribution margin (or gross profit) ÷ spend | a week the first time - COGS, shipping, payment fees, fulfillment, agreed with finance - then a standing job to keep those costs current | closest to actual profit per ad dollar |
| Platform-reported ROAS | platform-attributed value ÷ that platform's spend | near-zero - it is already on the dashboard | claimed revenue, not caused revenue; non-additive across platforms |
- efficiency: blended == MER > POAS > platform-reported
- value: POAS > blended == MER > platform-reported
- effort: POAS > MER > blended == platform-reported
- **Efficiency/value tie (blended == MER):** they are one metric with two denominators, and the tie holds only while agency, tooling and creative spend are small next to media. Once those are material, MER sits strictly above blended and the tie breaks.
- **Effort tie (blended == platform-reported):** both are already-reported totals, tied at exactly zero effort - precisely why the effort axis cannot choose between them and the value axis has to.
The axes disagree here, and the disagreement is the finding: platform-reported ROAS is the cheapest number in the building and the least worth acting on, while POAS is the most expensive and the only one that reflects actual profit. Spend the week on POAS anyway whenever the run is a standing check, or whenever contribution margin is unknown - the margin is rung 1's input, so without it the evidence gate withholds the verdict entirely (Step 4). That week buys the verdict itself, not a nicer number.
**Every ordering in this skill is a default, not a law.** It shifts with context and with who executes it, so re-rank it against what this business already owns before following it.
- A warehouse already joining ad spend to orders drops POAS to near-zero effort and puts it first.
- A finance team already publishing fully-loaded S&M each period does the same for fully-loaded CAC.
- An existing CRM cohort model does it for cohort CAC.
In the other direction, attribution nobody has a reason to trust does not demote paid CAC and platform-reported ROAS - it removes them.
Treat platform-reported ROAS as reported, never causal. Randomized-trial research (Gordon, Zettelmeyer, Bhargava & Chapsky 2019, Marketing Science, 15 large paid-social RCTs) found observational attribution overstating true lift by roughly 7-9x. Account audits show the two largest ad platforms jointly claiming 150-200% of real revenue (AdBeacon, vendor client audits - indicative, not audited).
As practitioner Ralph Burns puts it (Perpetual Traffic ep. 803, 2026): "The platforms themselves are going to over inflate and you can't really trust it."
If the variant of the user's numbers cannot be established, stop: the verdict is **insufficient evidence**, not a guess.
## Step 2 - Compute the metric set
Show each formula next to its result. Label every figure with variant, window, and source.
- `CAC (per variant) = included spend ÷ included new customers`
- `Blended ROAS = total revenue ÷ total paid spend`
- `MER = total revenue ÷ total marketing spend`
- `Contribution-margin rate = (revenue − COGS − variable costs: shipping, payment fees, fulfillment) ÷ revenue`
- `Break-even ROAS (= break-even MER) = 1 ÷ contribution-margin rate` - 60% margin ≈ 1.67x, 40% → 2.5x, 25% → exactly 4.0x. The same 4x ROAS is $1.40 profit per ad dollar at 60% margin and $0.00 at 25%.
- `Allowable CAC = contribution per sale` (B2C: AOV × contribution-margin rate; B2B: first-year ACV × gross margin, or the payback-window contribution)
- `Payback (months) = CAC ÷ (monthly gross profit per customer)` - keep the gross-margin term. About half of published versions drop it, which shortens the answer and flatters the verdict.
- `Discounted payback = CAC ÷ (monthly gross profit × annual retention)` - when retention is weak this blows past the raw figure. The gap is the early warning.
Three disciplines:
1. **Compute per plan, segment, and channel - never only blended.** The same $300 CAC is ~33 months of payback on a $9/month plan and 3 months at $99. One blended number describes neither.
2. **Mix-shift check before reading any blended trend.** Blended CAC or MER can move while every segment is flat, purely because the channel or segment mix moved.
3. **Missing data: renormalise the remaining channel weights, or withhold the blended figure entirely.**
- efficiency: renormalise > withhold
- Renormalising costs an hour and keeps a usable blend whenever the missing channel is small and its size is known.
- Withholding costs nothing and is always defensible but returns no number, so it is the fallback, not the first move.
- Never zero-fill a missing channel - a zero-filled channel silently corrupts every blend built on it.
## Step 3 - Build the comparison ladder
Three references, in strict priority order. Each rung is weaker than the one above it.
- efficiency: break-even > own history > external benchmark
One line covers every axis here, because cost runs exactly inverse to value: rung 1 is both the cheapest and the strongest, rung 3 both the most work to source properly and the weakest evidence. That alignment is unusual, and it is why resting a verdict on rung 3 alone is always the wrong trade - it costs the most and proves the least.
1. **The business's own break-even**, from its contribution margin (Step 2). An hour once the margin is known, and nothing after. Always computable from data the business owns, needs nothing external, and is non-negotiable: below it the spend loses money by arithmetic, whatever any benchmark says.
2. **The business's own trailing history** - 4-8 prior periods, same variant, same definition, same window. An hour when those periods already sit on one definition. A week to re-baseline when the definition moved. Direction and volatility matter more than the level: a vertical median cannot know this business's price point, margin, or sales motion, but its own last four quarters do.
3. **An external published benchmark, with provenance attached** - publisher, year, sample, and the metric variant it measured, every time (see [./references/benchmark-sources.md](./references/benchmark-sources.md)). A day of sourcing, and most searches end in a variant mismatch rather than a usable match. Use it to size the gap and set context, never to set the verdict alone. Match variants before comparing: a blended CAC judged against an agency's organic-weighted client figure, or a platform ROAS judged against a measured blended median, is a definitional mismatch, not a finding.
Published medians disagree with each other because panels and definitions differ (SaaS CAC payback medians of 16, 18, and 20 months coexist across 2024-2025 surveys).
- Present the conflict, never average it.
- Report ranges and quartiles, never a point estimate from a small sample.
## Step 4 - Assign the verdict
Four states, and the evidence gate comes first. Unknown is not the same as failing.
**Evidence gate - withhold the verdict (state: insufficient evidence) when any of these holds:**
- The CAC or ROAS variant in the user's data cannot be established (Step 1).
- Contribution margin is unknown, so rung 1 cannot be computed.
- The window is shorter than the business's conversion lag - the data has not matured.
- A channel's data is missing and the blend cannot be defensibly renormalised.
- Only rung 3 is available: no break-even and no usable history. A benchmark-only comparison is context, not a verdict.
**When the gate passes:**
- **Healthy** - at or above break-even with headroom on lag-matured data, and the trailing trend is flat or improving.
- **Watch** - above break-even, but the trend has deteriorated for two or more consecutive periods, or seasonality could plausibly erase the headroom, or the gap to a well-matched external benchmark is large and unexplained. Watch means "re-examine next period with a named question," not "act."
- **Unhealthy** - below the business's own break-even on lag-matured data, or a sustained deteriorating trend that has crossed break-even. This is arithmetic, not opinion.
- **Insufficient evidence** - the gate failed. Name exactly which input is missing and what would unlock the verdict. Withholding is the correct output, not a failure of the run.
An unhealthy verdict triggers a handoff, not an action: the pause/reallocate decision belongs to `mbfinotti/advertising-skills@ad-spend-guardrails` and `mbfinotti/advertising-skills@ad-spend-allocation`. One measurement-integrity exception worth flagging immediately: a server-side conversion API match rate below 60% is widely treated by practitioners as a measurement emergency - fix measurement before judging any number derived from it.
## Folklore - name it, don't delete it
Users will quote these targets. Deleting them from the report just moves the argument to the meeting. Name each one, give its origin, and put the business's own break-even next to it.
Full origins and the longer list: [./references/benchmark-sources.md](./references/benchmark-sources.md).
This table is deliberately unranked. Ordering it would imply one folk target is a better guide than another, when none of them is evidence at all - ranking here would be false precision. Row order tracks how often each one gets quoted, and carries no other meaning.
| Quoted rule | Origin | What it is worth |
| ------------------------- | ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------- |
| 3:1 LTV:CAC | David Skok (Matrix Partners, ~2011-2013), on record at SaaStr: "I guessed at that number" | a self-admitted guess from observing mature SaaS; secondary check at best |
| 12-month CAC payback | David Skok, 2011 | rule of thumb tied to 2011 fundraising conditions; never validated |
| 4:1 ROAS | no traceable author; often misattributed to a 2016 Nielsen study that says no such thing | simply break-even at a 25% margin, retroactively declared a target |
| MER > 4, 5-8 at scale | Taylor Holiday (Common Thread Collective) | stated heuristic, never measured across a sample |
| CAC ≈ 25% of gross profit | Taylor Holiday (CTC, 2022), the "fuel profit" framing | same status |
Blake Bartlett (OpenView) on the 3:1 rule: "Why is a 3x LTV:CAC ratio the appropriate benchmark? No one knows. It just is." Prefer payback over LTV:CAC as the affordability lens: LTV:CAC inherits whichever CAC variant fed it, hides per-plan variance under blended ARPU, and never asks when the cash comes back.
## B2B vs B2C
Identical for both - say so instead of hunting for a difference: the break-even arithmetic (Step 2), the comparison ladder, the evidence gate, the four verdict states, and the per-segment discipline.
| | B2C / ecommerce | B2B / SaaS |
| ------------------- | --------------------------------------- | ------------------------------------------------------------- |
| Headline metrics | MER, blended ROAS, contribution margin | cost per SQL, cost per closed-won, payback |
| Denominator inputs | AOV, contribution margin per order | ACV, lead-to-close rate |
| Window | weekly to monthly; data settles in days | 4-6 weeks minimum; cohorts immature until the cycle completes |
| Dominant distortion | platform-reported ROAS inflation | optimising to cheap form fills that never become revenue |
**B2B.** A CAC computed on a window shorter than the sales cycle counts this period's spend against last period's customers - that mismatch, not the spend, is usually what makes the number look bad. B2B cycles commonly run around three months, longer at enterprise, with multi-person buying committees.
- Group leads by the month generated and measure revenue at 90/180/365 days (cohort CAC/ROAS).
- Distinguish in-period from lagged CAC explicitly.
- Where revenue has not landed, fall back to pipeline-dollar metrics - a circulating convention is $0.10-0.20 cost per pipeline dollar at 180 days, a convention, not a measured benchmark.
Break-even shortcuts:
- `break-even CPL = ACV × lead-to-close rate`
- `break-even CPC = break-even CPL × landing-page conversion rate`
Order the funnel metrics by what each one is worth, not by how fast it arrives:
- efficiency: cost per closed-won > cost per SQL > cost per pipeline dollar > CPL
- effort: cost per closed-won (a quarter of waiting for the cohort, plus a CRM-to-billing join) > cost per SQL == cost per pipeline dollar (an hour each, once the CRM stages are clean) > CPL (near-zero)
Cost per SQL and cost per pipeline dollar tie because they are two readings off one artifact: clean the CRM stages once and the second costs nothing beyond the first, so nothing can separate them on effort.
Cost per closed-won is the metric that tells the truth. CPL is the cheapest and the most misleading, because it favours cheaper, worse-converting sources. While the cohort is immature, cost per SQL and cost per pipeline dollar are interim reads, never verdicts.
Cost per closed-won is also what this order starves: first on value, first on effort, so a ratio never reaches it. Promote it on purpose whenever the run is a standing check rather than a one-off. Treat the quarter of waiting as the price of the only honest number, not as a reason to settle for CPL.
**B2C/ecommerce.** Layer contribution margin, from narrowest to widest:
- CM1 = revenue − COGS
- CM2 adds delivery
- CM3 adds marketing - the closest view of cash profit per order
Anchor revenue on the order system net of refunds, never on platform-attributed value. Check first-order contribution against CAC: a CAC above first-order contribution makes the verdict depend on the repeat-purchase rate - say so explicitly. Seasonality alone can invert a verdict (a toys-vertical ROAS of 1.46 in July vs 2.90 in December in the same 2025 dataset - Billo, 80,000+ paid-social video ads), so compare periods to the same season, not the prior month.
## Cadence and reconciliation
- Weekly: pacing-level signals only (spend, CPL, ROAS direction). Not a verdict.
- Monthly: the verdict, reconciled with finance's booked actuals. Marketing's operational read and finance's closed number will differ by design - label each, show both side by side, and let finance's reconciled number govern any board-facing figure.
- Quarterly: refresh external benchmarks. Publications update yearly at best. Re-pulling more often adds noise, not information.
- Any methodology change (a spend line added, a customer definition changed): re-baseline before comparing to prior periods, and say so in the report.
## The report
Deliver these sections (full template and worked examples - one B2C and one B2B, each with a positive and a negative reading of the same numbers: [./references/worked-examples.md](./references/worked-examples.md)):
1. **Headline** - the verdict per segment/channel judged, one line each, with the rung(s) it rests on.
2. **Definitions record** - the seven definitional intake answers: variant, spend lines, customer basis, revenue basis, window, margin, history available. Add one line naming each variant the scoping answers (deadline, one-off vs standing, effort ceiling) deleted from this run, and one naming any high-value variant the efficiency order starved that the user should promote next period.
3. **Metric table** - one row per figure: metric | variant | value | window | data source.
4. **Comparison ladder** - rung 1, 2, 3 values and the gap to each, with a provenance line under every external figure.
5. **Verdict and evidence gate** - the state, the conditions checked, what was withheld and why.
6. **Folklore appendix** - only if the user raised a folk target: its origin next to the business's own break-even.
7. **Handoffs** - named next skill for anything out of scope that the check surfaced.
## Failure modes
- Setting the verdict from rung 3 alone - the core failure. A business can be "below the industry median" and comfortably profitable, or "above median" and losing money by arithmetic.
- Comparing across variants: blended vs paid, media-only vs fully-loaded, platform-reported vs blended. The gap measures definitions, not performance.
- Zero-filling a missing channel instead of renormalising or withholding.
- Judging B2B on a window shorter than the sales cycle, then blaming the spend.
- Dropping the gross-margin term from payback.
- Quoting any external number without publisher + year + sample + measured variant.
- Averaging conflicting published medians instead of presenting the conflict.
- Point estimates from small-sample surveys (several published SaaS metrics rest on n = 21-43) - report medians and quartiles as ranges.
- Reading a blended trend without the mix-shift check.
- Treating a verdict as a budget decision - that is a handoff, not a conclusion.
## Objective and pass condition
The run passes only when both conditions hold, audited line by line against the finished report:
- **Every figure is labeled.** 100% of reported figures carry their variant label, their window, and their data source, and every external benchmark cited carries publisher + year + sample + the variant it measured. A single bare number fails the run.
- **The verdict rests on rung 1 or rung 2.** If neither is computable, the only passing verdict is "insufficient evidence" with the missing inputs named.
Iterate until both conditions hold.
## References
- `mbfinotti/advertising-skills@ad-spend-guardrails` - turning this verdict into targets, floors, and kill-switch policy.
- `mbfinotti/advertising-skills@ad-spend-allocation` - acting on the verdict by moving budget.
- `mbfinotti/advertising-skills@ad-budget-pacing` - tracking spend against budget between checks.
- `mbfinotti/advertising-skills@ad-account-diagnostic` - finding out why, once this skill says something is unhealthy.
- `mbfinotti/advertising-skills@ad-attribution-gap` - reconciling the reporting systems whose outputs feed this check.
- [./references/benchmark-sources.md](./references/benchmark-sources.md) - published benchmark figures with full provenance, and the folklore origins list.
- [./references/worked-examples.md](./references/worked-examples.md) - report template plus end-to-end B2C and B2B examples, each with the same numbers read correctly and misread against a bare benchmark.