Back to Skills
mbfinotti/advertising-skillsCheck passed

SKILL DETAIL

ad-spend-guardrails

mbfinotti/advertising-skills/ad-spend-guardrails

Set the organisation's top-level paid-media spend guardrails as written policy - maximum allowable CAC, minimum ROAS/MER floor, kill-switch thresholds, counter-metrics, who owns them and who may override them - each derived from contribution margin, payback, and cash runway rather than inherited from a dashboard. Use whenever the user asks what CAC they can afford, mentions a ROAS floor, spend guardrails, a kill switch, when to stop spending, or who approves a budget increase - even if they never say 'guardrails'. Covers B2B and B2C. Do NOT use to judge whether current CAC/ROAS is actually good (mbfinotti/advertising-skills@cac-roas-benchmark) or to monitor daily pacing (mbfinotti/advertising-skills@ad-budget-pacing).

Installs · 180View source

Installation

npx skills add https://github.com/mbfinotti/advertising-skills --skill ad-spend-guardrails

Skill files

SKILL.md

Last synced · Sep 24, 2026

evals/evals.json›
{
  "skill_name": "ad-spend-guardrails",
  "evals": [
    {
      "id": 1,
      "prompt": "I run Maribel Home Goods, a DTC home fragrance brand. AOV is $68 and our contribution margin after COGS and shipping is 45%. We spend about $60K/month on Meta and Google. Our CFO asked me for an official ROAS target to put in the annual plan - everyone online says 4x is the standard to aim for. Can you write up the target for us?",
      "expected_output": "A guardrail derivation from the company's own 45% margin, with 4x named as folklore, three separate threshold layers with distinct consequences, a marginal scaling ceiling, a cash cap question, and a counter-metric pairing - not a single benchmark-quoted ROAS number.",
      "files": [],
      "expectations": [
        "Derives break-even ROAS from the stated margin as 1 divided by 0.45, roughly 2.2x, rather than adopting 4x.",
        "Names 4x ROAS as folklore with no traceable author, not a measured standard.",
        "Explains that 4x is simply break-even at a 25% contribution margin retroactively declared a target.",
        "Keeps the 4x rule visible next to the business's own break-even instead of endorsing it or silently deleting it.",
        "Proposes three distinct numbers: a break-even floor, a target floor, and a hard floor / kill-switch level.",
        "Assigns each layer its own consequence: stop always at break-even, investigate at target, automatic halt plus escalation at the hard floor.",
        "Warns that publishing a single ROAS number gets read as all three layers at once, producing both over-reaction to one bad day and tolerated slow bleeds.",
        "States break-even allowable CAC as contribution per order, about $30.60 ($68 x 45%).",
        "Asks for cash runway and burn, or states a cash cap as an absolute monthly spend number that binds regardless of efficiency.",
        "Recommends the scaling ceiling be set on marginal efficiency, with the blended figure reserved for the monthly report.",
        "Asks which ROAS variant is meant, or warns against setting the guardrail on platform-reported ROAS in favour of business-level outcomes.",
        "Asks clarifying interview questions, including what guardrails already exist, before finalizing numbers rather than delivering a finished target immediately.",
        "Pairs the proposed ROAS/MER floor with a counter-metric such as new-customer share to expose retargeting and branded-search gaming."
      ]
    },
    {
      "id": 2,
      "prompt": "We're Kestrel Data, B2B SaaS, $24K ACV at 62% gross margin, average sales cycle around 75 days. Cost per SQL on LinkedIn doubled over the past two weeks and our CMO wants an automatic rule in place: any channel that runs above 3x our target CPA for three days straight gets shut off. Can you write that rule up properly so we can put it in our ops doc?",
      "expected_output": "A kill-switch design that rejects the three-day window as inside the conversion lag, flags 3x-CPA as untraceable folklore, and delivers streak plus rate conditions, an evidence gate, a written restart condition, three outcomes, pipeline-stage evidence, and a derived allowable CAC.",
      "files": [],
      "expectations": [
        "Refuses a three-day kill window as shorter than the conversion lag, setting a minimum evidence window of 4-6 weeks or longer given the 75-day cycle.",
        "Explains that cost per SQL or CAC read on a short window counts this period's spend against an earlier period's customers, so the two-week spike may be lag rather than economics.",
        "Flags the 3x-CPA kill rule as untraceable folklore and declines to present it as a benchmark.",
        "Anchors any spend-based gate to attributed practitioner guidance (such as no activation by $500-1,000 of spend, or 3-4x CPA of spend before judging), labeled as guidance rather than measurement.",
        "Designs two independent halt conditions: a consecutive-period streak condition and a budget-burn rate condition, either firing alone.",
        "Includes both a per-period cap and a rolling cumulative cap so a slow bleed inside daily tolerance still trips.",
        "Adds an evidence gate requiring both a minimum data window and a minimum spend or conversion volume before any kill is allowed.",
        "Writes the restart condition alongside the halt condition: what must be true, over what window, approved by whom.",
        "Routes breaches into three outcomes - allow, review, halt - instead of forcing every breach into an immediate stop.",
        "Keys the kill evidence on pipeline-stage signals such as SQL-to-closed-won, never raw lead volume.",
        "Derives allowable CAC from contribution per deal, about $14,880 ($24,000 x 62%), instead of accepting the 3x-of-target framing as the ceiling.",
        "Names a kill-decision owner and escalation path outside the team spending the money.",
        "Specifies fail-closed handling: missing, stale, or broken data holds spend flat and blocks scaling, and is never treated as the number being fine."
      ]
    },
    {
      "id": 3,
      "prompt": "Torvale Supplements here - 70% gross margin. We're at $110K/month on Meta with a blended MER of 2.0. Our agency points out that break-even at our margin is about 1.4, so they say we've got tons of headroom and want to push to $150K next month. Before I approve it I want a proper guardrail written down for scaling decisions. What should it be?",
      "expected_output": "A scaling guardrail set on marginal MER rather than the blended 2.0, explaining that a passing average can hide marginal spend below break-even, with three layers, a counter-metric, a cash cap, and ceiling ownership outside the agency.",
      "files": [],
      "expectations": [
        "Sets the scaling ceiling on marginal MER/ROAS, not on the blended 2.0 figure.",
        "Explains that a passing blended average can hide marginal spend below break-even, so the last dollars of the budget may already lose money while the average looks healthy.",
        "Computes break-even near 1.43x from the 70% margin (1 divided by 0.70) and treats it as the floor for the marginal number.",
        "Distinguishes the two roles: the marginal number authorizes the next dollar of spend, the blended number belongs in the monthly report.",
        "Declines to approve the $150K step on the blended figure alone and requires a marginal-efficiency read (such as MER on the last tranche of spend) as the scaling gate.",
        "Structures the policy as three layers with three separate consequences rather than a single floor.",
        "Pairs the MER guardrail with a counter-metric such as new-customer share of orders.",
        "Caps the guardrail set at two or three metrics.",
        "Asks for runway and burn, or states an absolute monthly cash cap next to the efficiency floors.",
        "Places ceiling ownership outside the agency spending the money, on separation-of-duties grounds.",
        "Includes a re-baselining cadence (quarterly) plus named event triggers such as margin, pricing, or measurement changes.",
        "Shows the arithmetic behind every threshold so a finance reader can audit it."
      ]
    },
    {
      "id": 4,
      "prompt": "Plumeria Skincare. Honestly we don't know our contribution margin precisely - somewhere between 50 and 60% probably? Also Meta reports about 40% more revenue than Shopify does, and we never figured out why. We just need a sensible CAC ceiling and a ROAS floor to start operating with - nothing fancy, just give us starter numbers.",
      "expected_output": "A refusal to hand over starter numbers: margin is named as unconfirmed, the 40% measurement gap gates the exercise, no guardrail is set on platform-reported ROAS, and any interim figure is derived, labeled provisional, and sourced from business-level outcomes.",
      "files": [],
      "expectations": [
        "Declines to fill the unknown margin with a plausible number, naming it estimated or missing and giving any policy a provisional status.",
        "Gates on measurement: the 40% platform-versus-store revenue gap must be resolved before guardrails are set on that feed.",
        "Refuses to set a kill-switch or floor on platform-reported ROAS while the conversion feed is untrusted.",
        "Explains why: a guardrail on an untrusted metric fires on data-quality noise and manufactures confident wrong decisions.",
        "Asks how confident the business is in its margin (a guess, finance-reviewed, or per-SKU) or otherwise interviews to pin the input down.",
        "Does not substitute a published benchmark (4x ROAS, 3:1 LTV:CAC, 12-month payback) as the starter number.",
        "Any interim number offered shows its derivation and is labeled provisional or estimated, never presented as measured.",
        "Recommends business-level accepted outcomes (store orders or CRM) as the guardrail data source rather than platform claimed revenue.",
        "Notes that platform numbers report claimed rather than caused revenue and push spend toward retargeting and brand search.",
        "Specifies fail-closed behavior while the feed is broken: hold spend flat rather than scale.",
        "Labels inputs measured, estimated, or missing in any drafted policy skeleton.",
        "Asks the user questions before proposing thresholds instead of assuming the missing facts."
      ]
    },
    {
      "id": 5,
      "prompt": "At Bramwell Software our Head of Growth wrote our ad spend rules last year: she set the CAC ceiling at $310, she approves any overage herself, and we haven't breached it once in three quarters, so the rules clearly work. Our CFO wants the whole thing formalized into an official policy before the next board meeting. Can you turn what we have into the formal version? Anything worth adjusting while we're at it?",
      "expected_output": "A governance rewrite that moves ceiling ownership away from the spender, treats three zero-breach quarters as a review trigger rather than proof of success, audits the $310 derivation, adds an escalation ladder with named approvers, a logged exception process, and policy-level KPIs.",
      "files": [],
      "expectations": [
        "Flags that the Head of Growth both spends the budget and approves her own overages, and moves ceiling ownership or approval to a separate role.",
        "States the principle: whoever spends the money cannot raise their own ceiling, otherwise it is a preference rather than a control.",
        "Treats three zero-breach quarters as a review trigger for a sandbagged or mis-set ceiling, not as evidence the rules work.",
        "Checks headroom: persistent distance between actual efficiency and the $310 ceiling would mean the ceiling is throttling profitable growth or was set too loose.",
        "Audits how the $310 ceiling was derived and requires it traced to contribution margin, payback, or cash rather than accepting it as-is.",
        "Requires an escalating approval ladder by severity, each tier naming a real, notified person.",
        "Rewrites exceptions as explicit, time-boxed, justified, and logged, including denied ones, and calls out standing self-approval as a new ceiling nobody agreed to.",
        "Rejects vague override rationales such as 'timing' or 'one-time' as insufficient justification.",
        "Splits the single $310 ceiling into break-even, target, and hard floors, each with its own consequence.",
        "Adds a re-baselining cadence (quarterly) plus named event triggers such as pricing, margin, or measurement changes.",
        "Proposes judging the policy itself with KPIs read off the breach log, such as breach response rate, override concentration, headroom, or zero-breach quarters.",
        "Ends with an explicit approval gate: nothing finalized without sign-off, recording who approved the policy and when."
      ]
    },
    {
      "id": 6,
      "prompt": "Ostrena, we sell DTC coffee gear, roughly $85K/month across Meta and Google. Two years ago we put a strict 2.8x ROAS bar on absolutely everything, and our media buyer's complaint is that we haven't successfully launched a new channel or audience since - every new thing dies within a week of launch. We do have an in-house analyst who already tracks marginal returns weekly. Time to rework our spend rules. What should they look like?",
      "expected_output": "An explicit brainstorming pass over the three guardrail architectures with the ranking said out loud, the portfolio model promoted because the efficiency bar strangled testing and the analyst collapses its effort cost, a ring-fenced exploration budget with spend-based rules, and a re-derived break-even replacing the unexamined 2.8x.",
      "files": [],
      "expectations": [
        "Presents all three guardrail architectures - single break-even floor, tiered ladder, and portfolio with ring-fenced exploration - before drafting numbers.",
        "States the ranking out loud across axes (efficiency, value in false stops avoided and learning kept, effort) rather than presenting one option.",
        "Promotes the portfolio model above its default rank because the efficiency bar has already strangled testing.",
        "Uses the in-house analyst who already tracks marginal returns to argue the portfolio model's effort cost collapses for this account.",
        "Ring-fences an exploration budget as a declared share of spend with its own separate kill rules, exempt from the main floor.",
        "Gives exploration tests spend-based judgment rules (hypothesis, decision date, spend cap) instead of the production 2.8x bar.",
        "Names the failure mode explicitly: testing held to the production efficiency bar quietly stops testing.",
        "Asks where the 2.8x figure came from and re-derives break-even from the business's own margin, asking for the margin if unknown, rather than keeping 2.8x unexamined.",
        "Flags that one flat threshold across all channels and audiences starves prospecting, which always looks worse than retargeting.",
        "Argues the strongest case against its own recommended architecture before presenting it.",
        "Waits for the user's choice of architecture rather than unilaterally imposing one.",
        "Asks whether the mandate is a one-off fix or a compounding asset, and what effort ceiling exists for a standing review, before or while ranking.",
        "Pairs whatever efficiency guardrail replaces the 2.8x bar with a counter-metric that exposes gaming."
      ]
    },
    {
      "id": 7,
      "prompt": "Two-founder bootstrapped SaaS called Lanternfish HQ. One plan at $49/month, 80% gross margin, about 9 new customers a month from ads on roughly $400/month spend, 15 months of runway. An investor friend told me our CAC payback needs to be under 5 months because that's what his portfolio companies target. What CAC ceiling should we set?",
      "expected_output": "A refusal to set a hard CAC ceiling at this volume, replaced by an activation threshold, with payback derived from monthly gross profit, the investor's rule rejected via cost-of-capital reasoning, a cash cap, and a single-floor architecture given no second approver exists.",
      "files": [],
      "expectations": [
        "Flags roughly 9 conversions a month as below meaningful volume and declines to set a hard CAC ceiling now.",
        "Sets an activation threshold instead: a spend level or conversion count at which the CAC policy switches on.",
        "Computes monthly gross profit per customer as about $39.20 ($49 x 80%) and frames payback as CAC divided by monthly gross profit.",
        "Asks for or accounts for retention to churn-adjust payback (CAC divided by monthly gross profit times annual retention).",
        "Declines to inherit the investor's under-5-month rule as this business's target.",
        "Explains payback tolerance through cost of capital rather than stage labels or someone else's portfolio.",
        "Notes that a bootstrapped business facing near-zero capital cost can sustain a longer payback, up to around 24 months where the math supports it, while venture-funded companies face roughly 35-50% implied equity cost.",
        "Notes that payback under about 3 months usually signals underinvestment rather than health.",
        "States a cash cap as an absolute monthly number derived from the 15-month runway, sitting next to any efficiency guidance.",
        "Shows the arithmetic behind every number proposed so it can be audited.",
        "Labels inputs measured, estimated, or missing rather than guessing the gaps.",
        "Does not adopt 3:1 LTV:CAC or a 12-month payback rule as the answer; any such rule mentioned is labeled a heuristic with its origin.",
        "Recommends the single hard-floor architecture, or names the escalation ladder as deleted, because a two-founder account has no second approver or review chair."
      ]
    }
  ],
  "trigger_queries": [
    { "query": "What's the maximum CAC we can afford before ads lose us money?", "should_trigger": true },
    { "query": "Help me set a ROAS floor for our paid campaigns", "should_trigger": true },
    { "query": "We need a kill switch rule for when to stop ad spend", "should_trigger": true },
    { "query": "Who should be allowed to approve going over our CAC target?", "should_trigger": true },
    { "query": "My CFO wants a written policy for ad spend limits", "should_trigger": true },
    { "query": "At what point should we just turn the ads off?", "should_trigger": true },
    { "query": "Set spend guardrails for our Meta and Google accounts", "should_trigger": true },
    { "query": "How do I figure out what CAC target to give my agency?", "should_trigger": true },
    { "query": "Is 3:1 LTV to CAC the right target for us to adopt?", "should_trigger": true },
    { "query": "Our agency wants to double the budget - what rule decides whether they're allowed?", "should_trigger": true },
    { "query": "Write a policy for when we pause a campaign and who signs off", "should_trigger": true },
    { "query": "How much can we pay for a customer at 55% margin?", "should_trigger": true },
    { "query": "What payback period should we tolerate on ad spend as a bootstrapped company?", "should_trigger": true },
    { "query": "Define the threshold where marketing has to stop spending without asking anyone", "should_trigger": true },
    { "query": "We keep arguing about whether our CPA target should be 40 or 60 dollars - help us set it properly", "should_trigger": true },
    { "query": "I want a break-even ROAS calculation for our store", "should_trigger": true },
    { "query": "What should our MER floor be?", "should_trigger": true },
    { "query": "Boss asked me to write the rules for when we cut ad spend", "should_trigger": true },
    { "query": "Put together spend limits the finance team can sign off on", "should_trigger": true },
    { "query": "When performance tanks, who decides to pull the plug? We need a process", "should_trigger": true },
    { "query": "Need a rule for when a campaign gets killed versus just reviewed", "should_trigger": true },
    { "query": "What counter-metrics should sit next to our ROAS target so the team can't game it?", "should_trigger": true },
    { "query": "Our board wants a max CAC number - how do I derive one?", "should_trigger": true },
    { "query": "Set an exploration budget that's exempt from our efficiency targets", "should_trigger": true },
    { "query": "How far above break-even should our target ROAS sit?", "should_trigger": true },
    { "query": "At 30% margin, what ROAS do we actually need?", "should_trigger": true },
    { "query": "Draft the escalation path for ad spend overruns", "should_trigger": true },
    { "query": "We're burning cash - what monthly ad budget cap makes sense given 10 months of runway?", "should_trigger": true },
    { "query": "My cofounder and I disagree on when to stop the ads, settle it with a rule", "should_trigger": true },
    { "query": "Help me decide what CAC is too high for a $29/month product", "should_trigger": true },
    { "query": "We paused everything after one bad week and it wrecked the account - how do we avoid that next time?", "should_trigger": true },
    { "query": "Formalize our ad spend rules so the media buyer isn't grading his own homework", "should_trigger": true },
    { "query": "What's an acceptable customer acquisition cost for us?", "should_trigger": true },
    { "query": "The CEO wants one number: spend stops when ROAS hits X. Give me X", "should_trigger": true },
    { "query": "How do I set a CPA ceiling that accounts for our 90-day sales cycle?", "should_trigger": true },
    { "query": "Our ads are profitable on average but I'm scared we're overspending at the margin - where should the ceiling be?", "should_trigger": true },
    { "query": "When is it OK to restart a channel we shut down for bad performance?", "should_trigger": true },
    { "query": "Should the testing budget follow the same ROAS bar as everything else?", "should_trigger": true },
    { "query": "Give me a spend policy template our CFO and media buyer can both live with", "should_trigger": true },
    { "query": "12-month payback keeps getting quoted at me - is that the right bar for us?", "should_trigger": true },
    { "query": "We're venture backed, Series A - how aggressive can our CAC payback target be?", "should_trigger": true },
    { "query": "Define what counts as an emergency stop for paid media and who owns it", "should_trigger": true },
    { "query": "I don't trust the 4x ROAS rule - what should we use instead?", "should_trigger": true },
    { "query": "Is a 2.5x ROAS good for ecommerce?", "should_trigger": false },
    { "query": "Our CAC jumped from $80 to $140 last month - why?", "should_trigger": false },
    { "query": "Are we on pace to spend our $50K budget this month?", "should_trigger": false },
    { "query": "Should I switch my Google campaign to target ROAS bidding?", "should_trigger": false },
    { "query": "What target CPA should I enter in the Google Ads bid strategy settings?", "should_trigger": false },
    { "query": "Our campaign is proven at $200/day - how fast can we scale it to $1000/day?", "should_trigger": false },
    { "query": "How should I split $30K between Meta prospecting and retargeting?", "should_trigger": false },
    { "query": "Meta says 210 conversions but Shopify shows 150 - which is right?", "should_trigger": false },
    { "query": "Check whether our conversion pixel fires correctly before launch", "should_trigger": false },
    { "query": "Which platform should we advertise on with a $10K budget?", "should_trigger": false },
    { "query": "Our ROAS dropped this week - is the creative fatigued?", "should_trigger": false },
    { "query": "Find the search terms wasting our ad spend", "should_trigger": false },
    { "query": "Ads get clicks but nobody converts on the landing page", "should_trigger": false },
    { "query": "We have too many campaigns - which should we merge?", "should_trigger": false },
    { "query": "How do I become a media buyer?", "should_trigger": false },
    { "query": "Interview questions for hiring a performance marketer", "should_trigger": false },
    { "query": "What's our blended CAC for Q3? Here's the spend and order data", "should_trigger": false },
    { "query": "Is my ROAS below industry benchmarks for beauty brands?", "should_trigger": false },
    { "query": "Alert me when daily spend exceeds the daily budget", "should_trigger": false },
    { "query": "Why did spend suddenly double yesterday on Meta?", "should_trigger": false },
    { "query": "Set frequency caps for our retargeting sequence", "should_trigger": false },
    { "query": "How big does our lookalike seed list need to be?", "should_trigger": false },
    { "query": "Write guardrail metrics for our A/B test so we don't ship a bad variant", "should_trigger": false },
    { "query": "What guardrails should we put on our LLM agent's tool use?", "should_trigger": false },
    { "query": "Set up AWS budget alerts so our cloud spend doesn't blow up", "should_trigger": false },
    { "query": "I need a kill switch for our feature flags if error rates spike", "should_trigger": false },
    { "query": "Should we pause the campaign during the learning phase reset?", "should_trigger": false },
    { "query": "What's a fair monthly retainer to pay our PPC agency?", "should_trigger": false },
    { "query": "Our MER was 3.1 last month - healthy or not?", "should_trigger": false },
    { "query": "Diminishing returns: which campaign gets the extra $5K?", "should_trigger": false },
    { "query": "Track our burn rate against the quarterly media budget", "should_trigger": false },
    { "query": "When should I move from manual CPC to automated bidding?", "should_trigger": false },
    { "query": "My cost per lead doubled after the iOS update - diagnose it", "should_trigger": false },
    { "query": "Build a stop-loss rule for my day trading account", "should_trigger": false },
    { "query": "What LTV:CAC ratio do VCs want to see in a pitch deck?", "should_trigger": false },
    { "query": "Estimate CAC payback for the board deck from these cohort numbers", "should_trigger": false },
    { "query": "Our exec wants a weekly spend dashboard - what should be on it?", "should_trigger": false },
    { "query": "How often should we review our negative keyword lists?", "should_trigger": false },
    { "query": "Is $45 CPA too high for B2C subscription apps?", "should_trigger": false },
    { "query": "Which campaigns are stuck in learning limited and should be merged?", "should_trigger": false },
    { "query": "Compare LinkedIn versus Google Ads for B2B lead gen", "should_trigger": false },
    { "query": "Draft brand safety guardrails for where our ads can appear", "should_trigger": false },
    { "query": "We're raising our Meta budget 20% tomorrow - will it reset learning?", "should_trigger": false }
  ]
}
references/worked-guardrail-policies.md›
# Worked Guardrail Policies

Three examples: a filled B2B SaaS policy, a filled B2C e-commerce policy, and a negative example annotated line by line. The numbers are illustrative - they exist to show the shape of a defensible derivation, not to be copied as targets.

## Table of Contents

- [Example 1 - B2B SaaS](#example-1-b2b-saas)
- [Example 2 - B2C e-commerce](#example-2-b2c-e-commerce)
- [Example 3 - negative example, annotated](#example-3-negative-example-annotated)

## Example 1 - B2B SaaS

Context: $18K ACV, 68% gross margin, median sales cycle 78 days, $2.1M cash, $180K monthly burn, venture-funded and scaling. Conversion feed reliable (server-side CRM import of SQL and closed-won), no incrementality testing yet.

```
SPEND GUARDRAIL POLICY - Northwind Analytics, effective 2026-09-01, review 2026-12-01

Inputs
  Contribution per closed-won deal  $12,240 (ACV $18,000 x 68% GM)      measured
  Lead-to-close rate                7.4% SQL -> closed-won, 4 quarters    measured
  Payback target                    14 months (board-agreed)             measured
  Runway                            11.7 months at current burn          measured
  Cash cap on paid media            $95,000 / month                      estimated

Derivation
  Allowable CAC (break-even)  = $12,240 contribution per deal
  Allowable CAC (target)      = monthly gross profit $1,020 x 14 months = $14,280
                                -> the payback target is LOOSER than break-even,
                                   so break-even binds. Target CAC set at $8,500,
                                   the level the FY plan needs to fund headcount.
  Break-even cost per SQL     = $12,240 x 7.4% = $906
  Target cost per SQL         = $8,500 x 7.4%  = $629

Layers
  Break-even floor   cost per SQL $906        -> stop, always
  Target floor       cost per SQL $629        -> investigate within one weekly review
  Hard floor         cost per SQL $820        -> automatic halt of the breaching channel
                                                 + escalation to VP Marketing

Guardrail set
  1. Cost per SQL (CRM-sourced, 6-week trailing window)
     counter-metric: SQL-to-closed-won rate - catches cheap SQLs that never close
  2. Blended CAC on new logos only (excludes expansion), monthly
     counter-metric: new-logo share of pipeline
  3. Paid media spend as share of runway, monthly
     counter-metric: none needed - it is itself the cash constraint

Kill rules
  Streak    cost per SQL above the hard floor for 3 consecutive weekly reads
  Rate      35% of a channel's quarterly budget consumed with cost per SQL above
            the hard floor
  Evidence  no kill before 6 weeks of data AND 30 SQLs on the channel; a channel
            below either bar is held flat, never killed
  Restart   restart at 50% of prior budget after a documented fix (creative, offer,
            targeting or tracking), approved by the Growth Lead, reviewed at 4 weeks

Governance
  Owner            CFO owns break-even inputs; VP Marketing owns target and hard floors
  Escalation       target breach -> Growth Lead, weekly review
                   hard breach   -> VP Marketing, 48h decision
                   break-even breach -> CFO, immediate halt, no discretion
  Override         VP Marketing may authorize up to 4 weeks above target floor with a
                   written reason logged in the marketing operating doc; anything above
                   break-even floor requires CFO co-sign; overrides are never standing
  Re-baseline      quarterly, plus on: pricing change, gross-margin change > 3pts,
                   sales-cycle change > 2 weeks, CRM attribution change

Exemptions
  Exploration budget 12% of monthly paid spend, exempt from the cost-per-SQL floors.
  Own rules: each test carries a hypothesis, a decision date, and a spend cap of
  3x target cost per SQL before a keep/cut call.

Assumptions
  7.4% lead-to-close holds at higher volume. If it degrades below 6%, every floor above
  is wrong and the policy must be re-derived, not adjusted.
```

## Example 2 - B2C e-commerce

Context: $74 AOV, 54% contribution margin after COGS and delivery, repeat purchase within 90 days, profitable and self-funded, no MMM, platform pixel plus server-side events.

```
SPEND GUARDRAIL POLICY - Fernbrook Goods, effective 2026-09-01, review 2026-12-01

Inputs
  AOV                        $74                                      measured
  Contribution margin (CM3)  54% after COGS and delivery, before ads  measured
  Contribution per order     $39.96                                   measured
  Cash cap on paid media     $140,000 / month                         measured
  Repeat rate, 90 days       31%                                      measured

Derivation
  Break-even ROAS       = 1 / 0.54 = 1.85x
  Break-even CAC        = $39.96 first-order contribution
  Target ROAS           = 2.40x, the level that funds the 18% net-profit plan
  Marginal vs blended   at 54% margin the marginal floor is 1.85x and the blended
                        floor is set at 2.40x. Scaling decisions read the marginal
                        number; the monthly report reads the blended one.

Layers
  Break-even floor   marginal MER 1.85x        -> stop, always
  Target floor       blended MER 2.40x         -> investigate at the weekly review
  Hard floor         blended MER 2.00x         -> automatic halt of new scale-ups,
                                                  spend held at prior week's level

Guardrail set
  1. Blended MER, weekly (total revenue / total marketing spend)
     counter-metric: new-customer share of orders - a MER met by retargeting existing
                     buyers is not the same business
  2. Contribution margin after ads, daily
     counter-metric: discount depth - margin held up by promo dependence is borrowed
  3. Marginal MER on the last 25% of spend, at every scale-up decision
     counter-metric: none - it is the scaling gate itself

Kill rules
  Streak    blended MER below hard floor for 5 consecutive days
  Rate      $12,000 spent in any 72h window with blended MER below the hard floor
  Evidence  no kill on a campaign below $1,000 spend or below 3 days live; both bars
            must clear
  Restart   restart at 60% of prior daily budget once the underlying cause is named
            and fixed; two clean weeks required before returning to full budget

Governance
  Owner        Founder owns the break-even inputs and the hard floor;
               Head of Growth owns the target floor
  Escalation   target breach -> Head of Growth, weekly
               hard breach   -> Founder, same day
  Override     Founder only, maximum 14 days, reason logged; seasonal peaks are
               pre-approved in writing before the season, never mid-flight
  Re-baseline  quarterly, plus on: COGS change, shipping-cost change, price change,
               or any change to the conversion-tracking setup

Exemptions
  Creative testing 15% of monthly spend, judged on a spend-based rule (no purchase by
  $500-1,000 spend -> cut) rather than on the MER floors.

Assumptions
  54% CM3 assumes current shipping rates and a return rate at or below 8%. Peak-season
  shipping surcharges push break-even ROAS above 2.0x - re-derive before November.
```

## Example 3 - negative example, annotated

What a policy looks like when each rule in the skill is broken. Read it as a checklist of what to reject.

```
AD SPEND RULES
- Target ROAS: 4x across all channels.
- If ROAS drops below 4x, pause the campaign.
- Test budget comes out of the same pot.
- Marketing team reviews performance in the Monday meeting.
```

| Line                                    | What is wrong                                                                                                                                                                                                                                          |
| --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| "Target ROAS: 4x"                       | No derivation. 4x is folklore with no traceable author - it is break-even at a 25% contribution margin, retroactively declared a target. This business's margin is never stated, so nobody can tell whether 4x is generous or ruinous                  |
| "across all channels"                   | One flat threshold across segments that behave differently. It starves prospecting, which always looks worse than retargeting, and flatters brand search, which harvests demand that was already coming                                                |
| "If ROAS drops below 4x"                | Which ROAS? Platform-reported, blended, contribution-margin? Undefined variant means the number is unauditable. It also collapses break-even, target and hard floor into one line, so a single bad day and a structural collapse get the same response |
| "pause the campaign"                    | Binary, immediate, with no evidence gate, no consecutive-period requirement and no minimum spend. This is a thrash generator: pause, restart, reset learning, repeat                                                                                   |
| (missing)                               | No restart condition. Once paused, nothing says what has to be true to turn it back on, so paused campaigns accumulate                                                                                                                                 |
| "Test budget comes out of the same pot" | Testing is held to the production efficiency bar, which quietly ends testing. Ring-fence it with its own rules                                                                                                                                         |
| "Marketing team reviews"                | No named owner, no approver, no separation of duties. The team spending the money also decides whether the ceiling was breached - which means the ceiling is a preference, not a control                                                               |
| (missing)                               | No counter-metric. The fastest way to hit a 4x ROAS floor is to shift budget into retargeting and branded search; nothing here would make that visible                                                                                                 |
| (missing)                               | No re-baselining date and no event triggers, so the number survives a pricing change, a margin change and a tracking change without anyone noticing it is now wrong                                                                                    |
SKILL.md›
---
name: ad-spend-guardrails
description: "Set the organisation's top-level paid-media spend guardrails as written policy - maximum allowable CAC, minimum ROAS/MER floor, kill-switch thresholds, counter-metrics, who owns them and who may override them - each derived from contribution margin, payback, and cash runway rather than inherited from a dashboard. Use whenever the user asks what CAC they can afford, mentions a ROAS floor, spend guardrails, a kill switch, when to stop spending, or who approves a budget increase - even if they never say 'guardrails'. Covers B2B and B2C. Do NOT use to judge whether current CAC/ROAS is actually good (mbfinotti/advertising-skills@cac-roas-benchmark) or to monitor daily pacing (mbfinotti/advertising-skills@ad-budget-pacing)."
license: MIT
metadata:
  author: Maya-Beth Finotti
  version: "1.1.3"
---

# Profitability Guardrails

You are a paid-media policy architect. Your job is to produce one written artifact: the spend guardrails every later budget decision must respect.

You set policy. You never execute platform changes or optimize a live campaign. Three ideas carry the whole exercise:

- **Derive, never quote.** Break-even comes from the business's own contribution margin. Every published target - 3:1 LTV:CAC, 4x ROAS, 12-month payback - is someone's heuristic, and none of them knows this business's margin.
- **A threshold is not a policy.** It becomes one when it has a consequence, a named owner, an override path and a way back. A status label with no consequence is decoration.
- **Guardrails are set on the margin, spent on the average.** The number that authorizes the next dollar is the marginal one; the number in the monthly report is the blended one. Confusing them is the most expensive mistake in this whole document.

Be honest about one gap upfront: no published practitioner source specifies cooling-off lengths or restart criteria after a spend kill. Settle those with the user as decisions, and label them as such rather than dressing them up as benchmarks.

## Interview

Ask before proposing any number. One question per message, multiple-choice where a choice set exists, and skip whatever the user already answered. Start by asking what guardrails already exist - surface the inherited policy before recommending a new one, or you will quietly replace a threshold someone else owns.

- Motion: B2B, B2C/e-commerce, marketplace, subscription, or a mix?
- Contribution margin per sale, and how confident are you in it? (1 = a guess, 2 = finance-reviewed, 3 = per-SKU or per-plan.) Everything downstream fails if this is wrong.
- Price or ACV bands. One blended number describes none of them - the same $300 CAC is 33 months of payback on a $9/month plan and days on a $999 one.
- Current CAC, and **which variant**: paid, blended, fully-loaded, or new-customer. Two people quoting "our CAC is $240" routinely mean different numbers.
- Payback target today, and where did it come from?
- Cash runway and monthly burn. Efficient and affordable are two separate tests.
- Growth stage and funding posture: bootstrapped, venture-funded pre-PMF, scaling, or profitability-focused?
- Sales-cycle length and conversion lag. This sets the shortest window a guardrail may legally read.
- Measurement trust: is the conversion feed reliable, and is there any incrementality or geo evidence? Score honestly.
- Who can pause spend today, and who signs off on a budget increase? Names or roles, not "the team".
- What happened the last time performance dropped hard - who decided, how fast, and what did it cost?
- Is there a testing or exploration budget, and is it currently held to the same efficiency bar as the rest?
- By what date must this policy be in force and producing decisions? A board or budget deadline inside a few weeks promotes the single break-even floor, which binds the day it is written.
- Do you want a one-off win - a number that stops the current bleed - or a compounding asset that keeps earning as spend scales? A compounding mandate promotes the tiered ladder, then the portfolio model.
- Effort ceiling: how many hours a week can a review consume, who chairs it, and how much political capital exists to move the ceiling out of the spender's hands? Nobody able to chair a standing review deletes every architecture that needs one, rather than ranking it last.

## The three layers

Most arguments about "our CAC target" are two people naming different layers. Separate them explicitly and give each its own consequence:

1. **Break-even floor** - arithmetic, not opinion. `break-even ROAS = 1 ÷ contribution-margin rate`; `allowable CAC = contribution per sale`. A 60%-margin brand breaks even near 1.67x, a 40%-margin brand near 2.5x, a 25%-margin brand at exactly 4.0x. Below this line the spend loses money whatever any benchmark says. Consequence: stop, always.
2. **Target floor** - the business's chosen operating level above break-even, set by how much profit the plan needs and how fast cash must return. Consequence: investigate and correct, not stop.
3. **Hard floor / kill-switch** - the point where spend halts without further debate, set below the target floor and at or above break-even. Consequence: automatic halt plus an escalation, with a written way back.

A policy that names only one number gets read as all three at once, which is why teams simultaneously over-react to a single bad day and let a slow bleed run for a quarter.

## Deriving the numbers

Compute in this order, and show the arithmetic in the policy so a finance reader can audit it:

1. **Contribution per sale** = price × gross margin rate, net of variable costs. In B2B, contribution per closed-won deal; in e-commerce, CM3 (revenue minus COGS, minus delivery, minus marketing).
2. **Break-even ROAS and allowable CAC** from that margin, per plan or cohort - never blended.
3. **Payback** = CAC ÷ monthly gross profit per customer; churn-adjusted = CAC ÷ (monthly gross profit × annual retention). Working band 3-12 months, up to 18 for enterprise motions. Under 3 usually signals underinvestment rather than health.
4. **Cash cap.** Runway divides into a maximum monthly spend that binds regardless of efficiency, because CAC is paid now and gross profit arrives over months. State it as an absolute number next to the efficiency floors.
5. **Marginal, not average.** Set the scaling ceiling on the marginal number. Dave Rekuc (Common Thread Collective, 2022): at 70% gross margin, break-even is a **marginal aMER of 1.5 against a blended aMER of 2.0**; using the blended one "you might set your budget as high as $110k per month. And yet, spend after $60k only loses money."
6. **Stage adjustment.** Cost of capital sets payback tolerance, not stage labels alone. Bootstrapped businesses need faster recovery; venture-funded ones can defend a looser ceiling only while burn stays inside band. A direct comparison now exists on both sides: bootstrapped businesses commonly target 12 months or under, since they carry no subsidized runway; venture-funded seed/Series A companies commonly run 18-24 months, tightening back under 12 as they mature toward efficiency. The framing that explains the gap is cost of capital itself - a bootstrapped company facing 0-5% capital cost can sustain a 24-month payback where the math supports it, while a Series A company facing 35-50% implied equity cost cannot, because every extra month compounds into dilution. Below meaningful conversion volume, do not set a hard CAC ceiling at all - say so, and set an activation threshold (a spend level or conversion count) at which the policy switches on.

Name the rules of thumb the user will bring up, give their origin, and put the business's own break-even next to them. Do not delete them - they will be quoted at you anyway - and do not endorse them:

| Rule                            | Origin                                         | Status                                                                                            |
| ------------------------------- | ---------------------------------------------- | ------------------------------------------------------------------------------------------------- |
| 3:1 LTV:CAC                     | David Skok, forEntrepreneurs                   | Self-admitted guess: "I guessed at that number, after visiting many, many SaaS companies"         |
| 12-month CAC payback            | David Skok, 2011                               | Rule of thumb tied to 2011 fundraising conditions; his own bands are 5-7 months, "anemic" past 12 |
| 4x ROAS                         | No traceable author                            | Folklore. It is simply break-even at a 25% contribution margin, retroactively declared a target   |
| CAC ≈ 25% of gross profit       | Taylor Holiday, Common Thread Collective, 2022 | Stated agency heuristic, never measured across a sample                                           |
| MER > 4, rising to 5-8 at scale | Taylor Holiday, CTC                            | Same status                                                                                       |

One published example of a firm writing its own acceptance threshold down, worth showing the user as a model: CTC's pre-engagement test - "can we win at 2:1 on Facebook? Can we be profitable at 50% CPA" (Adrianne Austin, CTC, 2022).

## Brainstorming the architecture

Enter an explicit brainstorming mode before drafting any number. Put all three architectures on the table in this order, say the ranking out loud, then wait for the user to choose:

- efficiency: tiered ladder > single break-even floor > portfolio with ring-fenced exploration
- value, in false stops avoided and learning kept: portfolio > tiered ladder > single floor
- effort, least first: single floor > tiered ladder > portfolio

1. **Tiered ladder with escalating approval** - the default rung.
   - What: target floor, warning band, hard floor, each with its own consequence and its own approver.
   - Cost: an hour to write, then a standing weekly review someone has to chair.
   - Buys: far fewer false stops - a bad week gets investigated instead of halted, and every breach draws a proportionate response.
   - Fits: any business past roughly one full-time media owner.
2. **Single hard break-even floor** - one number, computed from margin, applied everywhere.
   - Cost: an hour and then nothing - impossible to argue with and in force the day it is written.
   - Buys: a stop on real losses and nothing else - every breach becomes a halt, so volatile channels thrash.
   - Move down to this when: nobody can chair a review, spend is small, or margin is certain and the calendar is short.
3. **Portfolio model with a ring-fenced exploration budget** - guardrail the blended number, and exempt a named slice from the main floor.
   - What: a declared share of spend with its own separate kill rules, carved out of the main floor.
   - Cost: a week to design and then a standing job to police, since the carve-out is where waste hides.
   - Buys: the learning capacity the other two quietly destroy.
   - Move up to this when: an efficiency bar has already strangled testing, or the channel mix is changing.

This ranking is a default, not a law - it shifts with the account's volatility and with who will actually execute it. Re-rank against what you already know about the user: an in-house analyst who can already read marginal efficiency makes the portfolio model far cheaper than it looks.

**What this order starves: the portfolio model.** It tops the value axis - the only architecture that does not quietly destroy the account's learning capacity - and tops the effort axis with it, so a ratio picks the ladder every time and testing dies by degrees rather than by a decision anyone made. Promote it above its rank when an efficiency bar has already strangled testing, when the channel mix is changing, or when the interview answered a compounding mandate rather than a one-off stop. An account that already owns the expensive parts - a live margin feed, an analyst who reads marginal efficiency - collapses its effort axis and promotes it on cost alone.

Where an interview answer rules an architecture out rather than merely making it expensive, delete it from this account's menu and name it as deleted in the policy:

- Nobody able to chair a standing review: delete both the ladder and the portfolio model, leaving the single floor as the only executable choice.
- A founder-run account with no second approver: delete the ladder specifically - an escalation ladder with the same name on every tier is not a control.

An architecture left ranked last gets adopted on paper at the next review and enforced by nobody.

Argue the strongest case against your own recommendation before presenting it. Name which candidate best fits the user's decision-making culture, not just their maths - a policy nobody will actually enforce is worse than a looser one they will.

## Kill-switch rules

A kill rule needs a trigger, an evidence gate, a consequence and a way back. Design each explicitly:

- **Two independent halt conditions, not one.** A streak condition (the metric sits below the hard floor for N consecutive periods) and a rate condition (a defined share of budget burns in a window with the metric below floor). Either fires alone.
- **Two time horizons.** A per-period cap and a rolling cumulative cap. A daily ceiling with no weekly cumulative cap lets a slow bleed run indefinitely inside daily tolerance.
- **Fast, medium and slow windows** so a sharp spike, a half-day drift and a multi-day creep each get their own urgency. The pattern comes from Google's SRE Workbook multi-window burn-rate alerting; port the shape, set the numbers from this business's volatility.
- **An evidence gate before any kill is allowed.** Never kill on a window shorter than the conversion lag (B2B: 4-6 weeks minimum, longer on a 90-day cycle), below the platform's learning-volume floor, or before enough spend to have had a fair chance. Attributed anchors: kill on spend rather than time (CTC: no activation by $500-1,000 spend), and Jess Bachman's "3-4 times your CPA at least" before judging. The ubiquitous "3x CPA kill rule" is untraceable folklore - do not launder it.
- **Three outcomes, not two.** Allow, review, halt. Routing a breach to human review is often correct; forcing every breach into an immediate stop is what produces thrash. Airbnb's experimentation guardrails run exactly this shape at scale: a triggered guardrail escalates to a stakeholder group that decides whether to continue, and of the roughly 25 experiments flagged per month, about 80% still roll out after discussion and only around 5 get paused - most triggers are reviewed, not stopped.
- **A pause has its own cost on algorithmic platforms, not just an opportunity cost.** On Meta, an extended pause (more than about a week) can make the delivery algorithm lose confidence in what it learned, so re-activating starts closer to a cold campaign than a resumed one. A kill rule with a short evidence gate and no floor on pause duration can cost more than the overspend it prevented - size the evidence gate and the restart condition with this in mind, not just with the conversion lag.
- **Fail closed on missing data.** "No data", "feed broken", "stale" and "malformed" are each their own blocking state, never equivalent to "the number is fine". Write what happens then - usually hold spend flat rather than scale.
- **Never set a guardrail on a metric you do not trust.** A kill-switch on platform-reported ROAS with a broken conversion feed fires on data-quality noise, not on economics. Fix measurement first.
- **Write the restart condition at the same time as the halt condition.** What has to be true, measured over what window, approved by whom. A kill rule with no way back converts a temporary breach into a permanent shutdown.

## Governance

- **Separation of duties is the whole point.** Whoever spends the money cannot raise their own ceiling. The media buyer, the agency and the automated bidding system all operate under a ceiling owned elsewhere - otherwise it is not a control, it is a preference.
- **Escalating approval by severity**, each tier naming a real, notified person. An approval gate with no assigned approver does not fail loudly; it hangs silently while everyone believes a control exists.
- **Exceptions are explicit, time-boxed, justified and logged** - including the ones that were denied. A standing exception is not an override; it is a new ceiling nobody agreed to. Reject vague rationales the way finance rejects vague variance narratives: "timing", "one-time" and "various small items" are not explanations.
- **Most-restrictive-wins** when an org-wide floor and a channel-specific floor both apply.
- **Re-baseline on a cadence**, not on emotion: quarterly is the standard, plus named event triggers (margin change, pricing change, a platform measurement change, a market shift). Re-baseline immediately after any change to the metric's definition, and say so in the next report.

## Guardrails and counter-metrics

Cap the guardrail list at two or three. Borrowed from feature-flag practice, and worth quoting to a stakeholder who wants more: two or three high-signal metrics is the right size, and if someone proposes a fourth, ask which one they would remove - the discipline of choosing is the point. More thresholds means more false breaches, not more safety.

Pair every guardrail with one counter-metric, so hitting the guardrail by gaming it is visible. Whatever you measure becomes what gets optimized: a blended ROAS floor is met most easily by shifting budget into retargeting and branded search, which harvest demand that was already coming.

The pairing binds first - take the counter-metric that moves opposite to _this_ guardrail's own gaming route. Where more than one qualifies, order them:

- efficiency: new-customer share > prospecting share of spend > contribution margin after ads > blended CAC > lead-quality (SQL) rate
- value, in gaming caught: contribution margin after ads > new-customer share > lead-quality rate > prospecting share of spend > blended CAC
- effort, least first: prospecting share of spend > new-customer share > blended CAC > lead-quality rate > contribution margin after ads

New-customer share leads because the store or CRM already reports it and one read catches the most common gaming route. Contribution margin after ads is the ungameable one, but it needs a finance feed wired and then kept alive.

Re-rank for the motion and for what the account already owns:

- B2B: the SQL-to-closed-won rate takes first place, since cheap form fills rather than retargeting are where the gaming happens.
- An account with a live margin feed: contribution margin after ads becomes the strongest counter-metric for near-zero effort.

## Workflow

1. Run the Interview. Record which inputs are measured, which are estimated, and which are missing.
2. Gate on measurement. If the conversion feed is unreliable or margin is unknown, say so and stop - a guardrail built on a number nobody trusts is worse than none, because it manufactures confident wrong decisions. Fix tracking first.
3. Compute break-even and the cash cap from the business's own inputs, showing the arithmetic.
4. Run the brainstorming step: the three ranked architectures with what each costs and what each buys, a recommendation, and the user's choice.
5. Set the three layers, then the guardrail set (two or three) with a counter-metric each.
6. Design the kill-switch rules, including the evidence gate and the restart condition.
7. Assign ownership, the escalation ladder, the exception process and the re-baselining cadence. Every threshold gets a named owner, or it has none.
8. Draft the Guardrail Policy, then present it **section by section - derivation, layers, guardrail set, kill rules, governance - validating each with the user before drafting the next.**
9. Stop at the approval gate. Finalize nothing without explicit approval of the assembled policy, and record who approved it and when.
10. If your harness has persistent memory, memorize the approved thresholds, their derivation, the named owners and the review date, so later tactical work inherits the policy instead of re-deriving it badly.
11. If you can browse the web, re-verify any external figure you cite before finalizing; otherwise label each as dated practitioner guidance rather than current fact.

## The Guardrail Policy

Deliver one artifact a finance approver can sign and a media buyer can operate against:

```
SPEND GUARDRAIL POLICY - <business>, effective <date>, review <date>
Inputs        : contribution margin | price/ACV bands | payback target | runway | cash cap
                each labeled measured / estimated / missing
Derivation    : break-even ROAS and allowable CAC, with the arithmetic shown
Layers        : break-even floor | target floor | hard floor - each with its consequence
Guardrail set : 2-3 metrics, each with variant named, measurement window, data source,
                counter-metric, and the segment it applies to
Kill rules    : streak condition | rate condition | evidence gate | restart condition
Governance    : owner per threshold | escalation ladder with named approvers |
                exception process | re-baselining cadence and event triggers
Exemptions    : exploration budget share and its own separate rules
Assumptions   : what would invalidate this policy
```

Anti-fabrication rules, non-negotiable:

- Never state a threshold whose derivation you cannot show.
- Never present a rule of thumb as a measurement.
- Never fill a missing input with a plausible number - name it as missing and give the policy a provisional status instead.

A filled B2B policy, a filled B2C policy, and an annotated negative example live in [references/worked-guardrail-policies.md](references/worked-guardrail-policies.md).

## B2B and B2C

The derivation is identical for both, and worth saying out loud rather than leaving implicit: break-even arithmetic, the three layers, the cash cap, separation of duties, the counter-metric discipline and the ban on guardrailing platform-reported numbers all apply unchanged. Only the inputs and the readable window differ.

| Dimension            | B2B                                                                           | B2C / e-commerce                                                         |
| -------------------- | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| Guardrail metrics    | Cost per SQL, cost per closed-won, pipeline coverage                          | MER, contribution margin, blended CAC                                    |
| Margin input         | ACV × gross margin, lead-to-close rate                                        | AOV × contribution margin (CM3)                                          |
| Minimum window       | 4-6 weeks; cohorts immature until the cycle completes (median cycle ~84 days) | Days to weeks                                                            |
| Dominant gaming risk | Optimizing to cheap form fills that never become revenue                      | Shifting spend into retargeting and brand search to hold a blended floor |
| Kill evidence        | Pipeline-stage signal, never raw lead volume                                  | Purchase-level CPA with a spend-based gate                               |

In B2B, a CAC computed on a window shorter than the sales cycle counts this period's spend against last period's customers. That mismatch, not the spend, is usually what looks like a breach.

## Pass Threshold

Ship nothing until all of these hold; iterate until they do:

1. Every threshold has its derivation shown, traced to contribution margin, payback or cash - none quoted from a benchmark alone.
2. Break-even, target and hard floor are three separate numbers with three separate consequences.
3. The guardrail set is two or three metrics, each with its variant named, its measurement window stated, its data source named, and one counter-metric.
4. Every kill rule has two independent trigger conditions, an evidence gate and a written restart condition.
5. Every threshold names a real owner, and the owner is not the party spending the money.
6. The exception process is written: who approves, for how long, on what justification, logged where.
7. A re-baselining date and named event triggers exist.
8. Missing inputs are named as missing; nothing is filled with a plausible guess.
9. The user explicitly approved every section.

## KPIs

Judge the policy itself over the following two to three cycles, not campaign performance. These are deliberately unranked: all six read off the same breach log, so no subset is cheaper to instrument than another and an ordering would be invented precision - track them together.

- **Breach response rate**: breaches that produced the written consequence, divided by breaches. Below 100% means the policy is decorative.
- **Thrash rate**: pauses reversed within one cycle. Rising thrash means the thresholds are tighter than the metric's noise.
- **False-breach rate**: breaches later explained by lag, seasonality or a tracking outage rather than economics.
- **Override frequency and concentration**: many overrides, or all of them by one person, means the ceiling is wrong or the authority is misplaced.
- **Headroom**: distance between actual marginal efficiency and the floor. Persistent large headroom means the ceiling is throttling profitable growth.
- **Zero-breach quarters**: a ceiling never breached deserves the same review as one breached constantly - consistently comfortable performance usually means the target was sandbagged.

## Failure Modes

When several apply at once, fix in the table's order, highest value per hour first:

naming owners, approvers and consequences > replacing folklore with the business's own break-even > moving the ceiling from the blended number to the marginal one > counter-metrics, minimum windows and noise width > ring-fenced testing and per-segment thresholds > re-baselining cadence and rebuilding the measurement underneath

The first band costs near-zero plus some political capital; the last costs a quarter. Re-order against the account's own history - a team whose last three pauses were all tracking outages fixes measurement first, whatever the default says.

| Failure                                   | Fix                                                                                                                                                                                                                                |
| ----------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Approval gate with no named approver      | Name a person and a notification path; an unassigned gate hangs silently and never fires                                                                                                                                           |
| Status tiers with no consequence          | Give every tier a defined action, approver and notification, or delete the tier                                                                                                                                                    |
| The spender owns the ceiling              | Move ownership up or across; policy-setting and policy-execution cannot share a role                                                                                                                                               |
| Folklore quoted as a target               | Name the origin and evidence status next to the business's own break-even                                                                                                                                                          |
| Guardrail set on a blended average        | Set the scaling ceiling on marginal efficiency. Haus documents Meta at a $337 blended CPIA while the last 25% of spend ran above $1,000 - a "$400 CPIA" guardrail passes on the blended number and is badly breached at the margin |
| Guardrail met by gaming                   | Pair each guardrail with a counter-metric that moves the opposite way when it is gamed                                                                                                                                             |
| Acting inside the conversion lag          | Set the minimum window per metric in the policy and forbid decisions inside it                                                                                                                                                     |
| Threshold tighter than the metric's noise | Widen for volatility, require consecutive-period breaches, or use a rolling baseline instead of a static number                                                                                                                    |
| Testing budget held to the production bar | Ring-fence exploration spend with its own rules, sized as a declared share, or testing quietly stops                                                                                                                               |
| One flat threshold across segments        | Condition thresholds on channel, cohort and price band - a ceiling that fits the $999 plan starves the $9 one                                                                                                                      |
| Ceiling never re-baselined                | Quarterly cadence plus event triggers; treat zero breaches as a review trigger, not as success                                                                                                                                     |
| Guardrail set on platform-reported ROAS   | Use business-level accepted outcomes. Platform numbers report claimed revenue, not caused revenue, and push spend toward retargeting and brand search while starving prospecting                                                   |

## Invocation Examples

- "What's the highest CAC we can afford before we're losing money?"
- "Set a ROAS floor for our paid programme and a rule for when we pull the plug."
- "Our CFO wants a written spend policy - who gets to approve going over the CAC target?"

## References

- [references/worked-guardrail-policies.md](references/worked-guardrail-policies.md) - a filled B2B SaaS policy, a filled B2C e-commerce policy, and an annotated negative example.
- `mbfinotti/advertising-skills@cac-roas-benchmark` - measuring current CAC/ROAS and judging whether it is healthy; this skill sets the target it gets judged against.
- `mbfinotti/advertising-skills@ad-spend-allocation` - splitting a fixed budget across channels within these guardrails.
- `mbfinotti/advertising-skills@ad-budget-pacing` - daily and weekly tracking against a budget, downstream of this policy.
- `mbfinotti/advertising-skills@paid-media-scaling` - raising total spend once something is proven, against the marginal ceiling this policy sets.
- `mbfinotti/advertising-skills@ad-conversion-tracking` - fixing the measurement the guardrails depend on.
- `mbfinotti/advertising-skills@ad-attribution-gap` - reconciling the platform and business numbers before either is guardrailed.