SKILL DETAIL
paid-media-scaling
mbfinotti/advertising-skills/paid-media-scaling
Decide when a proven paid campaign has earned a budget increase, how large each step should be, how the ramp sequences over weeks and months, and how to avoid performance collapse on the way up - including readiness gates, vertical vs horizontal scaling, and rollback triggers. Use whenever the user asks whether to scale a campaign, how fast ad spend can rise, mentions a budget ramp, a scaling ceiling, or says performance collapsed after the budget was raised - even if they never say 'scaling'. Covers B2B and B2C. Do NOT use to split a fixed total across campaigns (mbfinotti/advertising-skills@ad-spend-allocation) or to track daily spend against a set budget (mbfinotti/advertising-skills@ad-budget-pacing).
Installation
npx skills add https://github.com/mbfinotti/advertising-skills --skill paid-media-scaling
Skill-Dateien
SKILL.md
Zuletzt synchronisiert · 24.09.2026
evals/evals.json›
{
"skill_name": "paid-media-scaling",
"evals": [
{
"id": 1,
"prompt": "I run Meta ads for Velora Skincare. Our prospecting campaign has spent $500/day for the past 8 weeks with a stable $38 CPA against our $55 max. I read that Facebook's official rule is that you can raise budgets 20% every 72 hours without resetting learning, so I want to follow the official rule up to $2,000/day for the holiday push. For context I have the full change history: we bumped this campaign +30% in January, +15% in April, and +25% in June, and I can pull how CPA responded to each. Lay out the schedule.",
"expected_output": "A ramp response that corrects the '20% every 72 hours' premise as folklore rather than a platform rule, derives the step size from the account's own three past budget changes, sets learning-plus-lag holds, runs readiness checks, and pre-commits a rollback before any schedule.",
"files": [],
"expectations": [
"States that Meta has never published a budget-change percentage and that '20% every 72 hours' is not an official platform rule.",
"Labels the 20%-every-72-hours figure as folklore or unverified practitioner lore, not documented platform guidance.",
"Attributes the 20%/72h claim to second-hand practitioner relay of unverified account-rep advice rather than to Meta documentation.",
"Distinguishes the real platform artifacts behind the folklore (Meta's 'Increase budget by 20%' automated-rule UI preset, and/or Google's 20% guidance applying to bids on Display campaigns only) from an actual budget rule.",
"States what Meta actually documents: significant edits reset the learning phase, with no budget percentage named.",
"Derives the recommended step size from the account's own change history (the January +30%, April +15%, June +25% responses) rather than defaulting to a universal percentage.",
"Sets the hold between steps to at least the learning window plus conversion lag (about a week or more on Meta), not a flat 72 hours.",
"Runs or requests readiness checks (marginal economics, creative supply, cash, measurement health) before committing to the step schedule.",
"Pre-commits a rollback trigger, rollback action, and verification date before or alongside the first up-step.",
"No single step in the proposed schedule raises budget 30% or more in one move without an explicit learning-reset warning.",
"Any cited step-size or cadence number carries an evidence label (documented, research, or folklore).",
"If a concrete default percentage such as 15-20% appears, it travels with an explicit instruction to recalibrate from this account's own response data.",
"The schedule reaches $2,000/day through multiple held steps, and the response names what is monitored during each hold."
]
},
{
"id": 2,
"prompt": "Halberd Outfitters here, DTC menswear. Owner-approved break-even MER is 2.0. Blended MER is currently 2.7 and has been above break-even all summer, so we feel very safe. Main prospecting campaign monthly figures: June $40K spend, $122K revenue. July $48K spend, $140K revenue. August $56K spend, $152K revenue. For September we want to jump straight to $80K since blended has so much cushion. Sanity-check the plan and give us the go-ahead.",
"expected_output": "A refusal of the go-ahead grounded in marginal-band arithmetic: the newest band returns about 1.5 against a 2.0 break-even even though blended sits at 2.7, so the marginal gate fails and the response stops the ramp and names the fix instead of shipping a smaller one.",
"files": [],
"expectations": [
"Computes marginal return per spend band as delta revenue divided by delta spend, rather than judging the campaign on blended MER alone.",
"Identifies the June-to-July marginal band as approximately 2.25 ($18K additional revenue on $8K additional spend).",
"Identifies the July-to-August marginal band as approximately 1.5 ($12K additional revenue on $8K additional spend).",
"States that the most recent marginal band (about 1.5) is below the owner-approved 2.0 break-even boundary.",
"Declines to give the go-ahead on the $80K jump despite blended MER of about 2.7 sitting above break-even.",
"Explains that blended metrics always trail marginal: marginal efficiency turns unprofitable before blended does.",
"Frames the decision as whether the next dollar of spend still makes money, not how much the account can spend.",
"Notes the declining trend across bands (about 2.25 down to about 1.5) as evidence efficiency is deteriorating as spend grows.",
"Treats the failed marginal gate as stop-and-fix: names the failed gate and its fix rather than proposing a smaller budget increase as a consolation.",
"Flags the proposed $56K-to-$80K move (over 40% in one step) as an oversized single step that risks a learning-phase reset.",
"Asks whether the revenue figures are platform-attributed, triangulated with business data, or causally measured.",
"Recommends diagnosing why marginal efficiency is falling (for example saturation signals: frequency, CPM, unique reach, penetration) before adding any budget.",
"Any stop or exit rule offered references marginal contribution or the marginal band crossing the boundary, not blended MER falling below 2.0."
]
},
{
"id": 3,
"prompt": "Growth lead at Finchley Software, B2B SaaS with self-serve plus sales-assist. Our in-platform numbers: branded search runs an 11x ROAS, retargeting 7.8x, cold prospecting only 2.1x. Everything is platform reporting, we have never run any lift testing. The CFO just approved tripling total paid from $60K to $180K/month, and my plan is to put most of the increase behind branded search and retargeting since those are the proven winners. Draft the scale-up plan.",
"expected_output": "A plan that refuses to pour the tripled budget into branded search and retargeting on attributed numbers alone: it flags those two lines as exactly where attribution overstates most, requires a causal measurement step for a 3x ramp, and promotes measure-first, without ever claiming the current spend is wasted.",
"files": [],
"expectations": [
"Flags that platform-attributed ROAS overstates incremental value most on retargeting and branded search, the two lines the user wants to scale.",
"States that a ramp to roughly 3x current spend on purely attributed evidence requires a causal measurement step (holdout or incrementality test) or an explicit user-acknowledged risk line.",
"Recommends a geo-holdout or incrementality test on branded search and/or retargeting before the bulk of the increase lands there.",
"References the eBay paid-search experiment (attributed returns in the thousands of percent versus measured negative ROI) as the research-grade warning case for scaling on attribution alone.",
"Does not claim the account's current branded or retargeting spend is wasted and assigns no percentage of waste: the evidence motivates a test, never substitutes for one.",
"Classifies the account's current evidence bar as attributed (the weakest tier) and names the upgrade path toward triangulated or causal.",
"Questions branded search's headroom on demand-capture grounds: search captures existing demand and cannot expand the number of people searching.",
"Checks retargeting expansion against its finite pool: audience size and frequency limits rather than treating it as freely scalable.",
"Presents approach candidates and promotes measure-first ahead of immediately laddering budgets, given the 3x target and attributed-only evidence.",
"Says which condition promoted measure-first (target beyond 2x and/or attributed-only evidence) instead of presenting the choice unexplained.",
"Notes practitioner geo-holdout findings that reported cold ROAS around 3x often measures nearer 1.8-2.2x incrementally, labeled as vendor or practitioner data.",
"Any interim budget steps carry pre-committed hold periods and rollback triggers.",
"Labels cited figures by evidence class: platform numbers as attributed, the eBay result as research, holdout benchmarks as vendor or practitioner data.",
"The plan does not concentrate the entire increase on branded search and retargeting: some of the allocation logic accounts for the demand-capture ceiling or directs growth toward demand creation."
]
},
{
"id": 4,
"prompt": "We sell compliance software - ACV $28K, roughly 4 months from first click to closed deal. Six weeks ago we scaled LinkedIn from $12K to $20K/month and it's working great: CPL dropped from $95 to $61. I want to push to $45K/month next month, and we'll evaluate each budget bump after one week so we can move fast. Two side notes: sales mentioned discovery calls have felt lighter lately, but they always complain. And LinkedIn reports more conversions than our CRM shows - I trust LinkedIn's numbers since the CRM always lags. Build the ramp.",
"expected_output": "A response that treats the falling CPL as a possibly broken proxy and the sales remark as a lead-quality decay signal, judges steps on cost per SQL and lead quality over month-plus holds instead of one week, makes the CRM the arbiter over the platform, and gates the push to $45K on verifying quality first.",
"files": [],
"expectations": [
"Flags the falling CPL ($95 to $61 after scaling) as a possible broken proxy to verify, not as proof the scale-up is working.",
"Treats the sales team's 'discovery calls feel lighter' remark as a lead-quality decay signal to investigate rather than dismissing it.",
"Requires steps to be judged on cost per SQL and/or a lead-quality score, never on CPL alone.",
"Rejects the one-week evaluation window: holds must cover learning plus conversion lag, and B2B steps are judged on windows of a month or more using leading indicators.",
"Reverses the user's stated preference: when platform-reported conversions and the CRM disagree, the CRM wins.",
"Requires the offline conversion loop (CRM stage changes fed back to the platform) closed before scaling further on any lead metric.",
"Recommends reconciling platform conversions against the CRM on a recurring monthly basis.",
"Freezes or gates the push to $45K on verifying lead quality first: the proxy check precedes the raise.",
"Flags the $20K-to-$45K move (more than doubling in one step) as oversized and replaces it with multiple held steps.",
"Checks whether sales can absorb the added lead volume without speed-to-lead degrading.",
"Checks cash absorption: ad spend bills in days while B2B pipeline revenue lags by months (up to the 60-281 day range), so working capital must cover the gap.",
"Defines a per-step monitor set built on leading indicators (cost per SQL, lead-quality score) with closed-won as the lagging verdict.",
"Pre-commits a lead-quality-keyed rollback, for example cost per qualified lead above 1.5x target triggering a 20-30% cut, stabilization, then a slower resume.",
"Anchors affordability in a per-cohort or per-deal payback calculation from the $28K ACV and margin, or asks for the owner-approved CAC boundary, rather than inventing a target."
]
},
{
"id": 5,
"prompt": "The board committed us to $300K/month of paid media by end of Q3 - we're at $95K today. Ecommerce, Meta plus Google, blended ROAS 3.1, all measurement is platform reporting plus GA4. 30-day penetration on our core audience is about 12%, and our in-house studio ships 10 new creatives a month. My plan is simple: raise budgets 20% every week on both accounts until we hit the number. Write that up as a formal plan I can hand to the board.",
"expected_output": "Instead of formatting the user's ladder, the response interviews, brainstorms the three approaches with an explicit ranking, promotes measure-first because the 3.2x target rides on attributed evidence, and delivers a structured ramp-plan artifact with gates, dated held steps, rollbacks, ceilings, and an exit, pending the user's approval.",
"files": [],
"expectations": [
"Does not simply format the user's 20%-weekly ladder into a plan: the approach choice is examined before drafting.",
"Asks the user clarifying questions before finalizing the plan, covering at least one of: hardness of the Q3 date, one-off target versus compounding asset, or the effort ceiling.",
"Presents vertical ladder, measure-first, and horizontal expansion as named candidate approaches with trade-offs, and requests the user's pick.",
"States the candidate ordering explicitly rather than leaving the ranking implied by list position.",
"Promotes measure-first given the target is over 2x current spend (about 3.2x) on platform-attributed evidence, and names which condition triggered the promotion.",
"Re-ranks using account specifics: the in-house studio shipping 10 creatives a month is cited as lowering horizontal expansion's effort.",
"Applies the penetration bands: about 12% is under the 25% line, so vertical headroom remains available.",
"The ramp includes a causal measurement step (geo holdout or incrementality test) or an explicit user-acknowledged risk line for scaling to 3x on attributed evidence.",
"Each proposed step carries a date or hold-until point, a monitor set, a rollback trigger, and a rollback action.",
"The flat raise-and-judge-weekly cadence is not endorsed as-is: hold length is tied to learning window plus conversion lag, with longer evaluation windows noted where Google (for example Performance Max) is involved.",
"The plan names its exit condition and which ceiling is expected to bind first.",
"The deliverable follows a structured ramp-plan shape containing at least gates, evidence bar, approach, steps, ceilings, and exit or open items.",
"The plan is presented for the user's section-by-section validation or explicit approval, and the assistant does not execute or claim to execute budget changes.",
"Numbers carry documented/research/folklore labels, and any default step percentage travels with a recalibrate-from-account-history instruction."
]
},
{
"id": 6,
"prompt": "Bit of an emergency. We doubled our Google Ads budget on the 3rd - $800 to $1,600/day - on our roofing-leads campaign. It's the 8th and CPA has doubled from $90 to $180, although that's on just 7 conversions so far. Click to booked job normally takes about 3 weeks for us. My plan: kill the campaign today, then starting tomorrow make one fix per day - Monday tighten locations, Tuesday new headlines, Wednesday adjust the bid target - so we can isolate what actually works. Sound right?",
"expected_output": "A response that calls the doubled CPA on 7 conversions with a 3-week lag noise, refuses the same-day kill and the one-fix-per-day sequence, reverts to the prior budget as the smallest reversible action, batches any further edits into one session, and pre-commits rollback rules before a slower re-approach.",
"files": [],
"expectations": [
"States that a doubled CPA on 7 conversions is too small a sample to conclude performance collapsed.",
"Notes the roughly 3-week conversion lag means conversions attributable to the new spend have not landed yet, so the 5-day CPA read is structurally overstated.",
"Checks or asks about other false-positive causes before acting: tracking outages, seasonality, or running experiments.",
"Declines to kill the campaign today, reserving instant cuts for genuine emergencies (runaway spend, policy or legal exposure, broken destination, confirmed tracking corruption) that this situation does not match.",
"Recommends the smallest reversible action, such as reverting to the prior $800/day budget, over pausing or killing the campaign.",
"Rejects the one-fix-per-day sequence: drip-feeding edits across days can trigger repeated learning resets, and changes should be batched into one session.",
"Corrects the isolation logic: sequential daily edits during learning-phase instability do not isolate causes.",
"Identifies the original single +100% jump as an oversized step likely to have reset learning and contributed to the instability.",
"The re-approach plan uses smaller held steps with a slower resume (on the order of +10% per period after stabilization) rather than re-doubling.",
"A rollback trigger, rollback action, and verification date are pre-committed before any renewed up-move.",
"Any judgment window for the re-approach covers the learning window plus the roughly 3-week lag, not another 5-day read.",
"Distinguishes the persistence-based rollback convention (an efficiency drop persisting about 5-7 days at adequate sample before reverting) from an instant same-day kill, labeling it practitioner convention rather than platform rule.",
"Monitoring during the re-approach includes leading indicators and delivery/learning status rather than CPA alone inside the lag window."
]
},
{
"id": 7,
"prompt": "DTC supplements brand. One hero Meta prospecting campaign at $70K/month that we've scaled all year, still hitting our 2.2 target MER on blended. Latest 30-day data: audience penetration 37%, frequency 3.6 and climbing, CPMs up 31% year over year, CTR flat, and unique reach has now declined two months in a row. I want to take this same campaign to $140K over the next 8 weeks with 15% weekly bumps. We hold 9 proven ads and can produce 6 new ones a month. Lay out the ramp.",
"expected_output": "A refusal of the vertical ramp: penetration 37% with declining unique reach, elevated frequency, and rising CPM at flat CTR puts the audience past the 35% band, so vertical is deleted with its constraint named and the increase is redirected to horizontal expansion at proven per-unit budgets with its own creative and overlap checks.",
"files": [],
"expectations": [
"Reads 37% penetration against the saturation bands (under 25% headroom remains, 25-35% hold, 35%+ scale horizontally) and refuses the requested vertical ramp on this campaign.",
"Treats two consecutive months of declining unique reach as the leading saturation indicator.",
"Notes frequency 3.6 sits above the roughly 3.0 prospecting danger band, labeling the band as practitioner folklore rather than platform rule.",
"Reads rising CPM with flat CTR as saturation corroboration rather than as a bidding or creative problem.",
"States that doubling budget grows penetration only about 50-70%, not 100%, so the target cannot be bought vertically on this audience.",
"Deletes the vertical ladder for this account rather than merely demoting it, and names the saturation constraint that deleted it.",
"Redirects the increase to horizontal expansion: new audiences, geos, placements, or channels at proven per-unit budgets.",
"Notes each new horizontal unit runs its own learning phase and permanently multiplies creative demand, and checks the 6-new-ads-per-month pipeline against that demand.",
"Warns about audience overlap or cannibalization between new and existing units and calls for exclusions.",
"Grounds horizontal expansion in penetration-and-light-buyer growth evidence, citing Ehrenberg-Bass or labeling the rationale as research-supported.",
"Does not treat the campaign still hitting 2.2 blended MER as license to push vertically: asks for or computes marginal efficiency on the newest spend bands.",
"Each horizontal unit or step still carries a hold period, monitor set, and rollback trigger.",
"Names an exit condition for the ramp (such as marginal contribution crossing zero or a penetration band) rather than running increases indefinitely."
]
},
{
"id": 8,
"prompt": "I just took over paid media at Brontide Labs - B2B, webinar-lead campaigns, mostly LinkedIn and Meta. Spend is $25K/month, and I'm expected to be at $75K/month within three months; the CFO already signed off on the media budget. We have 4 ads that reliably work, and a freelancer produces about 2 new ones a month. We pay the ad accounts by card weekly, customers pay net-60. One quirk: the account has never had a budget change - same daily budgets since launch in March. Marginal numbers look decent and our tracking audit scored 12 out of 15. Give me the step plan.",
"expected_output": "A gate-first response: creative supply fails (about 15 proven ads needed at $75K against 4 on hand, and a 2-per-month pipeline cannot close it in time), so the ramp stops and names the fix; cash absorption is flagged on the weekly-card versus net-60 mismatch; and with an empty change log the step size starts at the labeled 15-20% folklore default with a recalibration instruction.",
"files": [],
"expectations": [
"Runs the creative-supply gate arithmetic: roughly 15 proven ads needed at $75K/month (monthly budget divided by $5,000) against the 4 on hand.",
"Labels the budget-divided-by-$5,000 proven-ad ratio as folklore with an instruction to calibrate it to this account.",
"Shows the pipeline cannot close the creative gap on the three-month timeline: about 2 new ads per month against a deficit of roughly 11 proven ads.",
"Treats the failed creative gate as stop-and-fix (fix the creative pipeline, let creative supply set ramp speed) rather than silently shipping a smaller ramp.",
"Flags the payment-terms mismatch (card billed weekly, customers pay net-60) as a business-absorption check, distinguishing CFO media sign-off from approval of the working capital the ramp consumes.",
"States the absorption principle that a profitable account can still break the company on cash, or quantifies the float the ramp consumes at the target spend.",
"Because the change log is empty, starts the step size from the concrete default (15-20% per step, never 30% or more in one move), explicitly because there is no account history to read.",
"Ships the default step with the instruction to recalibrate from the account's own response after the first few steps, moving to history-derived sizing.",
"Sets holds covering learning plus B2B conversion lag (weeks to a month or more), judged on leading indicators rather than closed revenue.",
"Every step carries a pre-committed rollback trigger, rollback action, and verification date.",
"Step-size and cadence figures carry documented/research/folklore labels, with nothing presented as a platform rule.",
"Steps are judged on lead quality or cost per qualified lead rather than raw cost-per-lead alone.",
"Names the exit condition and identifies creative supply as the first-binding ceiling given the arithmetic."
]
}
],
"trigger_queries": [
{ "query": "Our Meta prospecting campaign has held a $42 CPA for two months at $300/day. How fast can we get it to $1,500/day?", "should_trigger": true },
{ "query": "Can we double our Google Ads budget without wrecking performance?", "should_trigger": true },
{ "query": "Build a ramp plan to take paid social from $20K to $80K a month.", "should_trigger": true },
{ "query": "The board wants ad spend at $250K/month by Q2 and we're at $70K. Plan the ramp.", "should_trigger": true },
{ "query": "We tripled the budget last week and ROAS fell off a cliff. How do we re-approach?", "should_trigger": true },
{ "query": "Is it safe to increase the budget on our winning campaign yet?", "should_trigger": true },
{ "query": "How much can I raise my TikTok ad budget each week without resetting learning?", "should_trigger": true },
{ "query": "When has a campaign earned a bigger budget?", "should_trigger": true },
{ "query": "Is the 20% every 72 hours rule for Facebook budgets real?", "should_trigger": true },
{ "query": "Performance collapsed right after we raised spend. What went wrong and what now?", "should_trigger": true },
{ "query": "We keep hitting a ceiling around $50K/month on paid - every push above it tanks efficiency.", "should_trigger": true },
{ "query": "wanna pour way more money into this ad campaign, how fast can i go", "should_trigger": true },
{ "query": "What's a safe weekly budget increase for a LinkedIn campaign that's beating its CPL target?", "should_trigger": true },
{ "query": "Our CAC holds at target - should we scale vertically on the same campaign or expand to new audiences?", "should_trigger": true },
{ "query": "Draft a 12-week budget ramp for our lead-gen campaign, $8K to $30K a month.", "should_trigger": true },
{ "query": "My agency says raise budgets 15% every 4 days. Should I trust that cadence?", "should_trigger": true },
{ "query": "CEO wants us to 4x the ad budget by Black Friday. Talk me through doing that safely.", "should_trigger": true },
{ "query": "The campaign is printing money - how hard can we push it before it breaks?", "should_trigger": true },
{ "query": "At what point do I stop increasing this campaign's budget?", "should_trigger": true },
{ "query": "How long should I wait between budget increases on Meta?", "should_trigger": true },
{ "query": "We raised the daily budget and the CPA doubled. Should we roll back or wait it out?", "should_trigger": true },
{ "query": "What are the signs a campaign is ready for more budget?", "should_trigger": true },
{ "query": "Scale plan needed: $5K/day to $12K/day on our search campaigns before the seasonal peak.", "should_trigger": true },
{ "query": "If I bump the budget 50% in one go, will the algorithm freak out?", "should_trigger": true },
{ "query": "Our investor wants to know how quickly the paid channel can absorb another $1M a year.", "should_trigger": true },
{ "query": "Every time we scale past $2K/day the results fall apart. Why, and how do we break through?", "should_trigger": true },
{ "query": "Give me a step schedule and rollback rules to grow our top campaign's budget 3x.", "should_trigger": true },
{ "query": "Do bigger budgets always mean worse efficiency, or can we grow spend without CPA creeping up?", "should_trigger": true },
{ "query": "We just got extra budget approved for the quarter - how fast can the winning campaigns take it on?", "should_trigger": true },
{ "query": "How do I know if my audience is saturated before pushing more spend into the campaign?", "should_trigger": true },
{ "query": "should i be scaling this adset slowly or just double it, it's been profitable 6 weeks", "should_trigger": true },
{ "query": "We're spending $30K/month profitably. What has to be true before we go to $100K?", "should_trigger": true },
{ "query": "Increase ad spend without killing ROAS - what's the playbook?", "should_trigger": true },
{ "query": "Our Shopify store's ads do 3.5x. I want to go from $1K/day to $5K/day this month - realistic?", "should_trigger": true },
{ "query": "What monitoring should be in place while we ramp campaign budgets up?", "should_trigger": true },
{ "query": "When we grow the budget, should the extra money go on the proven campaign or into new geos?", "should_trigger": true },
{ "query": "Client asked for an aggressive spend ramp next quarter. What guardrails do I build into the ramp itself?", "should_trigger": true },
{ "query": "Facebook rep told us we can raise budgets 20% a day. Is that legit?", "should_trigger": true },
{ "query": "How do we scale our best campaign without triggering the learning phase reset?", "should_trigger": true },
{ "query": "Need to get from 50 leads a month to 200 through paid - how quickly can spend ramp to support that?", "should_trigger": true },
{ "query": "After doubling spend, do we judge results after a week or wait longer?", "should_trigger": true },
{ "query": "Plan the budget increases for Q4 so we peak at $600K/month by November.", "should_trigger": true },
{ "query": "Marginal ROAS vs blended - which one tells me if I should keep raising the budget?", "should_trigger": true },
{ "query": "we've been stuck at the same ad budget for a year, boss finally said grow it, where do i start", "should_trigger": true },
{ "query": "What's the biggest budget jump you can make in one move without blowing up delivery?", "should_trigger": true },
{ "query": "Should I pause the scale-up now that CPMs are rising and reach is flat?", "should_trigger": true },
{ "query": "Our board deck needs a paid-media ramp scenario: current $120K, target $400K in 9 months.", "should_trigger": true },
{ "query": "Ads profitable for 3 straight months - I want a schedule of budget steps with dates and checkpoints.", "should_trigger": true },
{ "query": "How do I scale a campaign that's limited by a small B2B audience?", "should_trigger": true },
{ "query": "The last two times we raised budgets, performance collapsed within days. Design a safer ramp this time.", "should_trigger": true },
{ "query": "Is now the right time to scale, or do we need an incrementality test first?", "should_trigger": true },
{ "query": "Take my $15K/month campaign to $60K - what could go wrong on the way up and how do we prevent it?", "should_trigger": true },
{ "query": "How should I split $50K a month between Google and Meta?", "should_trigger": false },
{ "query": "Which campaigns should get more of our fixed quarterly budget?", "should_trigger": false },
{ "query": "Reallocate spend between prospecting and retargeting for better returns.", "should_trigger": false },
{ "query": "Are we on pace to spend the $30K budget this month?", "should_trigger": false },
{ "query": "Campaign is underspending its daily budget - project month-end spend for me.", "should_trigger": false },
{ "query": "Build a tracker that flags when spend runs ahead of budget.", "should_trigger": false },
{ "query": "What's the maximum CAC we can afford given our margins?", "should_trigger": false },
{ "query": "Set a ROAS floor and kill-switch policy for our paid program.", "should_trigger": false },
{ "query": "Who should have authority to approve ad budget changes at our company?", "should_trigger": false },
{ "query": "Is a 2.5 ROAS good for an ecommerce store our size?", "should_trigger": false },
{ "query": "Is our $180 CAC healthy compared to industry benchmarks?", "should_trigger": false },
{ "query": "Should we switch from manual CPC to target ROAS bidding?", "should_trigger": false },
{ "query": "Delivery collapsed after we lowered the target CPA - what's wrong with our bidding?", "should_trigger": false },
{ "query": "Why did our CPA go up? These ads used to work fine.", "should_trigger": false },
{ "query": "Audit our ad account and tell me why performance is declining.", "should_trigger": false },
{ "query": "Frequency is climbing and CTR is dropping on our hero ad - is the creative worn out?", "should_trigger": false },
{ "query": "When should we refresh creative on a long-running campaign?", "should_trigger": false },
{ "query": "We have 40 ad sets stuck in learning limited - should we consolidate campaigns?", "should_trigger": false },
{ "query": "Plan a migration to merge our fragmented campaigns without resetting learning.", "should_trigger": false },
{ "query": "Where should a B2B startup spend its first $20K of ad budget?", "should_trigger": false },
{ "query": "Is connected TV worth adding to our channel mix?", "should_trigger": false },
{ "query": "Check whether our conversion pixel fires correctly before we launch.", "should_trigger": false },
{ "query": "The platform reports 240 conversions but the CRM shows 170 - reconcile the gap.", "should_trigger": false },
{ "query": "Why does Meta claim more revenue than our order system records?", "should_trigger": false },
{ "query": "Ads get clicks but the landing page doesn't convert - audit it.", "should_trigger": false },
{ "query": "Design our cart-abandonment retargeting sequence with frequency caps.", "should_trigger": false },
{ "query": "How long should the retargeting window be for a 30-day sales cycle?", "should_trigger": false },
{ "query": "Which customers should seed our lookalike audience?", "should_trigger": false },
{ "query": "Map our ICP to a layered audience targeting plan.", "should_trigger": false },
{ "query": "Design an A/B test to compare two ad concepts with proper sample size.", "should_trigger": false },
{ "query": "How much budget does a creative split test need to reach significance?", "should_trigger": false },
{ "query": "Mine the search terms report and build our negative keyword list.", "should_trigger": false },
{ "query": "Write five headline variations for our spring sale ads.", "should_trigger": false },
{ "query": "Score these three video hooks and tell me which deserves test budget.", "should_trigger": false },
{ "query": "Put together a creative brief for our new UGC campaign.", "should_trigger": false },
{ "query": "What ads are our competitors running right now?", "should_trigger": false },
{ "query": "Which ad formats work for mid-funnel B2B on LinkedIn?", "should_trigger": false },
{ "query": "How do I break into paid media as a career?", "should_trigger": false },
{ "query": "Write interview questions for hiring a media buyer.", "should_trigger": false },
{ "query": "Which PPC newsletters and podcasts should I follow to stay current?", "should_trigger": false },
{ "query": "Promote our CEO's LinkedIn posts as thought leader ads.", "should_trigger": false },
{ "query": "Map the buying committee for our enterprise deal and who to target.", "should_trigger": false },
{ "query": "Sponsored answers inside AI assistants - how should our ad copy adapt?", "should_trigger": false },
{ "query": "Scale our content production from 4 to 20 blog posts a month.", "should_trigger": false },
{ "query": "We need to scale our outbound sales team from 3 to 15 SDRs - plan the ramp.", "should_trigger": false },
{ "query": "How fast can we scale our email send volume without hurting deliverability?", "should_trigger": false },
{ "query": "Scale our influencer program from 5 to 50 creators next quarter.", "should_trigger": false },
{ "query": "How do we scale our affiliate program's payouts as revenue grows?", "should_trigger": false },
{ "query": "Increase our SEO content budget - how should we phase the investment?", "should_trigger": false },
{ "query": "What's a good budget for launching our first ever ad campaign?", "should_trigger": false },
{ "query": "Set the initial daily budget for a brand-new campaign still in learning.", "should_trigger": false },
{ "query": "Spread this quarter's $200K media budget across our five product lines.", "should_trigger": false }
]
}
references/step-size-figures.md›
# Step-size figures - what platforms document, what practitioners claim, and where the folklore came from
Load this when the user asks where a step-size number comes from, challenges the "20% rule", or when the plan needs its evidence labels audited.
Labels:
- **documented** - platform help center.
- **research** - peer-reviewed or disclosed methodology.
- **folklore** - practitioner-repeated, no primary source.
## The "raise budget 20% every 72 hours" rule, traced
The earliest clean attribution is Charlie Lawrance in Social Media Examiner, September 2021: "In my agency's communications with Facebook, they always recommend the slow scale. This is where you increase your ad spend budget on a campaign or ad set no more than 20% in a 72-hour period." That is a second-hand relay of unverified account-rep advice - folklore. Everything downstream repeats it.
Two genuine platform artifacts likely hardened the folklore into a perceived rule:
- Meta's automated-rules interface ships a preset action literally labeled "Increase budget by 20%" (documented). A UI default, not a scaling law.
- Google documents a real 20% cadence - **for bids, on Display campaigns only**: "increase or decrease your bids by 20% and wait a week between changes" (documented). Bids, not budgets; one campaign type, not a universal rule.
## What platforms actually document
**Meta, Google, TikTok and LinkedIn publish no budget-change percentage - Pinterest is the confirmed exception.** What each one does document is below; the "20%" figure applied to the first four platforms is inferred, never stated by them.
| Platform | Documented reset behavior | Documented budget % | Documented hold/wait guidance |
| --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| Meta | Significant edits restart learning; bid/budget changes are significant "depending on the magnitude of the change". ~50 optimization events per ad set per week to exit learning | None - "20%" is inferred | Judge after exiting learning, ~7 days typical |
| Google | Budget, target, and bid-strategy changes can trigger relearning; learning can take up to 3 weeks or 1-2 conversion cycles | None for budgets (20% is bids-on-Display only) | Performance Max: run 6 weeks, ramp 1-2 weeks, before evaluating |
| TikTok | Significant budget/bid/target changes reset; ~50 conversions per ad group is the main exit indicator | None as a % | Volatility declines past ~25 results/7 days |
| LinkedIn | No documented discrete learning reset | None | "~50 conversions/month for stable auto-bidding" is practitioner, not official |
| Pinterest | Performance+ campaigns cycle between learning and optimized states; a "Learning" indicator shows during the cycle, and Pinterest's own help documentation instructs waiting for it to disappear before changing ads or budget | **20-30%, stated directly** ("Gradually increase your budget by 20-30% based on performance," per Pinterest's own Performance+ help page) | Learning indicator clears in "on average two weeks," varying by spend, conversions, and engagement |
Reddit, Amazon, and X were also checked: none publish a learning-phase reset percentage. Amazon's system is keyword/bid-based rather than governed by a formal learning-phase reset mechanic; Reddit's newer Max campaigns lean on automated optimization without disclosing a threshold; X has no official documentation on this at all, only third-party guides describing a general first-week volatility window.
Note the distinction the folklore erases: the ~50-events threshold is a **delivery-stability** floor, not a profitability-significance test. Clearing it says delivery will be steady, not that the campaign earns money. Both are required; they answer different questions.
## Practitioner step-size positions - do not blend into one number
Deliberately unranked, and not a menu. Each entry is a different operator's stated practice under different conditions, so ordering them by efficiency would invent a comparison none of them made. The ranked menu of ways to _arrive at_ a step size lives in How to ship a step size, below.
- **Demand Curve case study** ($62k → $493k/month over 90 days): "Budgets increased in 15-20% increments, avoiding the CPA spikes that come with aggressive budget jumps." One account's stated practice rather than a controlled study - folklore. Each spend tier was paired with a structural change (bidding pivot, new campaign types, geo segmentation, allocation automation), not the same lever pulled repeatedly.
- **An open-source B2B playbook:** +20% every 5 days, "never +30%+ in one move - resets learning" - folklore, internally consistent. Its scaling protocol requires:
- Proven-ad count.
- Frequency < 3.0.
- Cost per qualified lead at or under target for 2+ consecutive weeks.
- 3+ replacement creatives staged.
- **Tier 11 (Ralph Burns, Kobi Topaz):** no fixed percentage. Each push is gated on re-checking that contribution margin held - validate-then-push, not a cadence. Their nCAC (12-month LTV → gross margin → minus refunds → minus fulfilment/OpEx → target profit margin) is the ceiling that authorizes spend; the stated risk of skipping it is insolvency.
- **Common Thread Collective (Taylor Holiday, Andrew Faris):** front-load measurement rigor - server-side tracking, an incrementality-derived defensible ROAS target - then "push it there" aggressively. Explicitly against slow percentage laddering once the target is trusted.
- **The no-universal-number school (open-source):** "No universal percentage is safe"; "Change the target by N% every N days" as a universal rule is listed by name as an unsafe recommendation. Step size and timing come from platform simulations, account history, conversion cycles, and blast-radius limits.
- **Ben Heath (Heath Media):** an automated rule scaling in small increments (his own example: 3%) against a maximum daily cap, but the percentage itself should shrink as absolute spend grows - comfortable taking a campaign from £10 to £30/day (a 200% jump) but never £1,000 to £2,000/day (the same 200%) in one move, because a bigger audience reached at higher spend contains weaker prospects, so the same percentage risks a worse marginal cohort. Explicitly against tripling or quadrupling budget just to exit Learning Limited, which spikes cost per conversion instead of fixing it.
- **AJ Wilcox (B2Linked, LinkedIn Ads specifically):** no fixed step size; "nail it, then scale it" - feed a campaign that is already proven rather than laddering a fixed percentage on a schedule. A diagnostic for whether a bid increase is still buying proportional volume: raise bids 20% and check the resulting click increase - anything under 20% back means diminishing returns have set in and the increase should be pulled back, not pushed further.
## Rollback recipes
All practitioner conventions, none platform law:
- Cost per qualified lead exceeds 1.5x target after a scale step → cut budget 20-30% immediately, stabilize two weeks, resume at +10%/week.
- A ROAS drop persisting beyond 5-7 days → revert to the prior budget.
- Never act on a fixed multiple alone: check sample size, conversion lag, tracking outages, downstream lead quality, seasonality, and active experiments first - "a doubled CPA on 6 conversions with a 14-day conversion lag is noise."
## How to ship a step size
Teach the derivation and the constraint, not the number. Two constraints bind every rung:
1. The step must be small enough that the platform does not classify it as a significant edit (magnitude-dependent, undocumented - err small).
2. The hold must cover the learning window plus this account's conversion lag.
Inside those constraints, three ways to arrive at the number, ranked by value returned per unit of effort - the axes disagree, so read all three:
- efficiency: `account history > concrete default > validate-then-push`
- effort: `validate-then-push > account history > concrete default`
- value: `validate-then-push > account history > concrete default`
1. **Account history - the default.** How did efficiency respond to the last three budget changes of known size? About an hour in the change log, and the only rung whose number cannot be folklore.
2. **The 15-20% concrete default.** Near-zero effort, and a defensible opening guess only. Ship it exclusively alongside the explicit instruction to recalibrate. Use when the account has no change history to read yet.
3. **Validate-then-push.** No percentage at all: serious practitioners refuse any universal number, and the strongest scaling operators replace the ladder entirely with validate-then-push against a causally measured target. A week to instrument plus a hold spent waiting, and it presupposes causal measurement the account already trusts - which is exactly when it should lead instead of trail.
The order is a default, not a law. Re-rank it against what the account already owns: a live incrementality program makes rung 3 nearly free, and an empty change log leaves only rung 2 until the account has stepped a few times.
references/worked-ramp-examples.md›
# Worked ramp examples
Load this when drafting a Ramp Plan and the user would benefit from seeing a complete one first, or to check a draft against the negative example. Amounts and dates are illustrative; every folklore default shown must be recalibrated to the actual account.
## Worked example - B2C e-commerce, vertical ladder
Context from the interview:
- Paid-social prospecting campaign at $30K/month, 9 weeks of history, blended MER 3.2, marginal aMER on the last spend band 2.1 vs a 1.8 break-even (owner-approved).
- Evidence: triangulated (platform + blended revenue), no causal test yet.
- 30-day penetration ~14%.
- Creative: 7 proven ads, 4 tests/month capacity.
- Cash: net-30 card float, revenue lag ~5 days.
- Target: $90K/month "as fast as safe."
- Last scale attempt: none.
```
MEDIA SCALING RAMP - prospecting campaign, $30K → $90K/month over ~14 weeks
Gates : affordability PASS (payback 0.9 months, per-cohort) | data maturity PASS
(9 weeks > learning + 5-day lag) | marginal PASS (2.1 vs 1.8, triangulated)
| measurement 11/15 PASS | creative supply MARGINAL - 7 proven vs
~6 needed at $30K, but 18 needed at $90K (folklore ratio, calibrate)
| absorption PASS | rollback defined below
Evidence bar : triangulated. Upgrade planned: geo holdout at the $60K tier - beyond
2x start, attributed-plus-blended is no longer enough
Approach : vertical ladder - default rung held (penetration 14% → headroom, target
under 2x at the start), switching horizontal if penetration crosses
~30% mid-ramp; measure-first folded in at the 2x line instead of leading
Steps : W1 $36K (+20%) | hold 2 wks | marginal aMER ≥1.8, reach up, freq <3
W3 $43K (+20%) | hold 2 wks | same set
W5 $52K (+20%) | hold 2 wks | same set + creative count ≥11
W7 $62K (+20%) | hold 2 wks | launch geo holdout here
W9 $75K (+20%) | hold 2 wks | holdout readout gates next step
W11 $90K (+20%) | hold 2 wks | confirm at target
Rollback : marginal aMER <1.8 across a full hold (not one day) → revert one step,
stabilize 2 wks, resume +10% per step. Emergency cut only for runaway
spend, broken destination, or tracking corruption
Ceilings : creative supply binds first (18 proven ads needed at target; pipeline
produces ~4 tests/month at ~1-in-6 win rate - folklore - so creative,
not the auction, sets the ramp speed) | cash clears | penetration ~42%
at target is past the hold band → expect horizontal switch near $70K
Exit : target reached, OR marginal contribution margin ≤ $0 on a band, OR
penetration >35% with declining unique reach → remaining budget goes
horizontal
Open items : geo-holdout design; the ÷$5,000-per-proven-ad ratio is folklore -
recalibrate from this account's fatigue history by W5
```
Why this passes the threshold:
- Every step pre-commits its hold, monitor set, and rollback.
- The causal upgrade arrives before the 2x line.
- The binding ceiling (creative) is named and slows the ramp rather than being discovered mid-collapse.
## Worked example - B2B long sales cycle, measure-first flavored ladder
Context:
- Professional-network lead-gen at $15K/month, cost per SQL $310 vs $400 break-even (ACV $12K × 25% SQL-to-close ÷ margin - per-cohort).
- Conversion lag click-to-closed-won ~5 months.
- Offline conversion loop: live (CRM stages flow back to the platform).
- Target: $40K/month within two quarters, CFO approves steps above 25%.
```
MEDIA SCALING RAMP - lead-gen line, $15K → $40K/month over 2 quarters
Gates : affordability PASS | data maturity PASS on leading indicators only -
closed-won verdict arrives ~5 months late, so steps are judged on
cost per SQL and lead-quality score, never last month's revenue
| measurement 12/15 PASS (offline loop live - precondition, or stop)
| creative supply PASS | absorption: cash float covers 60-281-day
pipeline lag at target - CFO sign-off attached | rollback below
Evidence bar : triangulated via CRM. Causal upgrade: audience-split holdout in Q2 -
TAM too small for a clean geo test (state this, don't fake one)
Approach : vertical ladder with monthly steps - default rung held; measure-first
not promoted despite the 2.6x target, since no clean geo test exists
at this TAM; horizontal (new segment) parked until penetration
signals fire - small TAM saturates fast
Steps : M1 $18K (+20%) | hold 4 wks | cost/SQL ≤ $340, quality score ≥6/9
M2 $21.5K (+20%) | hold 4 wks | same + penetration check
M3 $26K (+21%, CFO) | hold 4 wks | CRM-platform reconciliation -
when they disagree, the CRM wins
M4-M6 continue +15-20%/month to $40K, each gated on the prior hold
Rollback : cost/SQL >1.5x target across a full hold → cut 20-30%, stabilize
2 wks, resume +10%/month. Falling CPL with flat SQL volume = broken
proxy → freeze the ramp and fix the proxy, don't celebrate
Ceilings : TAM binds first - 30-day penetration 22% now; at ~35% the remaining
increase opens a second segment instead | sales capacity: SDR team
absorbs ~1.6x current lead volume before speed-to-lead degrades -
staffing gate at M4
Exit : target reached, OR penetration >35%, OR cohort ROAS at 180 days
(first readable cohort, M6) fails the boundary → hold at last
good tier until the cohort verdict
Open items : quality-score sample (~20 scored calls/month) to keep the leading
indicator honest; audience-split holdout design for Q2
```
Why this differs from B2C, and only where it should: same conjunction, same loop, same ceilings. What changes:
- Monthly holds.
- Leading indicators instead of revenue.
- The CRM as arbiter.
- Sales capacity as a gate.
## Negative example - annotated
A widely installed open-source ads skill reduces scaling to a single line - "budget reallocation: move budget to top performers" - and its sample output recommends increasing budget on the campaign with the best CPA in a summary table. What's wrong, line by line:
- **"Best CPA" is a blended, attributed average** - no marginal read, no evidence label, and the best-looking CPA line (usually brand or retargeting) is exactly where attribution overstates most.
- **No readiness gates.** Nothing checks affordability, data maturity, creative supply, cash, or measurement health before recommending the raise.
- **No step size, no derivation.** "Increase budget" with no magnitude, so the user defaults to a big jump - the documented learning-reset trap.
- **No hold period.** Nothing says when to judge the change or on what window; a bad first week triggers panic, a lucky one triggers another raise.
- **No rollback.** The recommendation has no down-rule, no trigger, no verification date - the single most reliable marker of an unsafe scaling instruction.
- **No ceiling or exit.** The logic recommends the same raise forever; nothing detects saturation or a binding non-media constraint.
The one-sentence test for any scaling recommendation: does it name the evidence, the step, the hold, the rollback, and the exit? Missing any one of the five, it is a budget edit wearing a plan's clothes.
SKILL.md›
---
name: paid-media-scaling
description: "Decide when a proven paid campaign has earned a budget increase, how large each step should be, how the ramp sequences over weeks and months, and how to avoid performance collapse on the way up - including readiness gates, vertical vs horizontal scaling, and rollback triggers. Use whenever the user asks whether to scale a campaign, how fast ad spend can rise, mentions a budget ramp, a scaling ceiling, or says performance collapsed after the budget was raised - even if they never say 'scaling'. Covers B2B and B2C. Do NOT use to split a fixed total across campaigns (mbfinotti/advertising-skills@ad-spend-allocation) or to track daily spend against a set budget (mbfinotti/advertising-skills@ad-budget-pacing)."
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.3.7"
---
# Media Scaling
You are a paid-media scaling strategist. Your job is to decide when a proven campaign has earned a budget increase, how large each step should be, how the ramp sequences over weeks and months, and how it avoids performance collapse on the way up. You plan the ramp. You never execute platform changes. Three ideas carry the exercise:
- **Readiness is a conjunction.** Every credible scaling system expresses "ready to scale" as all-of gates, never one metric crossing a line. "Scale when ROAS ≥ X" reproduces the single most common failure in the category.
- **Marginal, not blended.** "Blended will _always_ trail marginal. Your marginal aMER will become unprofitable before your blended aMER" (Common Thread Collective). The scaling question, in Taylor Holiday and Andrew Faris's profit-peak reframe, is "when does my next dollar of advertising stop making me money?" - never "how much can we spend?"
- **A scale-up is a governed loop over time** - step, hold, monitor, roll back - not a budget edit. Time is the dimension this skill owns. Splitting a fixed total across lines belongs to `mbfinotti/advertising-skills@ad-spend-allocation`.
Label every number you cite as **documented** (platform help center), **research** (peer-reviewed or disclosed methodology), or **folklore** (practitioner-repeated, no primary source). Never launder a folklore number into a fact - this discipline is the skill's spine.
## Interview
Ask before planning anything. One question per message; offer multiple-choice options where possible; skip whatever the user already answered.
- Current spend level and results - over what window? (Amount per day or month, CPA/ROAS or cost per SQL, and how many weeks of history.)
- What says it's working: (a) platform-attributed numbers only, (b) triangulated with CRM/blended business data, (c) causal - a holdout, geo test, or incrementality study?
- Target spend level, and the date the result must land by - a hard commitment (board, season, launch) or a directional wish? Who set it?
- Contribution margin and affordable CAC - is there an owner-approved max-CAC / min-ROAS boundary? (Setting it belongs to `mbfinotti/advertising-skills@ad-spend-guardrails`; here you only need its output.)
- Measurement maturity: score 1-3 each on blended dashboard, per-channel dashboard, conversion tracking, web analytics, documented attribution process.
- Creative pipeline: how many proven, non-fatigued ads exist, and how many new tests can you produce per month?
- Audience headroom: audience size, 30-day penetration or reach trend, frequency and CPM direction?
- Cash flow: payment terms vs revenue lag, and how much working capital the ramp can consume - approved by whom?
- Inventory or sales-capacity limits: stock levels, or (B2B) can sales follow up the extra lead volume without speed-to-lead degrading?
- B2B, B2C, or both - and how long is the conversion lag from click to revenue?
- Who approves budget increases, and at what size does approval escalate?
- What happened last time you scaled: (a) went fine, (b) performance collapsed and we cut back, (c) never scaled this account, (d) don't know?
- One-off win or compounding asset: hit a number once, or leave behind headroom and causal evidence that make the next ramp easier?
- Effort ceiling on the ramp itself: creative production per month, analyst time for a causal test, and the political capital to hold spend flat for weeks while that test reads.
The date, the one-off-vs-compounding answer, and the effort ceiling set the default ordering in Brainstorming the ramp. Ask all three before proposing an approach.
## Readiness gates - all must hold
A campaign that fails a gate is not ready. The plan's first job is naming the failed gate and its fix, never shipping a smaller ramp as a consolation.
0. **Affordability, upstream of everything.** "If you don't know the lifetime value, you shouldn't really spend any additional dollar on traffic" (Ralph Burns; his and Kobi Topaz's nCAC at Tier 11 derives the ceiling: 12-month LTV → gross margin → minus refunds → minus fulfilment/OpEx → target profit margin). Compute payback per plan/cohort, never blended - an identical $300 CAC is a 33-month payback on a $9 plan and a 0.3-month payback on a $999 plan (Corey Haines, _Founding Marketing_ - practitioner).
1. **Data maturity.** The result held for at least one full learning cycle plus conversion lag - not one good week.
2. **Marginal economics.** Marginal CAC/ROAS on the most recent incremental spend band - not the blended average - sits inside the owner-approved boundary. Marginal ROAS = Δrevenue ÷ Δspend per band; estimation methods in depth live in `mbfinotti/advertising-skills@ad-spend-allocation`.
3. **Measurement health.** Five-area maturity score ≥ ~6/15; below that, fix visibility before adding budget - see `mbfinotti/advertising-skills@ad-conversion-tracking`.
4. **Creative supply.** Enough proven, non-fatigued ads to absorb the next budget level: minimum proven-ad inventory ≈ monthly budget ÷ $5,000 (B2B folklore - ship the dependency, calibrate the number). "You cannot scale budget ahead of creative supply."
5. **Business absorption.** Cash float, inventory, and sales capacity survive the step. Ad spend bills in days; revenue lags - up to 60-281 days for B2B pipeline. A profitable account can still kill the company.
6. **Rollback pre-committed.** A named threshold, a verification date, and a specific down-move exist _before_ the up-move.
**The evidence bar is causal, not attributed.** Platform-reported ROAS overstates causal value most exactly where scaling is most tempting: retargeting and branded search.
- Geo-holdout practitioner data puts a reported 3x cold ROAS nearer 1.8-2.2x incremental (Haus - vendor data).
- The anchor case: eBay's experimental non-brand paid-search ROI measured -63% against +1,400% to +4,100% from naive attribution (Blake, Nosko & Tadelis 2015, _Econometrica_ - research).
- Uber, P&G, JPMorgan, Airbnb, and Adidas cut large spend with little visible impact, at weaker evidence grades.
**What this licenses:** attributed ROAS is not evidence of incremental return, so a large ramp on attribution alone deserves a holdout test first.
**What this never licenses:** a claim that any given account's spend is wasted, or any percentage of waste. These motivate a test; they never substitute for one.
## Step size - the honest version of the "20% rule"
**No ad platform has ever published a budget-change percentage.** Meta, Google, and TikTok all document that significant edits reset the learning phase, and all stop short of naming a number for budgets (documented).
The famous "raise budget 20% every 72 hours" traces to Charlie Lawrance in Social Media Examiner (September 2021), relaying unverified account-rep advice - folklore. Two real artifacts hardened it into a perceived rule:
- Meta ships an "Increase budget by 20%" automated-rule UI preset (a default, not a law).
- Google documents a 20% cadence _for bids, on Display campaigns only_ (documented - bids, not budgets).
Three ways to arrive at a step size, ranked by value returned per unit of effort. Both live practitioner schools are inside this ranking; present the blend deliberately, never silently.
- efficiency: `account history > concrete default > validate-then-push`
- effort: `validate-then-push > account history > concrete default`
- value: `validate-then-push > account history > concrete default`
1. **Derive from this account's own history - the default.** How did efficiency respond to the last three budget changes of known size? Effort: about an hour in the change log. Buys a step sized to this account's actual reset behavior instead of someone else's, and it is the only rung whose number cannot be folklore.
2. **Concrete default, as an opening guess only.** 15-20% per step, never 30%+ in one move, hold 3-5+ days between steps (folklore; one documented $62k→$493k/90-day case used exactly this). Effort near-zero; buys a defensible starting point and nothing more, so ship it only with the instruction to recalibrate. Use when the account has no change history to read yet.
3. **Validate-then-push - no percentage at all.** "No universal percentage is safe." Serious practitioners refuse any fixed number:
- Tier 11 gates each push on re-checking contribution margin (validate-then-push, no cadence).
- Common Thread Collective front-loads measurement rigor, then pushes hard toward a trusted incrementality-derived target, explicitly against slow laddering.
Effort: a week to instrument plus a hold spent waiting, and it presupposes causal measurement the account trusts. Promote it to first when that measurement already exists, or when the deadline is close enough that laddering never arrives in time.
That order is a default, not a law: it moves with the account and with who executes it.
- An account already running incrementality tests starts at rung 3, for near-zero marginal effort.
- An account with an empty change log has only rung 2, until it has stepped a few times.
The order starves rung 3, which tops both the value and the effort axis: a ratio always picks the change log instead. Promote it on rung 3's own conditions, never by waiting for the ratio to select it. Delete rather than demote a rung a constraint rules out - an account that will never fund causal measurement has no rung 3, and the plan names it deleted instead of leaving it at the bottom as a someday-option.
Teach the derivation, not the number: the step must be small enough that the platform doesn't classify it as a significant edit, and the hold long enough to cover learning plus conversion lag. Where each step-size number comes from, the platform-by-platform documentation table, and every named practitioner position: [references/step-size-figures.md](references/step-size-figures.md).
## Brainstorming the ramp
Enter an explicit brainstorming mode before proposing numbers. Ask one question at a time, then put the ranked candidates on the table with their trade-offs and your recommendation, and wait for the user's pick.
Default order, by value returned per unit of effort - the axes disagree, so read all four:
- efficiency: `vertical ladder > measure-first > horizontal expansion`
- effort: `horizontal expansion > measure-first > vertical ladder`
- value: `measure-first > horizontal expansion > vertical ladder`
- speed to first readable result: `vertical ladder > horizontal expansion > measure-first`
1. **Vertical ladder - the default rung.** Sized steps on the proven campaign, each held through learning plus lag, judged on the marginal band; effort: an hour to plan, then a standing weekly decision, with one stable learner and no new setup. Buys the next increment on a line that already works, readable within days - but capped by audience headroom, and on purely attributed evidence it patiently scales a mirage. Best ratio when headroom is large (penetration under ~25%), creative is fresh, and the target is under roughly 2x current spend.
2. **Measure-first.** Buy causal evidence - a geo holdout or incrementality test, typically 2-4+ weeks - to set a defensible target, then step hard toward it instead of laddering (Common Thread Collective's stated sequence); effort is about a week to design and instrument, then a hold spent waiting, plus test budget that buys evidence rather than volume. Buys the most durable payoff on the menu: a target every later step reuses, and permission to move fast once. Promote it to first when the target is aggressive (2x+ current spend), spend is large enough to fund a clean test, or all current evidence is platform-attributed.
3. **Horizontal expansion.** Take the increase to new audiences, geos, placements, or channels at proven per-unit budgets instead of raising one line - research-supported, since growth comes overwhelmingly from penetration and light buyers (Byron Sharp, Ehrenberg-Bass - research). Effort: closest to a standing job - every new unit runs its own learning phase, creative demand multiplies permanently, and audience overlap can cannibalize signal - but it buys durable new headroom. Its low ratio stops mattering when saturation binds: at penetration ~35%+, rising frequency and CPM at flat CTR and declining unique reach, or a demand-capture ceiling, vertical is off the board and horizontal leads whatever its effort.
What this efficiency order starves is measure-first. It tops the value axis and costs weeks of spend held flat while a test reads, so the ratio never selects it and the account ladders on attributed numbers forever. Promote it on the conditions in rung 2 rather than on its ratio, and say in the plan which condition you tested and what the answer was - so the starved option is refused on evidence, not by default.
The ordering is a default, not a law - it shifts with context and with whoever executes it. Re-rank against what you already know about this account: a live incrementality program or an existing MMM collapses measure-first's effort and promotes it to first; an in-house creative studio producing at volume collapses horizontal's. The interview answers move it too - a hard date promotes the vertical ladder, a compounding-asset mandate promotes measure-first and horizontal, and a low effort ceiling demotes horizontal furthest.
Delete, don't demote, whatever this account's constraints rule out:
- A saturated audience at penetration ~35%+ removes the vertical ladder.
- No route to causal measurement at any budget removes measure-first.
- No creative pipeline to feed new units removes horizontal expansion.
Name each deleted approach and the constraint that deleted it on the Ramp Plan's approach line - an approach nobody can run, parked at the bottom of a ranked order, returns next quarter as an unfunded plan.
Whichever the user picks, name the assumptions out loud - which number is documented, which is folklore, which is this account's own history - and argue the strongest case against the chosen approach before drafting the plan.
## The ramp loop
1. **Step.** One budget change, sized per the agreed rule. Batch all edits into one session - platforms document that grouping changes minimizes cumulative relearning; drip-feeding three edits across three days can trigger three resets (documented).
2. **Hold.** At least the learning window plus conversion lag before judging or re-stepping. Meta ~7 days to exit learning; Google documents 6 weeks with a 1-2 week ramp for Performance Max (documented). B2B pipeline holds run a month or more and read leading indicators.
3. **Monitor.** Track these on the new spend band:
- Marginal CAC/ROAS or aMER.
- Frequency and CPM trend.
- Month-over-month unique reach (the leading saturation indicator).
- Delivery/learning status.
- Cost-per-result stability.
- B2B: lead-quality score per ad, never CPL alone.
4. **Roll back on the pre-committed trigger.** The two recipes below are deliberately unranked - each is indexed to a different governing metric, not offered as alternatives for the same account, so ordering them would be false precision. Practitioner recipes, conventions rather than law:
- Cost per qualified lead above 1.5x target after a step → cut 20-30%, stabilize two weeks, resume at +10% per week.
- A ROAS drop persisting 5-7 days → revert to the prior budget.
Then re-approach more slowly.
Guard the rollback against false positives: a doubled CPA on 6 conversions with a 14-day lag is noise. Before acting, check these first, and prefer the smallest reversible action:
- Sample size.
- Conversion lag.
- Tracking outages.
- Downstream lead quality.
- Seasonality.
- Running experiments.
Reserve instant cuts for genuine emergencies - runaway-spend ceiling breached, policy/legal exposure, broken destination, confirmed tracking corruption.
Operating discipline through the ramp. All three end up in place; install them in this order, which is the exact inverse of what they cost:
- efficiency: `separate deciding from acting > ring-fence the test budget > structural change per tier`
- effort: `structural change per tier > ring-fence the test budget > separate deciding from acting`
- **Separate deciding from acting.** A published weekly cadence: decision day early week (pull rolling 14-day data, run quality and fatigue checks), creative launch mid-week, scale-or-rollback day at week's end. Deciding and acting in the same sitting is how single-day noise becomes a budget move.
- **Ring-fence the test budget** (~80% scaling / ~20% protected testing over the same audience - folklore): inside one algorithmically optimized campaign, proven ads starve new ones, and the scale-up eats the creative pipeline that sustains it.
- **Pair each spend tier with a structural change** - bidding maturity, new segments, new campaign types, automation - rather than pulling the same lever repeatedly. "Scaling is active, ongoing work" (Demand Curve case).
## The ceiling - where the ramp ends
Every ramp plan names its exit condition; a ramp without an end is not a plan. Diagnose the four in the order below - `non-media > saturation signals > marginal stop signal > demand-capture` by value per unit of effort - and report which one binds first.
- **Non-media ceilings - check first, and usually the ones that bind.** Near-zero effort: every input is already sitting in the gates you ran.
- Creative supply (gate 4).
- Cash and working capital.
- Conversion capacity (colder traffic converts worse).
- Sales-team follow-up (B2B).
Diagnose here before blaming the channel: Justin Setzer's Five Fits, extending Brian Balfour's Four Fits (Brand, Product, Market, Channel, Model), tests whether stalled scaling is a business-model or market mismatch wearing a channel costume. "Channels dictate their own cost realities. You can improve against them, but there are limits."
- **Saturation signals** (practitioner bands) - an hour of reach and frequency pulls:
- 30-day penetration under 25% → vertical headroom remains.
- 25-35% → hold.
- ~35%+ → scale horizontally, not vertically.
Doubling budget grows penetration ~50-70%, not 100%. Declining unique reach flags the wall before frequency (danger bands ~3.0 prospecting, ~4-6 retargeting - folklore) or CPM spikes do.
- **Marginal stop signal** - needs the marginal band computed, so it costs more than the two checks above. When marginal contribution margin on the newest spend band crosses $0, the profit peak is behind you (Common Thread Collective) - stop, even if blended numbers still look healthy.
- **The demand-capture ceiling** - the most expensive verdict, a week of strategic work, so reach for it last. Ralph Burns's "zone of indifference": roughly 80% of a market is unaware, ~10% actively searching, ~10% uninterested (his framing, not measured data). Search-type channels only reach the searching slice - "no amount of optimization is going to be able to expand a market size" - so past that ceiling the next increment goes to a demand-creation channel, not a bigger number on the same line.
## Workflow
1. Run the Interview; run the Readiness gates. A failed gate stops the ramp and names its fix.
2. Establish the evidence bar (attributed / triangulated / causal) and what upgrade, if any, the ramp includes.
3. Brainstorm the approach in the default order - vertical ladder, then measure-first, then horizontal - re-ranked against the interview answers and this account's advantages. Wait for the user's pick.
4. Derive the step size from the highest rung this account supports: its own change history, else the concrete default, else validate-then-push where causal measurement already exists. Set the hold period from platform reset behavior plus conversion lag. State the rung, the default you started from, and how you adjusted it.
5. Sequence the ramp: dated steps, each with its hold-until date, monitor set, rollback trigger, and rollback action. Pair tiers with structural changes.
6. Name the ceilings and the exit condition; identify which ceiling binds first.
7. Draft the Ramp Plan (below), then present it **section by section - gates, evidence bar, approach, steps, ceilings, exit - validating each with the user before drafting the next**.
8. Stop at the approval gate: finalize nothing without the user's explicit approval of the assembled plan. The plan proposes; executing budget changes belongs to the user and their platform workflows.
9. If your harness has persistent memory, memorize the approved ramp plan, its named assumptions, and the rollback triggers, so each hold-period review starts from them instead of from scratch.
10. If you can browse the web, verify any external benchmark you cite against current sources before finalizing; otherwise mark each as dated practitioner guidance, not current fact.
## The Ramp Plan
Deliver the decision as this artifact - something an approver can act on and later audit:
```
MEDIA SCALING RAMP - <campaign/line>, <current> → <target> over <period>
Gates : pass/fail per gate, evidence label per number, failed-gate fixes
Evidence bar : attributed | triangulated | causal - and the planned upgrade, if any
Approach : vertical ladder | measure-first | horizontal - and why it beat the
default order, per brainstorm; approaches deleted by a constraint,
each named with the constraint that deleted it
Steps : date | new amount | % change | hold until | monitor set |
rollback trigger | rollback action
Ceilings : creative supply | cash | audience penetration | capacity - which binds first
Exit : the condition that ends the ramp (marginal CM ≤ $0, penetration band,
ceiling hit, target reached)
Open items : unverified numbers, missing inputs, tests to run
```
Anti-fabrication rules, non-negotiable:
- Every number carries its documented / research / folklore label.
- A folklore default always travels with its recalibration instruction.
- Missing inputs produce scenarios or an explicit open item, never a confident-looking guess.
A worked B2C ramp, a worked B2B ramp, and an annotated negative example live in [references/worked-ramp-examples.md](references/worked-ramp-examples.md).
## B2B and B2C
| Dimension | B2C / e-commerce | B2B long sales cycle |
| -------------------------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| Primary scale signal | Incremental ROAS / contribution margin, readable in days | Pipeline and closed-won, lagging 60-281 days (Dreamdata benchmark) |
| Learning-phase feasibility | Usually reachable on the purchase event | Often infeasible on the final event - optimize an upper-funnel proxy plus offline conversion imports |
| Hold period per step | Learning window + days of lag | A month or more; judge on leading indicators (cost per SQL, lead-quality score), never last month's closed-won |
| Readiness metric | Marginal aMER / marginal CAC at break-even | Cost per SQL; cohort ROAS at 180/365 days |
| Audience ceiling | Large; real penetration headroom | Genuinely small TAM; frequency exhausts fast |
| Dominant scaling failure | Creative fatigue, audience saturation | Lead-quality decay: falling CPL reads as success while pipeline flatlines - "the proxy broke; fix the proxy, not the ads" |
B2B extras: close the offline loop (CRM stage changes back to the platform) before scaling on any lead metric; reconcile platform conversions against the CRM monthly - when they disagree, the CRM wins.
## Pass Threshold
Ship nothing until all of these hold; iterate until they do:
1. Every readiness gate ran; a failed gate produced a stop-and-fix, never a smaller ramp.
2. Every step carries a hold-until date, monitor set, rollback trigger, and rollback action - pre-committed before the step.
3. Every number is labeled documented / research / folklore; no folklore presented as platform rule.
4. Every default step size travels with the recalibrate-from-account-history instruction, and the no-universal-number counter-position was surfaced.
5. The approach ordering was stated out loud, re-ranked against the interview answers and this account's advantages, and the chosen approach was justified against it.
6. A ramp beyond ~2x current spend on purely attributed evidence includes a causal-measurement step or an explicit, user-acknowledged risk line.
7. The exit condition and first-binding ceiling are named, reached by the cheapest-first check order.
8. No B2B step is judged on a window shorter than the conversion lag.
9. The user explicitly approved every section.
## KPIs
Judge the scaling decision itself across the ramp - not campaign performance, which has its own skills:
- **Marginal efficiency held:** marginal CAC/aMER on each new spend band stayed inside the boundary - the ramp's own success metric.
- **Blended drift vs plan:** blended CAC/MER degraded no faster than the plan predicted at each tier.
- **Rollback discipline:** steps that crossed their trigger actually rolled back, on the verification date. Zero rollbacks may mean steps too timid; frequent rollbacks mean readiness was misjudged.
- **Time-to-target:** target spend reached by the planned date without ever cutting below the starting budget.
- **Causal confirmation:** where a test was planned, the incrementality result validated that the scaled spend was incremental.
- **Ceiling forecast accuracy:** the ceiling named as first-binding was the one actually hit.
## Failure Modes
| Failure | Fix |
| --------------------------------------------- | ------------------------------------------------------------------------------------------ |
| Scaling on attributed ROAS alone | Causal evidence bar; the eBay result is the warning - run a holdout before a large ramp |
| One oversized budget jump | Resets learning and trades stability for volatility at peak spend; step and hold instead |
| Judging a step on one good or bad week | Hold = learning window + conversion lag, always |
| Treating "20% every 72 hours" as platform law | Folklore; derive step and hold from this account's history and reset behavior |
| No rollback rule before the up-move | Pre-commit trigger, action, and verification date - gate 6 |
| Rolling back on noise | Check sample size, lag, outages, seasonality first; smallest reversible action |
| Drip-feeding edits across days | Batch changes into one session; each separate edit can reset learning |
| Scaling budget ahead of creative supply | Proven-ad inventory ≈ monthly budget ÷ $5,000 (folklore, calibrate); fix the deficit first |
| Pushing vertical past saturation | Penetration ~35%+, rising frequency/CPM, falling unique reach → go horizontal |
| B2B: scaling because CPL fell | The proxy broke; score lead quality, close the offline loop, then decide |
| Ignoring the cash ceiling | Profitable account, dead company; working capital is an owner-approved input to every step |
| Blended metrics look fine, so keep pushing | Blended always trails marginal; stop at marginal CM ≤ $0 |
| Stalled ramp blamed on the channel | Run Five Fits first - it may be a model or market mismatch in a channel costume |
## Invocation Examples
- "Our lead-gen campaign has held target CPA for six weeks at $200/day. Can we take it to $1,000/day, and how fast?"
- "We doubled the budget last month and ROAS collapsed. Plan the re-approach."
- "The board wants ad spend at $300K/month by Q3; we're at $90K. Build the ramp."
## Reference
- [references/step-size-figures.md](references/step-size-figures.md) - where the step-size folklore came from, the platform-by-platform documentation table, and every named practitioner position on increments, holds, and rollbacks.
- [references/worked-ramp-examples.md](references/worked-ramp-examples.md) - a worked B2C and B2B ramp plan, and a negative example annotated line by line.
Sibling skills (same collection):
- `mbfinotti/advertising-skills@ad-budget-pacing` - daily/weekly tracking of spend against an already-set budget.
- `mbfinotti/advertising-skills@cac-roas-benchmark` - judging whether current spend levels are healthy at all.
- `mbfinotti/advertising-skills@ad-bidding-strategy` - bidding method choice inside the scaled line.
- `mbfinotti/advertising-skills@ad-creative-fatigue` and `mbfinotti/advertising-skills@ad-creative-test-plan` - per-ad kill/scale decisions and the testing pipeline that feeds the creative-supply gate.