SKILL DETAIL
ad-creative-fatigue
mbfinotti/advertising-skills/ad-creative-fatigue
Decide whether a running ad creative is genuinely wearing out, or whether a confounder - budget change, learning-phase reset, audience saturation, auction CPM inflation, seasonality, tracking breakage - explains the decline, and return a verdict with confidence plus the highest-return remedy per unit of effort. Use whenever the user mentions creative or ad fatigue, wear-out, climbing frequency, dropping CTR, a rising CPA on an older ad, when to refresh creative, or whether to kill an ad - even if they never say 'fatigue'. Covers B2B and B2C across search, social, video, and native. Ends at the verdict: replacement creative belongs to mbfinotti/advertising-skills@ad-copy-variants and mbfinotti/advertising-skills@ugc-ad-scripts.
Installation
npx skills add https://github.com/mbfinotti/advertising-skills --skill ad-creative-fatigue
技能文件
SKILL.md
最近同步 · 2026年9月24日
evals/evals.json›
{
"skill_name": "ad-creative-fatigue",
"evals": [
{
"id": 1,
"prompt": "I run paid social for Lumo Home, a DTC home-fitness brand. Our hero static 'before-after-hero-v2' has been our best Meta ad for five weeks. Eleven days ago we raised the ad set budget from $600/day to $1,050/day in one step. Since then: link CTR 1.19% vs 1.38% before, CPA $44 vs $33, frequency 2.6 vs 1.8. The two other creatives in the ad set dropped the same way. Our growth lead says the hero ad is fatigued and wants me to brief two replacement UGC videos today. Can you confirm the fatigue read and help me plan the refresh?",
"expected_output": "A confounded verdict attributing the decline to the +75% budget step and its learning-phase reset, refusing the creative refresh and re-baselining at the new spend level.",
"files": [],
"expectations": [
"Runs a confounder screen before issuing any fatigue verdict",
"Marks the budget/bid change confounder as failing, tying the decline onset to the budget step from $600/day to $1,050/day eleven days ago",
"Identifies the budget step as a significant edit that reset the platform learning phase",
"Cites all three creatives in the ad set declining together as evidence against fatigue, because fatigue is creative-specific",
"Issues a verdict of confounded, not fatigued",
"Refuses to brief or commission the two replacement UGC videos",
"Recommends re-baselining the creative at the new spend level after delivery stabilises, rather than judging it against the pre-increase peak",
"States that some efficiency loss at higher spend is the normal cost of scale, not decay",
"Treats the frequency rise from 1.8 to 2.6 as a lagging, confirming signal that cannot carry the call",
"Flags that the current comparison window overlaps post-edit volatility and requires a clean window after stabilisation before any future fatigue read",
"Delivers a structured verdict block containing at least verdict, confidence, action, and re-check fields",
"Sets a re-check date one full comparison window after delivery stabilises",
"Applies no universal frequency or CTR-drop threshold as justification for a verdict"
]
},
{
"id": 2,
"prompt": "Brightfern here, we sell skincare DTC. Our Meta cold-prospecting video hit frequency 3.4 this week and every guide I have read says to kill creatives at 3.0. Link CTR is 1.24% this week vs 1.21% for the trailing month, CVR is 3.1% vs 3.2%, CPA is up about 4%. I have a replacement video edited and ready to go. Should I pause the current one tomorrow morning and swap in the new one?",
"expected_output": "A healthy verdict that rejects the universal frequency-3.0 kill rule, explains frequency's structural flaws as a signal, and leaves the ad running.",
"files": [],
"expectations": [
"Declines to call the creative fatigued",
"Rejects the universal 3.0 frequency kill threshold and states triggers derive from the creative's own trailing baseline and the account's own history",
"States frequency is a lagging, confirming signal that never carries the fatigue call by itself",
"Notes reported frequency is measured at ad or ad-set level while fatigue happens at creative level",
"Notes reported frequency is a period average, not the marginal effect of the next impression",
"Applies the two-signal rule requiring at least two signals, at least one leading, across two or more consecutive periods, and finds no decayed signal here",
"Reads the stable link CTR (1.24% vs 1.21%) and stable CVR as evidence the creative is healthy",
"Issues a healthy verdict and recommends leaving the ad alone",
"Distinguishes cold from warm exposure tolerance before reading the frequency number, citing the higher warm tolerance as an attributed practitioner heuristic rather than a rule",
"Advises against pausing or swapping the creative tomorrow; the ready replacement may enter as an additional test at most, not a substitution",
"Delivers a verdict block including a next scheduled review or re-check date",
"Does not present the frequency-3.0 rule as a platform rule or research finding"
]
},
{
"id": 3,
"prompt": "Peakform sells sports supplements. Our top Meta ad's numbers fell apart in the last ten days: CVR went from 3.2% to 1.1% while link CTR is steady around 1.4% and hook rate has not moved. The ad is seven weeks old so I assume it has just worn out. We also shipped a new checkout flow on the 2nd, and our 20%-off launch code expired around then, but the dev team says the release was minor. I need a fresh creative angle - can you help me figure out what the new ad should say?",
"expected_output": "A confounded verdict identifying a downstream funnel/offer problem (healthy CTR, collapsed CVR), with funnel checks instead of a new creative angle.",
"files": [],
"expectations": [
"Identifies healthy CTR with collapsing CVR as the inverse of the fatigue signature, pointing at landing page, offer, or tracking",
"Issues a confounded verdict, not fatigued",
"Refuses to develop the requested new creative angle",
"Ties the decline onset to the checkout deploy and/or the expired 20%-off code dates",
"Recommends comparing conversion rate for other traffic sources hitting the same page",
"Recommends reconciling platform-reported conversions against the order system, checking for a stable ratio between the two rather than equality",
"Reports the confounder screen before reaching any fatigue conclusion",
"States that no creative refresh can fix a problem downstream of the ad",
"Recommends handing the structural or tracking root cause to an account-level diagnostic follow-up",
"States the fatigue signature for contrast: costs rise while conversion rate holds",
"Delivers a verdict block with confidence and re-check fields",
"Frames expected recovery as tied to fixing the funnel or offer, not to any new creative"
]
},
{
"id": 4,
"prompt": "Small B2B account here - Kindler & Co, we sell compliance training. Our Meta webinar-signup ad's link CTR was 0.95% over the last two days, about 8,000 impressions total, against 1.1% for the trailing two weeks. We got 3 signups this week vs 5 last week. Spend is about $70/day. My manager wants the creative killed and replaced by Monday because the data says it's done. Is he right?",
"expected_output": "An insufficient-data verdict backed by the sampling-noise band, refusing the kill and stating what extra data would clear the gate.",
"files": [],
"expectations": [
"Applies the sampling-noise band p plus-or-minus 2 times sqrt(p(1-p)/n) to the baseline rate",
"Concludes the 0.95% read sits inside the noise band around 1.1% at roughly 8,000 impressions (band approximately 0.87%-1.33%), so the observed decline is noise",
"Notes two days of data fails the requirement of movement across two or more consecutive periods",
"Issues an insufficient data verdict",
"Refuses to recommend killing or replacing the creative by Monday",
"States the additional days, impressions, or spend needed before the noise band becomes narrower than the observed delta",
"Notes that 3-5 weekly conversions cannot clear the conversion floor for any CPA- or CVR-based claim",
"Recommends gating this low-conversion account on leading engagement signals and saying so in the verdict's confidence line",
"Warns that retiring a healthy creative on an underpowered read is the expensive false positive to avoid",
"Uses a moving-average comparison against the creative's own trailing baseline instead of comparing single days",
"Delivers a verdict block with confidence marked low and the gate math given as its basis",
"Sets a re-check date after enough data accumulates"
]
},
{
"id": 5,
"prompt": "Ledgerline - fintech, ABM motion on LinkedIn. Our matched account list renders about 38,000 members. We have run five sponsored-content ads for ten weeks. All five are decaying together - even the one we launched fresh three weeks ago, which fell to the pack's numbers within two weeks. CTR 0.29% vs 0.43% earlier, reach has been flat for a month while impressions keep climbing, frequency 7.8, demo requests down from about 6 a week to 2-3, and click-to-demo rate slid too. Leadership already approved budget for a full new creative batch from our agency. How many new ads should we brief and what angles?",
"expected_output": "A saturating verdict: the 38k pool is depleting, the fresh-creative test already failed on the same audience, so the remedy is audience expansion, not the approved creative batch.",
"files": [],
"expectations": [
"Issues a saturating verdict, not fatigued",
"Cites all five creatives decaying together, including the fresh one, as pointing at the pool rather than the assets",
"Reads the fresh creative's failure to recover performance on the same audience as the discriminating test resolving toward saturation",
"Reads flat reach with climbing impressions as the platform recycling the same pool",
"Reads conversion rate degrading alongside engagement as the saturation signature, contrasting it with fatigue where CVR typically holds",
"Recommends audience expansion, list refresh, or exclusions as the remedy instead of new creative",
"Explicitly advises against briefing the approved new creative batch, stating creative work is wasted effort under saturation",
"Notes ABM and list-based audiences saturate by design, so B2B should expect saturating verdicts more often than fatigued",
"Gates confidence on leading engagement signals because single-digit weekly conversions cannot clear the conversion floor, and says so in the confidence line",
"Promotes audience expansion to the first action under the saturating verdict",
"Uses new-member reach or first-time-exposure trend as a leading saturation signal",
"Delivers a verdict block including the confounder screen results",
"Sets a re-check one full comparison window after the audience expansion"
]
},
{
"id": 6,
"prompt": "Trailnut, snack brand. Our Meta UGC video 'trail-taste-test' has run nine weeks and is still our volume driver. Over the last three weekly reads: hook rate 24% then 19% then 16%; hold rate at 15 seconds basically flat (9.1% vs 9.4% baseline); link CTR down about 20%; CVR stable at 2.8%. Hide/report rate has tripled vs our account norm. No budget, audience, or landing-page changes, and the deltas are way outside noise on about 150k weekly impressions. We have an in-house editor with spare capacity this week. My plan is to start over: brief an entirely new concept with our agency, about six weeks out. Sound right?",
"expected_output": "A fatigued verdict localised to the hook, prescribing a hook swap shipped as a new ad ID (because negative feedback is elevated) instead of the six-week new concept.",
"files": [],
"expectations": [
"Confirms a fatigued verdict from two or more signals moving together across consecutive periods with the gate cleared",
"Localises the decay to the hook: hook rate falling while hold rate and CVR stay stable means the opening is tired, not the whole ad",
"Attributes the hook-fatigue framing to Ben Heath or labels it as a practitioner heuristic",
"Recommends a hook/thumbnail swap on the same body as the first action, ahead of any new concept",
"States the six-week new concept over-serves the evidence, being the highest-effort remedy reserved for an exhausted concept",
"Because negative feedback is elevated, requires the swap to ship as a new ad/asset ID, since the algorithmic penalty rides the asset ID and a hook swap on the same ID does not clear it",
"Instructs never to edit the live ad because edits reset learning and destroy the baseline; the swap launches alongside",
"Keeps the current ad running while the swap ramps, never pausing a producer with nothing staged",
"Assigns the hook swap to the in-house editor, matching its effort profile",
"States expected recovery as hook rate and link CTR back toward baseline on the same audience, which also confirms fatigue over saturation",
"Sets a re-check one full comparison window after the swap launches",
"Delivers a verdict block including the confounder screen and a confidence line",
"Positions iterating the winner, not the new concept, as the fallback if the hook swap fails"
]
},
{
"id": 7,
"prompt": "Norvane is a B2B SaaS running Google Ads. Our main RSA's Ad Strength dropped from 'Good' to 'Average' last week and our new account manager says that means the ad is fatiguing and we must rewrite all the assets this sprint. CTR and conversion rate have been flat for two months. The asset report shows nine assets at 'Good' or 'Best' and one headline stuck at 'Low' for five weeks. Do we rewrite everything?",
"expected_output": "Rejects Ad Strength as a fatigue signal, reads the flat CTR/CVR as healthy, and recommends replacing only the persistent Low headline based on asset performance labels.",
"files": [],
"expectations": [
"Rejects Ad Strength as a fatigue or performance signal",
"States Ad Strength measures asset diversity and completeness, not performance",
"Cites Optmyzr's analysis finding 'Average'-strength ads with the best CPA/CVR as evidence",
"Points to the asset performance labels (Learning/Low/Good/Best) as the performance-based signal on RSAs",
"Recommends replacing only the headline stuck at 'Low' for five weeks, not a full rewrite",
"Labels any replacement cadence such as 2-4 weeks as practitioner advice, not a Google rule",
"Reads two months of flat CTR and conversion rate as no decay, yielding a healthy verdict for the ad overall",
"Refuses the account manager's full asset rewrite",
"Bases any future fatigue read on the ad's own trailing performance baseline, not on platform score movement",
"Delivers a verdict block with confidence and re-check fields",
"Does not recommend pausing or restructuring the RSA"
]
},
{
"id": 8,
"prompt": "Aurelia Waters is a premium bottled-water brand. We run a 60-second brand film on YouTube and Meta with a brand-awareness objective - we optimize for lift and view metrics, not conversions. Eight weeks in, view rate and our aided-recall survey numbers keep improving while frequency crept to 4.1. Our agency's policy is mandatory creative rotation every four weeks because all ads wear out, and they say we are overdue and risking burnout. Should we force the rotation?",
"expected_output": "A wear-in verdict valid because the objective is brand: leave the ad alone, refuse rotation on a frequency number or calendar policy, keep monitoring.",
"files": [],
"expectations": [
"Identifies performance improving under repeated exposure as wear-in",
"Notes wear-in is a plausible verdict for brand objectives, attributed to Les Binet",
"Notes Meta's 2023 research found no wear-in for direct-response objectives, so the verdict is available only because this objective is brand",
"Issues a wear-in verdict and recommends leaving the ad alone",
"Refuses to rotate the creative on a frequency number",
"Rejects the agency's universal four-week rotation policy as not derived from this creative's own data",
"Verifies the improvement holds across multiple consecutive periods against the creative's own baseline",
"Still screens confounders such as seasonality, auction, or mix shift before accepting the improving trend",
"States the same improving read would be treated with suspicion if the objective were direct response",
"Delivers a verdict block with confidence and a re-check or monitoring date",
"Continues scheduled monitoring rather than declaring the creative permanently immune to wear-out"
]
},
{
"id": 9,
"prompt": "Fernwald Outfitters, outdoor apparel. Our Meta Advantage+ shopping campaign's CPA is up 22% over three weeks. We have 28 live ads that boil down to six concepts. Our head of growth wants a full creative refresh - all 28 replaced next month - because the campaign is fatigued. Our exports can go per-creative per-day. What should the refresh roadmap look like?",
"expected_output": "Rejects the campaign-average diagnosis, moves the analysis to per-creative and per-concept level with asset breakdowns under Advantage+, and refuses the blanket 28-ad refresh.",
"files": [],
"expectations": [
"Rejects the campaign-level average as the unit of analysis",
"Requires analysis at creative and concept level, grouping the 28 ads under their six concepts",
"States that one or a few fatigued ads can drag the campaign average while most creatives stay healthy",
"Notes Advantage+ automated rotation silently shifts budget away from fatigued assets, masking per-asset decay, so asset-level breakdowns must be read",
"Flags a creative silently losing budget share with no manual change as itself a decay signal to check per creative",
"Adopts the offered per-creative per-day export as the working data granularity",
"Refuses the blanket refresh of all 28 ads",
"Builds each creative's own trailing baseline before judging any of them",
"Runs the confounder screen on the CPA rise before attributing anything to fatigue",
"Defers any fatigue verdict until the per-creative read is complete, or issues insufficient data at campaign grain",
"Treats CPA as a lagging signal and requires leading signals per creative before any call",
"Scopes the output to identifying which creatives, if any, need action rather than scheduling a wholesale refresh"
]
},
{
"id": 10,
"prompt": "Bexley Books, DTC bookseller. Our analyst's weekly deck shows our evergreen carousel fatiguing: every Monday we compare the last seven days against the prior seven, CPA is always up around 25%, and then the gap shrinks by the next pull. Mid-month we also moved the dashboard from 7-day click to 1-day click attribution to be conservative. CTR (all) is stable so engagement seems fine. Marketing wants a refresh plan for the carousel. What's the plan?",
"expected_output": "Identifies the decline as an attribution-lag artefact compounded by the attribution-setting switch and the all-clicks CTR read; no refresh, fix the measurement instead.",
"files": [],
"expectations": [
"Identifies trailing-window attribution lag: the most recent days of any trailing window under-report conversions by construction",
"States every trailing window ends in an apparent decline, consistent with the gap shrinking on the next pull as conversions backfill",
"Requires both comparison windows to be lag-mature before any delta is read",
"Flags the mid-month switch from 7-day click to 1-day click as manufacturing an artificial decline and requires the same attribution setting on both windows",
"Rejects CTR (all) and requires link/outbound CTR, because all-clicks counts reactions, comments, and expands and can flatter a dying ad",
"Issues a confounded or insufficient-data verdict rather than fatigued",
"Refuses to produce the requested refresh plan",
"Recommends re-pulling the same windows after the attribution window closes to confirm the backfill",
"Runs the remaining confounder screen before any verdict",
"Requires the two-signal rule with at least one leading signal before any future fatigue call",
"Delivers a verdict block with confidence and re-check fields",
"Names fixing the measurement method - matched, mature windows on one consistent attribution setting - as the action"
]
},
{
"id": 11,
"prompt": "Glintwork sells jewelry. It is mid-November and our nine-week-old Meta prospecting video suddenly got expensive: CPA up 28% and CPM up 24% over three weeks. Link CTR is 1.35% vs 1.33% baseline and CVR is flat. The whole team is convinced the ad wore out right before Black Friday and wants an emergency two-week creative sprint. Green-light it?",
"expected_output": "A confounded verdict: CPM up with CTR flat is Q4 auction inflation, not wear-out; no emergency sprint.",
"files": [],
"expectations": [
"Reads CPM rising while CTR holds flat as auction-side inflation, not creative wear-out",
"States CPM is diagnostic for fatigue only when paired with falling CTR",
"Attributes the cost rise to seasonal Q4/Black Friday auction pressure raising the price of every impression",
"Recommends comparing the creative's CPM trend against the account's other ad sets, and category benchmarks where available, over the same weeks",
"Notes stable CTR and CVR mean the audience responds exactly as before and each response simply costs more",
"Issues a confounded verdict",
"Refuses to green-light the emergency creative sprint",
"Notes auction inflation typically shows account-wide and industry-wide, not on one creative",
"Checks comparison-window composition for promo days and day-of-week match given the season",
"Delivers a verdict block with confidence and re-check fields",
"Treats any fixed CPM-inflation percentage threshold as unsourced folklore rather than a decision rule"
]
},
{
"id": 12,
"prompt": "Mapleworks, B2B HR software. Our main Meta lead-gen video looks genuinely done: link CTR down 27% and hook rate down 24% vs its own 30-day baseline across four consecutive weekly reads, about 180k impressions in the window so the deltas are far outside noise, CVR steady, and its spend share keeps falling with no manual change. Change history is clean - no edits since June, budget flat - the landing page is untouched, CPM is in line with our other ad sets, first-time impression ratio sits at 71%, and tracking reconciles with our CRM. Constraints: our board demo is in four days and the CEO wants the CPL bleeding stopped before it; our only other live ad is a static that performs fine; our video editor left last month; agency lead time is five weeks minimum; and legal is out this month so no new audience-list upload can be signed off. What do we do?",
"expected_output": "A fatigued verdict whose action ladder is pruned by the stated constraints: budget shift to the healthy static now, every deleted rung named with its deleting constraint, production staged for after the deadline.",
"files": [],
"expectations": [
"Issues a fatigued verdict from the supplied evidence: two leading signals across four consecutive periods, confounders passing, gate cleared",
"Deletes the rungs the constraints rule out rather than merely demoting them to the bottom of the list",
"Rules out the hook swap, naming the departed video editor as the deleting constraint",
"Rules out iterating and the new concept, naming the five-week agency lead time against the four-day deadline",
"Rules out audience expansion, naming the unavailable legal/lawful-basis sign-off",
"Treats rotation as limited to the single healthy static or rules it out for lack of a challenger bench",
"Selects a same-day surviving action - shifting budget toward the healthy static, optionally with a frequency-cap check - as the highest-ranked remaining rung",
"Keeps the still-profitable fatigued video running rather than pausing a producer with nothing staged",
"States the same-day action buys time and CPA protection, not a fix",
"Names every deleted rung with its deleting constraint on the verdict's ruled-out line",
"Does not recommend editing the live video",
"Recommends starting replacement production for after the deadline so the deleted rungs return as funded scope rather than silently disappearing",
"Delivers a verdict block with a re-check one full comparison window after the action"
]
}
],
"trigger_queries": [
{ "query": "Is my Facebook ad fatigued? CTR dropped 20% over the last two weeks", "should_trigger": true },
{ "query": "our best meta ad is dying, do we need new creative?", "should_trigger": true },
{ "query": "frequency on our prospecting campaign hit 4.2, should I rotate the ads?", "should_trigger": true },
{ "query": "how do I know when an ad creative is worn out?", "should_trigger": true },
{ "query": "CPA on our 8-week-old video ad keeps climbing, is the creative done?", "should_trigger": true },
{ "query": "creative wear-out check on our TikTok campaign please", "should_trigger": true },
{ "query": "when should we refresh our ad creatives?", "should_trigger": true },
{ "query": "Meta is showing 'creative fatigue' status on two ads, what now?", "should_trigger": true },
{ "query": "this ad crushed it for two months and now it doesn't - why?", "should_trigger": true },
{ "query": "same ad, same audience, same budget, worse results every week. what's going on?", "should_trigger": true },
{ "query": "should I kill this ad or let it ride?", "should_trigger": true },
{ "query": "the video that carried our account all summer suddenly costs twice as much per sale", "should_trigger": true },
{ "query": "hook rate on our top UGC ad went from 25% to 17%, is it tired?", "should_trigger": true },
{ "query": "do I really need to refresh creatives every 2 weeks like my agency says?", "should_trigger": true },
{ "query": "our LinkedIn ads have been running 8 weeks and engagement keeps sliding", "should_trigger": true },
{ "query": "ad fatigue analysis for our holiday campaign", "should_trigger": true },
{ "query": "is rising frequency a reason to swap creative?", "should_trigger": true },
{ "query": "CTR decay on a long-running search ad - fatigue or something else?", "should_trigger": true },
{ "query": "my winning ad stopped winning, help me figure out if it's worn out", "should_trigger": true },
{ "query": "how long can one ad creative keep performing before it burns out?", "should_trigger": true },
{ "query": "our retargeting ads feel stale - people must be sick of seeing them", "should_trigger": true },
{ "query": "results tanked on an ad we haven't touched in 6 weeks", "should_trigger": true },
{ "query": "is it time to retire our hero ad?", "should_trigger": true },
{ "query": "has our audience seen this ad too many times? how do I tell?", "should_trigger": true },
{ "query": "ROAS on our oldest creative keeps drifting down month over month", "should_trigger": true },
{ "query": "diagnose whether this creative decline is real or just noise", "should_trigger": true },
{ "query": "my boss says the ad is burned out because frequency is 3.5 - true?", "should_trigger": true },
{ "query": "when does ad wear-out actually start on Meta cold audiences?", "should_trigger": true },
{ "query": "should we refresh creative before Black Friday or hold what's running?", "should_trigger": true },
{ "query": "the same TikTok spark ad has run 3 weeks and delivery is trending down", "should_trigger": true },
{ "query": "CPL on our lead ad rose 30% but nothing changed in the account - is the ad spent?", "should_trigger": true },
{ "query": "our ad's thumbstop rate keeps dropping week over week", "should_trigger": true },
{ "query": "people are hiding our ad more than usual and results are dipping - worn out?", "should_trigger": true },
{ "query": "how do I tell ad fatigue apart from audience saturation?", "should_trigger": true },
{ "query": "is this creative decline just seasonality or actual fatigue?", "should_trigger": true },
{ "query": "first-time impression ratio fell below 50%, do we have a fatigue problem?", "should_trigger": true },
{ "query": "my Google RSA has run for 6 months - does search creative fatigue exist?", "should_trigger": true },
{ "query": "which metrics actually prove an ad is wearing out?", "should_trigger": true },
{ "query": "give me a keep, refresh, or kill verdict on this long-running ad", "should_trigger": true },
{ "query": "our YouTube ad's view rate has slid three weeks straight", "should_trigger": true },
{ "query": "old ad is still profitable but declining - when do I pull the plug?", "should_trigger": true },
{ "query": "conversion costs creeping up on an ad that's been live since March", "should_trigger": true },
{ "query": "I think our audience is tired of this ad but my cofounder disagrees - settle it", "should_trigger": true },
{ "query": "what's the right refresh cadence for our Meta creatives?", "should_trigger": true },
{ "query": "here are weekly CTRs for our main ad: 1.4, 1.3, 1.1, 0.9 - creative fatigue?", "should_trigger": true },
{ "query": "ad performance decay - how do I confirm it's actually the creative?", "should_trigger": true },
{ "query": "our agency wants to rebuild all creatives because performance dipped last week - overkill?", "should_trigger": true },
{ "query": "why is my ad getting more expensive the longer it runs?", "should_trigger": true },
{ "query": "ad ran great for 6 weeks, now CTR is down and CPM is up - worn out?", "should_trigger": true },
{ "query": "should a still-working ad be rotated out preventively?", "should_trigger": true },
{ "query": "how many weeks does a UGC ad usually last before it wears out?", "should_trigger": true },
{ "query": "spend share of our best ad keeps shrinking even though we changed nothing", "should_trigger": true },
{ "query": "everyone's seen our ad already, right? reach flatlined", "should_trigger": true },
{ "query": "my ad is on a slow decline - is it the creative or the auction?", "should_trigger": true },
{ "query": "is my ad wearing in or wearing out? numbers actually improved with frequency", "should_trigger": true },
{ "query": "getting worse results from the same creative in month 3 - is that normal?", "should_trigger": true },
{ "query": "check if our top 3 ads are fatiguing", "should_trigger": true },
{ "query": "when a proven ad slips, how long should I wait before replacing it?", "should_trigger": true },
{ "query": "the creative that used to print money is bleeding now - diagnosis?", "should_trigger": true },
{ "query": "help me decide between refreshing the hook or making a whole new ad for a declining video", "should_trigger": true },
{ "query": "our ad's cost per result doubled and Meta flagged it creative limited", "should_trigger": true },
{ "query": "banner has been in rotation 10 weeks on native placements - past its shelf life?", "should_trigger": true },
{ "query": "customers keep seeing the same ad and I'm worried about burnout - what would the data need to show?", "should_trigger": true },
{ "query": "one creative in the ad set fell off a cliff while the others are fine", "should_trigger": true },
{ "query": "prospecting video slowing down: fatigue call or budget issue?", "should_trigger": true },
{ "query": "how do I baseline an ad so I can spot wear-out early?", "should_trigger": true },
{ "query": "is a 25% CTR decline over 3 weeks enough to declare fatigue?", "should_trigger": true },
{ "query": "ad performance dropped, nothing changed, audience is small - fatigue or saturation?", "should_trigger": true },
{ "query": "our carousel's engagement halved since last month and it's run since spring", "should_trigger": true },
{ "query": "declining hook rate but stable hold rate on our main video - what does that mean?", "should_trigger": true },
{ "query": "score these five hook options before we launch the campaign", "should_trigger": false },
{ "query": "which of these three video openings deserves test budget?", "should_trigger": false },
{ "query": "write 10 headline variants for our spring sale ads", "should_trigger": false },
{ "query": "turn this value prop into distinct ad copy angles", "should_trigger": false },
{ "query": "write a UGC script for our skincare brand's new serum", "should_trigger": false },
{ "query": "draft the replacement UGC video script for the ad we're retiring", "should_trigger": false },
{ "query": "put together a creative brief our designer can execute", "should_trigger": false },
{ "query": "brief for a short-form video shoot targeting gym owners", "should_trigger": false },
{ "query": "design an A/B test to compare our two new ad concepts", "should_trigger": false },
{ "query": "how much budget does a clean creative split test need?", "should_trigger": false },
{ "query": "our entire ad account's ROAS dropped across every campaign this month - find the root cause", "should_trigger": false },
{ "query": "audit my ad account, everything underperforms since the iOS update", "should_trigger": false },
{ "query": "are we pacing to spend our $60k monthly ad budget?", "should_trigger": false },
{ "query": "we're underspending against the Q4 budget - fix the pacing", "should_trigger": false },
{ "query": "our campaign is crushing it - how fast can I scale the budget?", "should_trigger": false },
{ "query": "what's a safe weekly budget ramp for a winning ad set?", "should_trigger": false },
{ "query": "what frequency caps should each stage of my retargeting sequence get?", "should_trigger": false },
{ "query": "design a cart-abandonment retargeting funnel with exclusions", "should_trigger": false },
{ "query": "build a cold prospecting targeting plan for our SaaS", "should_trigger": false },
{ "query": "which interest and lookalike segments should we layer for launch?", "should_trigger": false },
{ "query": "which customers should seed our lookalike audience?", "should_trigger": false },
{ "query": "is 800 matched users enough for a value-based lookalike?", "should_trigger": false },
{ "query": "Meta says 240 purchases, Shopify says 180 - reconcile the gap", "should_trigger": false },
{ "query": "why does the ad platform report more revenue than our CRM?", "should_trigger": false },
{ "query": "verify our pixel isn't double-firing before we launch Monday", "should_trigger": false },
{ "query": "check conversion event dedup between browser and CAPI", "should_trigger": false },
{ "query": "is a $95 CAC acceptable for a $40/month subscription?", "should_trigger": false },
{ "query": "benchmark our ROAS - healthy or not?", "should_trigger": false },
{ "query": "ads get clicks but the landing page doesn't convert - audit the page", "should_trigger": false },
{ "query": "review my landing page for message match with the ad", "should_trigger": false },
{ "query": "carousel or video for a cold awareness campaign?", "should_trigger": false },
{ "query": "which ad formats fit a lead-gen objective on LinkedIn?", "should_trigger": false },
{ "query": "organize these competitor ads into a swipe file", "should_trigger": false },
{ "query": "catalog what competitors are running in the ad library right now", "should_trigger": false },
{ "query": "should we move from cost cap to bid cap on this campaign?", "should_trigger": false },
{ "query": "pick a target ROAS bidding setup for our new campaign", "should_trigger": false },
{ "query": "we have 40 campaigns doing the same thing - plan a consolidation", "should_trigger": false },
{ "query": "merge these ad sets without resetting learning", "should_trigger": false },
{ "query": "split $50k monthly between Google, Meta and TikTok", "should_trigger": false },
{ "query": "how much of the budget should prospecting vs retargeting get?", "should_trigger": false },
{ "query": "set a maximum allowable CAC policy for the company", "should_trigger": false },
{ "query": "define kill-switch spend thresholds for our paid program", "should_trigger": false },
{ "query": "mine our search terms report for negative keywords", "should_trigger": false },
{ "query": "our search campaign matches junk queries - build exclusion lists", "should_trigger": false },
{ "query": "which paid channels fit a $3k/month budget for a B2B tool?", "should_trigger": false },
{ "query": "is connected TV worth it for our audience?", "should_trigger": false },
{ "query": "promote our CEO's viral LinkedIn post as a thought leader ad", "should_trigger": false },
{ "query": "plan a founder-fronted ads program from his organic posts", "should_trigger": false },
{ "query": "write copy for an ad slot inside an AI assistant's answer", "should_trigger": false },
{ "query": "map the buying committee for our six-figure ERP deal", "should_trigger": false },
{ "query": "how do I break into media buying as a career?", "should_trigger": false },
{ "query": "review my performance marketing portfolio for job applications", "should_trigger": false },
{ "query": "write a job description for a senior media buyer", "should_trigger": false },
{ "query": "which PPC newsletters and podcasts should I follow?", "should_trigger": false },
{ "query": "where do I start with paid ads for my startup?", "should_trigger": false },
{ "query": "my email open rates have been declining for months", "should_trigger": false },
{ "query": "newsletter subject line fatigue - subscribers stopped opening", "should_trigger": false },
{ "query": "our organic Instagram reach keeps dropping", "should_trigger": false },
{ "query": "SEO traffic to our blog has decayed since the core update", "should_trigger": false },
{ "query": "I'm creatively burned out as a designer, how do I recover?", "should_trigger": false },
{ "query": "act as a creative director and set the visual direction for our rebrand", "should_trigger": false },
{ "query": "our sales team is fatigued by the new CRM workflow", "should_trigger": false },
{ "query": "our app store screenshots need a refresh - what converts best?", "should_trigger": false },
{ "query": "website hero image has been the same for a year, should we change it?", "should_trigger": false },
{ "query": "our billboard has been up for 6 months - is OOH wear-out measurable?", "should_trigger": false },
{ "query": "we made two new ads - which one should we launch first?", "should_trigger": false },
{ "query": "estimate the sample size for testing a new hook against the control", "should_trigger": false },
{ "query": "how often should I post new organic TikToks to avoid follower fatigue?", "should_trigger": false },
{ "query": "set up automated rules to pause ads when CPA exceeds target", "should_trigger": false },
{ "query": "our push notification opt-outs are climbing - are we messaging too often?", "should_trigger": false }
]
}
references/confounders.md›
# Confounder Screen - The Differential Diagnosis Playbook
Run every section before any fatigue verdict. Each entry gives the telltale pattern that distinguishes the confounder from true fatigue, and the check that confirms it (screen FAILS, the confounder explains the decline) or clears it (screen PASSES, move on). A single confirmed confounder that accounts for the decline ends the fatigue inquiry: the verdict is `confounded` (or `saturating` for the audience case) and the fix targets the cause, not the creative.
True fatigue's signature, for contrast:
- Gradual, sustained across two-plus periods.
- Specific to individual creatives rather than the whole account.
- Visible in two or more signals at once.
- Conversion rate typically holding while costs rise.
Run all eleven regardless; this order only decides what to check first when the clock is short, ranked by decline explained per minute of checking:
- efficiency: change-history reads (§1 budget/bid, §2 learning-phase reset, §11 sibling mix, §10 landing page/offer) > window-composition recomputes (§5 seasonality, §7 attribution maturity, §9 noise) > breakdown pulls (§8 placement/device, §4 auction CPM against the account's other ad sets) > pool trending (§3 saturation)
- effort, lightest first: change-history reads (one log, minutes) > window recomputes (re-run the delta on matched windows) > breakdown pulls (one export per dimension) > pool trending (a reach series, and sometimes a live fresh-creative test)
The top group leads on both axes because it is where this file's own biggest false-positive generator sits - a learning-phase reset is answered by a date, not by analysis. Re-rank for the account: an account nobody has touched in six weeks makes the change-history group cheap but empty, which promotes the window recomputes to first.
## 1. Budget or bid change
More spend forces the platform into broader, cheaper-to-lose auctions and lower-intent pockets of the audience; efficiency degrades in a way that looks identical to fatigue - CTR down, CPA up - within days of the increase.
- Telltale: the decline starts at, or within the learning re-stabilisation window after, a budget/bid step - not gradually. The whole ad set moves together, not one creative.
- Check: pull the change history and overlay change dates on the metric series. Decline onset aligned to a budget/bid event → FAIL. If the budget rose, compare against the _pre-increase_ efficiency at the old spend level, not against the peak; some efficiency loss at higher spend is the normal cost of scale, not decay.
## 2. Learning-phase reset
Any significant edit - creative tweak, budget step beyond the platform's tolerance, audience or optimization-event change, bid-strategy change, a long pause - resets algorithmic learning, and delivery swings wildly for days afterwards. New launches do the same in their first days. This is the single biggest false-positive generator.
- Telltale: high day-to-day variance rather than a steady slide; onset exactly at an edit or launch; platform UI showing "learning" status.
- Check: interview answer plus change history for the last-edit date. Any significant edit inside the comparison window → FAIL; re-measure only on a clean window starting after delivery stabilises. Meta's guidance: allow at least 7 days after a significant edit before evaluating.
## 3. Audience overlap / audience saturation
The pool is depleting - through small audience size, overlapping ad sets bidding against each other, or simply having reached everyone - and no creative, however fresh, fixes a tapped-out pool. Fatigue and saturation need opposite remedies (new creative vs audience expansion), which is why this split gets its own verdict.
- Telltale: first-time impression ratio falling, reach flattening while impressions climb, frequency grinding upward account-wide, and - the key separator - conversion rate degrading along with engagement (fatigue usually leaves CVR stable while costs rise). Multiple creatives in the same audience decaying together points at the pool, not the assets.
- Check: trend first-time impression ratio (or new-user reach) and rolling reach. Where feasible, run the discriminating test from Meta's research: launch a fresh creative to the _same_ audience - recovery means it was fatigue; no recovery, but performance returning on a _fresh_ audience, means saturation → verdict `saturating`. Also check audience-overlap tooling for sibling ad sets competing for the same people.
## 4. Auction-side CPM inflation
Competitor entry, seasonal auction pressure (Q4, sales periods, elections) raises the price of every impression. Costs rise with the creative doing nothing wrong.
- Telltale: CPM up while CTR holds steady - the audience responds exactly as before; each response just costs more. Usually visible account-wide and industry-wide, not on one creative.
- Check: compare the creative's CPM trend against the account's other ad sets and, if available, category benchmarks for the same weeks. CPM up, CTR flat → FAIL (auction, not fatigue). CPM up _and_ CTR down remains consistent with fatigue - keep screening.
## 5. Seasonality and comparison-window composition
Weekends, holidays, paydays, and seasonal demand shifts change who is online and how they behave. A comparison window with different day-of-week or holiday composition than its baseline manufactures a decline from nothing.
- Telltale: the "decline" disappears when comparing like-for-like days (this Tuesday-to-Monday vs last Tuesday-to-Monday), or mirrors last year's seasonal curve.
- Check: recompute the delta on day-of-week-matched windows excluding holidays and promo days. Delta shrinks into the noise band → FAIL.
## 6. Tracking breakage, deduplication faults, consent-mode shifts
A pixel outage, a broken event, a server-side/browser dedup fault, or a consent-banner change under-reports conversions. Reported CPA rises while real performance is unchanged - measurement decayed, not the creative.
- Telltale: conversions drop suddenly (often to a step-function new level) while clicks and engagement hold; the drop coincides with a site release, tag change, or consent update; platform-reported conversions diverge from the order system.
- Check: reconcile platform-reported conversions against the source of truth (orders, CRM) for the window. The healthy state is a _stable ratio_ between the two, not equality; a ratio that lurches at the decline's onset → FAIL. Also check the platform's event diagnostics for match-quality or event-volume drops.
## 7. Attribution-window skew
Platforms attribute conversions to the click date and keep crediting for days afterwards, so the most recent days of any trailing window are always under-reported. Every trailing window therefore ends in an apparent decline - phantom decay by construction. Comparing metrics across different attribution windows (7-day-click vs 1-day-click) manufactures the same artefact.
- Telltale: the "decay" is concentrated in the last few days of the window and backfills upward when re-pulled a week later; or the baseline and comparison windows use different attribution settings.
- Check: compare only lag-mature windows (both windows old enough that the attribution window has fully closed), same attribution setting on both. Decline evaporates on mature windows → FAIL.
## 8. Placement or device mix shift
The same ad delivering into a different placement/device mix (more Audience Network, more right-column, more mobile) shows different blended CTR/CPM without any change in creative effectiveness - the mix moved, not the response within any cell.
- Telltale: blended metrics decline while per-placement metrics are stable; the placement/device breakdown shows share shifting toward structurally-lower-CTR inventory.
- Check: break the metric down by placement and device and compare within cells against the same cells in the baseline. Within-cell stability with shifted mix → FAIL.
## 9. Normal statistical noise on small numbers
Small daily denominators make rates jump around; a two-day dip on a few thousand impressions is a coin flip, not a trend.
- Telltale: the observed delta sits inside the baseline's sampling-noise band; the "trend" is one or two periods old; single-day granularity on a low-volume creative.
- Check: run the Confidence Gate's noise check (`p ± 2 × sqrt(p(1-p)/n)` on the baseline window). Delta inside the band, or fewer than two consecutive periods of movement → FAIL (as in: not evidence of anything). This is the confounder the gate exists to formalise.
## 10. Landing page or offer change downstream of the ad
A page redesign, price change, stock-out, expired promo, slower page, or broken form collapses conversion downstream of a perfectly healthy ad.
- Telltale: CTR and engagement healthy, CVR down - the exact inverse of the fatigue signature. Onset aligned to a site deploy or offer calendar. All traffic sources to that page degrade together, not just this creative.
- Check: overlay site/offer change dates; compare CVR for other traffic hitting the same page. Page-wide CVR drop → FAIL; the fix is the funnel, not the creative.
## 11. Sibling-mix shift inside the ad set
Adding or removing creatives changes how the platform splits delivery. A new sibling siphons spend and the best pockets of the audience; a removed sibling dumps its (possibly worse-fit) delivery onto the survivors. The measured creative's blended numbers move without its effectiveness changing. Cosmetic resizes count as siblings too - near-duplicates cannibalise each other's reach.
- Telltale: the decline's onset matches a sibling launch/pause in the same ad set; spend share moved at the same moment; the creative's within-segment response is stable where it still delivers.
- Check: change history for sibling adds/removals inside the window; trend each sibling's spend share. Mix event aligned with the decline → FAIL - re-baseline from the new mix's stabilisation date instead of refreshing the creative.
references/examples.md›
# Worked Examples
Three illustrative cases with filled Fatigue Verdict blocks. All figures are invented for illustration - internally consistent, but not benchmarks; every real verdict derives its numbers from the account's own data.
## Table of Contents
- [Example A - B2C ecommerce, Meta: this one IS fatigue](#example-a---b2c-ecommerce-meta-this-one-is-fatigue)
- [Example B - B2B, LinkedIn: the audience is saturating, not the creative fatiguing](#example-b---b2b-linkedin-the-audience-is-saturating-not-the-creative-fatiguing)
- [Example C - NEGATIVE example: a convincing fatigue call that is actually a budget scale-up](#example-c---negative-example-a-convincing-fatigue-call-that-is-actually-a-budget-scale-up)
## Example A - B2C ecommerce, Meta: this one IS fatigue
Skincare brand, cold prospecting, broad audience (~2.1M), UGC video "morning routine" running 6 weeks in a 4-creative ad set. No edits, no budget changes, stable $900/day at ad-set level. Data: per-creative per-day export; 3-day moving average vs the creative's own trailing 14-day baseline.
Noise check on the lead signal: baseline link CTR 1.32% on ~210,000 comparison-window impressions gives a noise band of 1.32% ± 2×sqrt(0.0132×0.9868/210000) ≈ 1.27%-1.37%. Observed 0.98% sits far outside it. Two leading signals plus one confirming signal moved together across four consecutive 3-day periods; CVR held - the classic fatigue signature (costs rise, people who still click still buy).
```
FATIGUE VERDICT - "UGC-morning-routine-v3", 2026-08-20
platform : Meta | funnel stage: cold prospecting
window : Aug 08-20 (3-day MA) vs baseline Jul 25-Aug 07 (trailing 14-day)
volume : $4,120 spent, 209,700 impressions, 118 conversions in window
signals
link CTR : 0.98% vs 1.32% (-26%) [leading]
hook rate (3s) : 17.9% vs 24.1% (-26%) [leading]
hold rate (15s) : 8.7% vs 9.0% (-3%, flat) [leading]
spend share (ad set) : 22% vs 34% (no manual change) [leading]
frequency : 2.3 vs 1.9 [lagging]
CVR : 3.4% vs 3.5% (stable) [lagging]
CPA : $34.90 vs $26.10 (+34%) [lagging]
confounder screen
budget/bid change : pass - no changes in 6 weeks
learning-phase reset : pass - last significant edit Jul 02
audience saturation : pass - first-time impression ratio 68%, reach still growing
auction CPM inflation : pass - CPM +4%, in line with account's other ad sets
seasonality/window mix : pass - day-of-week-matched windows, no promo days
tracking breakage : pass - platform-to-orders ratio stable at ~0.83
attribution-window skew : pass - both windows lag-mature, same 7-day-click setting
placement/device mix : pass - placement shares within 2pts of baseline
statistical noise : pass - delta far outside the 2-SE band (see above)
landing page/offer change: pass - no deploys; other traffic to page converting normally
sibling-mix shift : pass - same 4 creatives all window
confidence : high - 2 leading + 2 confirming signals, 4 consecutive periods, gate cleared with margin
verdict : fatigued
action : rung 1 - hook/thumbnail swap on the same body (hook rate fell, hold rate held:
the opening is tired, not the ad). Negative feedback normal, so same body is fine,
but launch as a NEW ad alongside - never edit the live one. Asset via
mbfinotti/advertising-skills@ugc-ad-scripts; keep v3 running while the swap ramps.
ruled out : new concept - stated 6-week production lead time runs past the account's Q4
asset freeze, so it is off this account's ladder, not merely last on it.
Iterate survives (in-house editor, days) and is the fallback if the swap fails.
expected : hook rate and link CTR back toward baseline on the same audience within one
window; CPA follows. Recovery on same audience also confirms fatigue over saturation.
re-check : 2026-09-03 (one full 14-day window after launch)
```
## Example B - B2B, LinkedIn: the audience is saturating, not the creative fatiguing
Cybersecurity vendor, cold ABM motion, matched company list rendering ~46,000 members, 5 sponsored-content ads per LinkedIn's guidance, 11 weeks in. Conversion volume (demo requests) is single-digit weekly, so the conversion floor cannot clear - the call gates on leading engagement signals, and the confidence line says so.
The tell: all five creatives - including one launched fresh 3 weeks ago - decay together, reach has been flat for a month while impressions climb, and CVR degrades alongside CTR. Fatigue is creative-specific and leaves CVR stable; this is pool depletion. The fresh creative's failure to recover performance on the same audience is the discriminating test failing in the saturation direction.
```
FATIGUE VERDICT - "ROI-report-static-A" (pattern shared by all 5 ads), 2026-08-20
platform : LinkedIn | funnel stage: cold prospecting (ABM list)
window : Jul 21-Aug 17 (rolling 7-day) vs baseline May 26-Jun 22 (30-day)
volume : $18,400 spent, 342,000 impressions, 14 demo requests in window
signals
link CTR : 0.31% vs 0.44% (-30%, all 5 ads within 4pts of each other) [leading]
reach (rolling) : flat 4 weeks; impressions +38% [leading]
new-member reach : falling steadily since early July [leading]
frequency : 8.1 vs 4.9 [lagging]
CVR (click->demo) : 1.1% vs 1.6% (degrading WITH engagement) [lagging]
CPL : $1,310 vs $780 (+68%) - low volume, directional only [lagging]
confounder screen
budget/bid change : pass - stable budget, manual bid unchanged
learning-phase reset : pass - no edits inside window (1 sibling added Jul 28, see mix)
audience saturation : FAIL - list rendered ~46k; reach plateaued at ~41k; fresh
creative (Jul 28) decayed to the pack within 2 weeks on the
same audience - the discriminating test points to the pool
auction CPM inflation : pass - CPM +6%, normal for the category
seasonality/window mix : pass - matched windows; B2B summer dip checked vs last year, smaller
tracking breakage : pass - form fills reconcile with CRM
attribution-window skew : pass - lag-mature windows
placement/device mix : pass - feed-only placement
statistical noise : pass - CTR delta outside 2-SE band on 342k impressions
landing page/offer change: pass - no changes; other channels' CVR to page stable
sibling-mix shift : noted, not causal - Jul 28 add redistributed spend but decay predates it
confidence : medium - engagement-gated (14 conversions cannot clear the conversion floor);
signals many and consistent, but CVR/CPL read is directional
verdict : saturating (audience), not fatigued (creative)
action : rung 6, promoted to first under a `saturating` verdict - expand the pool:
refresh/extend the account list, add lookalike-style expansion off
closed-won, and add a frequency cap via the engagement-exclusion
mechanic; audience work via mbfinotti/advertising-skills@ad-audience-targeting.
Do NOT commission new creative - the Jul 28 test already showed it changes nothing.
ruled out : none - no stated constraint deletes a rung here. The creative rungs are ruled
out by the verdict, not by capacity; that is an evidence call, and it reverses
the moment the pool grows again.
expected : new-member reach resumes growth; frequency falls below its Jun level;
CTR recovers only as fresh members enter - not before
re-check : 2026-09-17 (one full 30-day window after list expansion)
```
## Example C - NEGATIVE example: a convincing fatigue call that is actually a budget scale-up
Home-fitness brand, Meta, cold prospecting. The account's best static has run 5 weeks; 12 days ago the team scaled the ad set budget +80% in one step. Now: CTR -14%, CPA +32%, frequency up.
Pattern-matched against a fatigue checklist this "confirms": two signals down, multiple periods, frequency rising. The screen catches it in two lines: the decline starts exactly at the budget step, the whole ad set (all 3 creatives) moved together, and the comparison is being made against the _pre-scale_ peak, which the platform can no longer buy at 1.8x the spend.
```
FATIGUE VERDICT - "before-after-static-hero", 2026-08-20
platform : Meta | funnel stage: cold prospecting
window : Aug 08-20 (3-day MA) vs baseline Jul 25-Aug 07 (trailing 14-day)
volume : $9,700 spent, 407,000 impressions, 221 conversions in window
signals
link CTR : 1.19% vs 1.38% (-14%) [leading]
hook rate : n/a (static)
frequency : 2.6 vs 1.8 [lagging]
CVR : 3.1% vs 3.2% (stable) [lagging]
CPA : $43.90 vs $33.30 (+32%) [lagging]
confounder screen
budget/bid change : FAIL - ad-set budget +80% on Aug 08; decline onset same day;
all 3 sibling creatives degraded in lockstep (fatigue is
creative-specific; scale effects hit the whole ad set)
learning-phase reset : FAIL - the +80% step is a significant edit; delivery volatile
through ~Aug 15; window overlaps the reset
audience saturation : pass - first-time impression ratio 74%, reach growing fast
auction CPM inflation : pass - CPM +5% only; the cost move is CTR-and-mix-driven
seasonality/window mix : pass - matched windows
tracking breakage : pass - platform-to-orders ratio stable
attribution-window skew : partial - last 3 days not lag-mature; CPA overstated at the margin
placement/device mix : FAIL (secondary) - Audience Network share 9% -> 17% post-scale;
within-placement CTR nearly flat
statistical noise : pass - deltas outside noise band (but explained above)
landing page/offer change: pass
sibling-mix shift : pass - same 3 creatives
confidence : high - in the confounder, not in fatigue: three screen lines explain the decline
verdict : confounded (budget scale-up + learning reset + placement mix), NOT fatigued
action : no refresh. Re-baseline the creative at the new spend level from Aug 16 (post-
stabilisation); judge future decay against that baseline, not the pre-scale peak.
Accept that marginal efficiency at 1.8x spend sits below the old average, and
hand the "how fast to scale" question to mbfinotti/advertising-skills@paid-media-scaling.
ruled out : n/a - no rung was selected. A `confounded` verdict acts on the cause and never
enters the Action Ladder, so nothing was deleted from it.
expected : metrics stabilise at a new, slightly worse-than-peak level; if decay then
resumes against the NEW baseline with confounders clean, re-open the fatigue case
re-check : 2026-09-01 (first full clean window on the new baseline)
```
The lesson: commissioning a refresh here would have burned production capacity, reset learning again with the new upload, and "validated" the fatigue call when metrics stabilised for reasons that had nothing to do with the new creative.
references/signal-reference.md›
# Signal Reference
Every signal is read against the creative's own trailing baseline (see SKILL.md workflow step 3), never against an industry number. "Alone?" states whether a move in that signal, by itself, justifies any conclusion. Evidence tiers for the reference points at the bottom: **platform-documented**, **published research**, **practitioner heuristic** (no traceable primary source - always label it when quoting).
Which signals to pull first, when you cannot pull them all:
- diagnostic value per unit of effort to obtain: link CTR > hook rate > spend share within the ad set > first-time impression ratio > hold rate > CPM > CVR > frequency > CPA/ROAS
- effort, lightest first: link CTR == CPM == frequency == CPA (a real tie - all four sit in the same rows of the same default export, so pulling any one of them hands you the other three) > hook rate == hold rate (also a real tie - one video-retention breakdown produces both) > CVR (reconcile tracking first) > spend share (compute per creative from the export) > first-time impression ratio (Meta breakdown, or a new-user-reach proxy elsewhere)
The two orders disagree on purpose: the numbers that cost nothing to pull are the ones that mislead alone, which is what the two-signal rule exists to stop. Re-rank for the account before pulling anything - no video deletes hook and hold rate outright, a B2B account with single-digit weekly conversions demotes CVR and CPA below every engagement signal, and a non-Meta account demotes first-time impression ratio to whatever proxy the platform offers.
## Delivery and exposure signals
**Frequency** - average impressions per person reached in the period: `impressions / reach`.
- Class: lagging, confirming. Alone? No - the most misused signal in the field. Meta's own analytics team notes two structural flaws: it is reported at ad/ad-set/campaign level while fatigue happens at creative level, and it is a period average, not the marginal effect of the next impression. A creative at frequency 3.0 can be spent while another at 5.0 is healthy.
- Read: rising frequency plus decaying response confirms exposure pressure; rising frequency with stable response means nothing is wrong yet. Always split cold vs warm - warm audiences tolerate several times the cold-audience exposure.
- Platforms: all major platforms report it; only Meta exposes anything near creative-level granularity via breakdowns.
**Reach vs impressions** - unique people vs total deliveries; their ratio is frequency, but their _trends_ are separately useful.
- Class: leading for saturation. Alone? Directional only.
- Read: impressions climbing while reach flattens means the platform is recycling the same pool - the account is running out of new people before any creative wears out. Rolling month-over-month reach falling is an early saturation tell that moves before frequency looks alarming.
- Platforms: Meta, LinkedIn, TikTok natively; Google Display/YouTube via reach reports; not meaningful on search.
**First-time impression ratio** - share of the period's impressions that are someone's first exposure to the ad.
- Class: leading, primarily for saturation onset. Alone? Good early warning, still needs a response signal beside it.
- Read: falling ratio means deliveries are increasingly repeats. Practitioner interpretation of healthy prospecting sits around 65-80%, with below ~50% read as saturation approaching (Triple Whale; Flighted) - heuristic bands derived from Meta's Delivery Insights metric, not platform rules. On retargeting the ratio is low by design; do not read it there.
- Platforms: Meta (Delivery Insights). Elsewhere, approximate with new-user reach (TikTok reports daily new-user reach directly).
**Spend share within the ad set** - the creative's share of its ad set's spend, trended.
- Class: leading. Alone? Suggestive, needs confirmation.
- Read: under algorithmic delivery, a creative silently losing budget share with no manual change means the platform's own models are deprioritising it - an implicit fatigue read that arrives before your dashboards show it. Also check the mirror image: a sibling launch can take share from a still-healthy creative (that is mix shift, not decay - see the confounder file).
- Platforms: computable everywhere from per-creative spend; no platform labels it.
## Response signals
**Link CTR (outbound CTR)** - `link clicks / impressions`. Use link/outbound clicks, not "all clicks": all-clicks CTR counts reactions, comments, profile taps, and expands, which can hold steady or rise while actual intent collapses - it flatters a dying ad.
- Class: leading - usually the first response signal to move, high-volume enough to be statistically stable daily on most budgets. Alone? The cleanest single signal, but still needs a partner: a dip has many non-fatigue causes (placement mix, one bad day), and a _naturally low_ CTR (weak creative from day one) is not a _declining_ CTR (proven creative losing steam).
- Platforms: all. On Meta explicitly separate link CTR from CTR (all).
**Hook rate (3-second rate)** - `3-second video views / impressions` (Dara Denney's definition; Ben Heath frames it as the share watching past the first 3 seconds).
- Class: leading - on video it moves before clicks, and on TikTok it is typically the first thing to move. Alone? Strong for diagnosing _where_ decay lives: hook rate down with hold rate stable means the opening is tired, not the ad - rung 1 of the action ladder, its highest return per hour spent. Ben Heath reports that for many advertisers over 90% of viewers drop off before the four-second mark, which is why the opening carries so much of the fatigue load.
- Platforms: Meta, TikTok, YouTube, LinkedIn video - computed from 3-second (or platform-equivalent) view counts.
**Hold rate** - viewers still watching at a mid-point checkpoint over those who started; Denney uses viewers reaching 15 seconds; completion quartiles (25/50/75/100%) serve the same role.
- Class: leading/secondary - confirming, not primary. Alone? No.
- Read: hook rate stable with hold rate decaying points at the body/on-ramp, not the opening. Both decaying together is generic wear-out.
- Platforms: any platform with video quartile or watch-time reporting.
**Thumbstop** - static-ad and feed shorthand for the same construct as hook rate: the share of impressions that stop scrolling (3-second views on video, sometimes engagement-based proxies on statics).
- Class: leading. Alone? Same caveat as hook rate - a great thumbstop is not a great ad; clickbait shows high thumbstop with collapsed downstream metrics, so read the whole funnel.
- Platforms: computed, mainly Meta/TikTok vocabulary.
**CPM** - `cost per 1,000 impressions`.
- Class: leading but ambiguous. Alone? Misleading - CPM is set by the auction, not by your creative alone. Rising CPM with flat CTR is auction competition or seasonality (Q4, elections), not fatigue. Diagnostic only when paired with falling CTR: the platform pricing your deliveries up _while_ response falls is consistent with the system downgrading the creative.
- Platforms: all.
**CPC** - `spend / link clicks`. Arithmetic composite of CPM and CTR: CPC rising decomposes into "auction got pricier" (CPM up) or "creative stopped earning clicks" (CTR down). Always decompose before reading it.
- Class: derived. Alone? No - read its components instead.
- Platforms: all.
**Conversion rate (CVR)** - `conversions / link clicks`.
- Class: lagging, and the key _separator_. Alone? Its stability is the information.
- Read:
- Costs rising with CVR stable is the fatigue pattern (people who still click still buy).
- CVR degrading alongside engagement suggests saturation (the remaining pool is lower-intent).
- CVR collapsing while CTR holds is a landing-page/offer/tracking problem that no creative refresh will touch.
- Platforms: all, subject to tracking integrity - screen tracking first.
**CPA / ROAS drift** - `spend / conversions`, `conversion value / spend`.
- Class: lagging - last to move; by the time CPA spikes the decay has been building for days. Alone? Never - low conversion volume makes daily CPA the noisiest number on the dashboard, and attribution lag makes trailing windows under-report recent days structurally.
- Read: confirmatory only, on lag-matched windows, after the leading signals have made the case.
- Platforms: all; B2B accounts should demote it below engagement signals entirely (see SKILL.md, B2B vs B2C).
## Platform rating systems
**Meta delivery statuses - "Creative fatigue" / "Creative limited"** - platform-documented statuses; industry sources (Jon Loomer; AdSights) report the fatigue status firing around a doubling of cost per result vs history, with "Creative limited" the milder tier.
- Class: lagging by construction. Read: useful as a backstop, not a detector - if these fire routinely, detection upstream is too slow.
**Google RSA / PMax asset labels - Learning / Low / Good / Best** - platform-documented, performance-based per-asset ratings. Read: the legitimate freshness signal on Google; replacing persistent "Low" assets is standard practice (the common 2-4 week cadence is practitioner advice, not a Google rule).
**Google Ad Strength (Poor-Excellent)** - measures asset diversity and completeness, **not performance**. Optmyzr's analysis found "Average"-strength ads with the best CPA/CVR and "Poor" with the best ROAS. Never treat Ad Strength movement as a fatigue signal.
## Attributed reference points
Quote these only as attributed starting references; the account's own baseline overrides all of them.
**Platform-documented**
- Meta "Creative fatigue"/"Creative limited" delivery statuses exist (Meta Business Help Center); the ~2x cost-per-result trigger is as reported by Jon Loomer and AdSights, Meta's page being closed to automated verification.
- Meta learning phase: roughly 50 optimization events per week per ad set for stable delivery; significant edits reset it (Meta Business Help Center).
- LinkedIn official guidance: rotate the lowest-engagement ad every 1-2 weeks; run 4-5 ads per campaign (LinkedIn Marketing Solutions, Sponsored Content best practices).
- TikTok guidance: refresh on the order of every 7 days / when delivery trends consistently down.
**Published research**
- Meta 2023 creative-repetition study (Analytics at Meta, Medium): mean exposure 4.2 per creative; ~45% drop in conversion likelihood at 4 exposures; click likelihood decaying as (N+1)^-0.43; adding a new creative in high-fatigue cases produced an average ~8% conversion-rate improvement; **no wear-in effect found for direct-response objectives**.
- Wear-in/wear-out lineage: Pechmann & Stewart (1988) - effectiveness rises, peaks, declines. Les Binet's wear-in effect (via Motion, motionapp.com/blog/why-ads-resist-creative-fatigue): some brand ads gain effectiveness with exposure.
**Practitioner heuristics (label as such every time)**
- Ben Heath: results drop-off often starts around frequency 2.0-2.5 on cold Meta audiences, with warm audiences commonly fine at 10+ (heathmedia.co.uk/facebook-ad-frequency/); "ad fatigue is usually hook fatigue", and over 90% of viewers drop before the 4-second mark for many advertisers (heathmedia.co.uk/scale-meta-ads-faster-by-testing-hooks/).
- AJ Wilcox / B2Linked: LinkedIn's delivery caps exposure by creative count - up to "7 in 48 hours" with 7+ creatives, roughly one per 24h with 1-3 creatives (b2linked.com/blog-page/linkedin-ads-frequency-caps-how-often-can-someone-see-your-ads).
- Baseline methods: 3-day moving average vs the creative's own trailing 14-day baseline (AdSights); rolling 7-day vs 30-day (Segwise). Sustained ~20-25% decline as the action band and 10-20% as watch-list (AdSights et al.) - adjustable defaults, not laws.
- Widely repeated but with no traceable primary source: frequency 2.5-3.0 prospecting flag, 4-6 retargeting tolerance, 20-30% CTR-drop band, ~18-25% CPM inflation, first-time-impression ratio <50%. Treat every one as folklore-grade until the account's own data confirms or replaces it.
SKILL.md›
---
name: ad-creative-fatigue
description: "Decide whether a running ad creative is genuinely wearing out, or whether a confounder - budget change, learning-phase reset, audience saturation, auction CPM inflation, seasonality, tracking breakage - explains the decline, and return a verdict with confidence plus the highest-return remedy per unit of effort. Use whenever the user mentions creative or ad fatigue, wear-out, climbing frequency, dropping CTR, a rising CPA on an older ad, when to refresh creative, or whether to kill an ad - even if they never say 'fatigue'. Covers B2B and B2C across search, social, video, and native. Ends at the verdict: replacement creative belongs to mbfinotti/advertising-skills@ad-copy-variants and mbfinotti/advertising-skills@ugc-ad-scripts."
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.1.4"
---
# Creative Fatigue
Read a creative's performance-over-time data and decide whether it is genuinely wearing out, how confident that call is, and what to do about it. The core discipline is refusal: many things can cause a declining metric, and creative wear-out is only one of them. Most refresh decisions are made too early, off one metric, against no baseline.
This skill runs a differential diagnosis in a fixed order, confounder screen before the fatigue call, never after:
1. Baseline
2. Confounder screen
3. Decay measurement
4. Confidence gate
5. Verdict
Two further tensions stay live throughout:
- Fatigue is often hook fatigue rather than whole-ad fatigue (Ben Heath).
- Some ads improve with repeated exposure, the "wear-in" effect Les Binet describes in brand advertising, though Meta's 2023 research found no wear-in for direct-response objectives - wear-in is a real verdict only where the objective is brand, not DR.
There are no universal thresholds here: every trigger is derived from the creative's own trailing baseline and the account's own history. Named practitioner numbers appear only as attributed starting reference points.
This skill ends at the verdict and the recommended action.
- Producing the replacement asset belongs to `mbfinotti/advertising-skills@ad-copy-variants`, `mbfinotti/advertising-skills@ugc-ad-scripts`, and `mbfinotti/advertising-skills@ad-creative-brief`.
- Sample-sizing a pre-launch test belongs to `mbfinotti/advertising-skills@ad-creative-test-plan`.
- Scoring a video opening before launch belongs to `mbfinotti/advertising-skills@ad-hook-analyzer` (here, hook rate is only read as a decay signal on ads already running).
- When the confounder screen shows the problem is not the creative at all (structure, targeting, tracking), hand off to `mbfinotti/advertising-skills@ad-account-diagnostic`.
- Retargeting-sequence and frequency-cap architecture design is `mbfinotti/advertising-skills@retargeting-funnel`.
- Scaling and pacing decisions after a win are `mbfinotti/advertising-skills@paid-media-scaling` and `mbfinotti/advertising-skills@ad-budget-pacing`.
## Interview
Ask before diagnosing anything. One question per message; offer multiple-choice answers where possible; skip anything already answered or visible in supplied data.
- Which platform(s) is the creative running on? (Meta / Google Ads - search, Performance Max, or YouTube / LinkedIn / TikTok / other)
- B2B or B2C/ecommerce?
- Funnel stage: cold prospecting or retargeting/warm? (Tolerated exposure differs sharply between the two.)
- What can you actually export, and at what granularity: per-creative per-day, per-creative weekly aggregate, or only ad-set/campaign level? Per-creative per-day is the working assumption; anything coarser weakens every step downstream.
- Daily spend and conversion volume on the creative in question? (Decides whether the confidence gate can clear at all.)
- Rough audience size for the ad set, and is the audience list-based/ABM or broad?
- How many creatives are live in the same ad set? (Sibling mix changes what "losing spend share" means.)
- When was the ad, or its parent ad set/campaign, last edited - creative, budget, bid, audience, optimization event? Exact date matters.
- Did budget change during the window under suspicion? By how much?
- By what date does the result have to land? (A hard deadline promotes the Action Ladder's same-day rungs - rotation, budget shift, frequency cap - and demotes iteration and new concepts.)
- Do you want a one-off win on this creative, or a compounding asset? (A compounding mandate promotes iteration and new concepts even where a configuration fix would hold this week.)
- What is your effort ceiling: creative production capacity and lead time, in-house editing hours, whose sign-off list work needs, and how reversible the change has to be? (A refresh recommendation nobody can execute is worthless - this answer deletes rungs from the Action Ladder rather than reordering them.)
- What is the target CPA or ROAS the creative is judged against?
## Workflow
1. Run the Interview; collect every answer before touching the data.
2. Fix the unit of analysis. Analyse at creative and concept level, never at campaign averages: forty rows for forty variants of five concepts hides the real pattern, and one fatigued ad dragging an ad-set average looks like a campaign-wide problem. Where the platform rotates variants automatically (Advantage+-style delivery, dynamic creative), read the asset-level breakdown - automated rotation masks individual asset decay. Conversely, a creative silently losing budget share with no manual change is itself a decay signal.
3. Build the creative's own baseline. Never compare single days.
Sourced methods to offer as starting defaults, picked per the user's data granularity and overridden by their own history - the two are not ranked against each other, since both cost the same single export and only granularity decides:
- 3-day moving average against the creative's own trailing 14-day baseline (AdSights).
- Rolling 7-day window against a 30-day baseline (Segwise).
Exclude known outages and match the comparison window's day-of-week composition to the baseline's. If your harness can read the export or run the calculation, compute it; otherwise emit the exact export steps and spreadsheet formulas (columns, moving-average ranges, delta formula `(current - baseline) / baseline`) for the user to run and report back.
4. Run the confounder screen, before any fatigue talk. Work through [references/confounders.md](references/confounders.md) and mark every confounder pass (ruled out) or fail (present):
- budget/bid change
- learning-phase reset
- audience overlap/saturation
- auction CPM inflation
- seasonality and window composition
- tracking breakage
- attribution-window skew
- placement/device mix shift
- statistical noise
- landing-page/offer change
- sibling-mix shift
Any fail that explains the decline ends the fatigue inquiry: the verdict is `confounded` (or `saturating` for the audience case) and the fix targets the actual cause. Structure/targeting/tracking root causes hand off to `mbfinotti/advertising-skills@ad-account-diagnostic`.
5. Measure decay against the baseline using [references/signal-reference.md](references/signal-reference.md). Apply the field's consensus decision rule: no single metric confirms fatigue - require at least two signals moving together across two or more consecutive periods (Segwise; AdSights). Weight leading signals (link CTR decay, hook-rate decay on video, falling first-time-impression ratio, silent budget-share loss) over lagging ones. Frequency is a lagging, confirming signal only - Meta's own analytics team notes reported frequency is measured at ad/ad-set level while fatigue happens at creative level, and is a period average, not the marginal effect of the next impression. Do not build the call on frequency.
6. Apply the Confidence Gate (below). If it fails, the verdict is `insufficient data` - state what extra spend or days would clear it and stop; refuse to recommend a refresh on noise.
7. Separate fatigue from saturation before finalising. The discriminating logic:
- Fatigue: costs rise while conversion rate holds.
- Saturation: both degrade, alongside a falling first-time-impression ratio and flattening reach.
Where the user can run it, the definitive test is Meta's own:
- A new creative restoring performance on the same audience proves fatigue.
- Performance recovering only on a fresh audience proves saturation.
Their fixes are opposite, new creative versus audience expansion, so the split matters.
8. Issue the verdict from the Verdict Ladder, then pick the remedy from the Action Ladder - top of its efficiency ordering first, re-ranked against the Interview's deadline, one-off-versus-compounding and effort-ceiling answers - and fill one Fatigue Verdict block per creative (shape below; worked versions in [references/examples.md](references/examples.md)).
9. Set the re-check date (one full comparison window after any action) and log the decision for the Measuring section's scorecard.
10. If your harness has persistent memory, memorize the account's baselines, per-creative verdicts, actions taken, and re-check dates, so the next run starts from history instead of re-deriving it.
## The Fatigue Verdict
Deliver one block per creative analysed:
```
FATIGUE VERDICT - <creative id/name>, <date>
platform : <platform> | funnel stage: <cold prospecting | retargeting>
window : <comparison window> vs baseline <baseline definition>
volume : <spend> spent, <impressions> impressions, <conversions> conversions in window
signals
<signal> : <current> vs <baseline> (<+/-x%>) [leading|lagging]
<signal> : <current> vs <baseline> (<+/-x%>) [leading|lagging]
...
confounder screen
budget/bid change : pass | FAIL - <note>
learning-phase reset : pass | FAIL - <note>
audience saturation : pass | FAIL - <note>
auction CPM inflation : pass | FAIL - <note>
seasonality/window mix : pass | FAIL - <note>
tracking breakage : pass | FAIL - <note>
attribution-window skew : pass | FAIL - <note>
placement/device mix : pass | FAIL - <note>
statistical noise : pass | FAIL - <note>
landing page/offer change: pass | FAIL - <note>
sibling-mix shift : pass | FAIL - <note>
confidence : high | medium | low - <one-line basis: gate math, signal count, window length>
verdict : fatigued | saturating | confounded | insufficient data | healthy | wear-in
action : <chosen action-ladder rung + who produces the asset, if any>
ruled out : <every rung the account's stated constraints delete, each with the constraint that deleted it - or "none">
expected : <what should recover, by roughly how much, based on which evidence>
re-check : <date - one full comparison window after action>
```
## Confidence Gate
No verdict may be issued below the data floor, and the floor is the user's own math, not a copied constant:
- **Noise check.** For a rate signal (CTR, hook rate, CVR) the observed decline must exceed sampling noise. Approximate the baseline rate's noise band as `p ± 2 × sqrt(p × (1-p) / n)`, with `n` the impressions (or clicks, for CVR) in the comparison window. A "decline" still inside that band is noise, whatever it looks like on a chart.
- **Two-signal rule.** At least two signals, at least one of them leading, moving the same direction across two or more consecutive periods (Segwise; AdSights). One metric, one period, never clears the gate.
- **Conversion floor.** For CPA/ROAS-based claims, enough conversions in both baseline and comparison windows that the delta survives the same noise check on CVR. Low-conversion accounts (most B2B) should gate on leading engagement signals instead and say so in the verdict's confidence line.
When the gate is not met: report `insufficient data`, compute the additional days or spend needed (days ≈ what it takes for `n` to make the noise band narrower than the observed delta at current traffic), and refuse to recommend a refresh. Killing a creative on an underpowered read is the most expensive false positive this skill exists to prevent.
## Verdict Ladder
Each verdict maps to its own action - they never all collapse to "make new creative". These verdicts are deliberately not ranked against each other: the evidence picks exactly one, so an efficiency ordering over them would be false precision. The ranking happens one level down, inside the Action Ladder.
| Verdict | Meaning | Action |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `fatigued` | Two-plus signals decayed vs own baseline, confounders ruled out, gate cleared | Action Ladder, highest return per unit of effort first - rung 1 unless the evidence rules it out |
| `saturating` | The audience pool is depleting - falling first-time-impression ratio, flattening reach, CVR degrading with costs | Audience expansion, fresh seed, or exclusions - not new creative; Action Ladder rung 6, promoted to first under this verdict; see `mbfinotti/advertising-skills@ad-audience-targeting` |
| `confounded` | A non-creative cause explains the decline | Fix the actual cause; hand structural/tracking causes to `mbfinotti/advertising-skills@ad-account-diagnostic` |
| `insufficient data` | Confidence Gate not met | Report the gap, state days/spend needed, re-check then; no refresh |
| `healthy` | Signals inside the noise band of the creative's own baseline | Leave it alone; note next scheduled review |
| `wear-in` | Performance improving with exposure - plausible for brand objectives (Les Binet); Meta found no wear-in for direct response, so treat a DR "wear-in" read with suspicion | Leave it alone; do not rotate it on a frequency number |
## Action Ladder
Ranked by return per unit of effort - what each rung buys, against what it costs you to ship it. Not by price: the cheapest rung and the one worth doing first are rarely the same rung, and "buys a week" is not the same purchase as "restores the ad".
- efficiency (the default order, and the numbering below): hook swap > rotate & rest > budget shift > iterate the winner > frequency cap > audience expansion > new concept > pause
- effort, lightest first: rotate == budget shift == pause > frequency cap > hook swap > audience expansion > iterate > new concept - the three-way tie is a real equality, not indecision: each is one console change, produces no asset, needs no sign-off, and is reversible the same day
- durability of what it buys: new concept > iterate > audience expansion > hook swap > frequency cap > rotate > budget shift > pause - a concept is the parent of every iteration run off it, so it outlives any single one of them
- compliance cost, heaviest first: audience expansion (customer-list and ABM uploads need a lawful-basis attestation and a data-processing review, and an uploaded list is not easily un-shared) > every other rung (near-zero - all are platform settings or new assets)
Default to rung 1 and take the highest-placed rung the evidence supports. Move one rung down the list when the rung above has already been tried and the decay resumed, or when the signal pattern rules it out - hook, hold rate and CVR decaying together means the whole ad is spent, not its opening.
**What the efficiency order starves.** New concept is high on value and high on effort, so a ratio buries it at rung 7 every round - and it is the only rung that fixes an exhausted concept, which no hook swap or iteration can touch. Promote it above everything else when iteration stops recovering performance across successive attempts on the same concept: that pattern says the concept is spent rather than its execution, and every cheaper rung above is then buying nothing. Audience expansion has the same value-and-effort shape but is not starved here, because the `saturating` verdict already promotes it to rung 1 by rule - its failure mode is the reverse, reaching for it under `fatigued`, where it wastes a good audience.
Every ordering above is a default, not a law: it shifts with the account's context and with who executes it. Re-rank against what you already know about this account, and against the Interview's deadline and mandate answers:
- An always-on challenger bench promotes rotation to first.
- An in-house editor who can ship a new opening the same day keeps the hook swap ahead of every configuration fix.
- A `saturating` verdict overrides the whole order: audience expansion becomes rung 1 and creative work is wasted effort.
- A hard deadline promotes the same-day rungs.
- A compounding mandate promotes iteration and new concepts.
The effort ceiling does something different: it **deletes** rungs from this account's ladder rather than demoting them.
- No production capacity, or a lead time longer than the deadline, deletes new concept and iterate.
- No editor deletes the hook swap.
- A locked or empty creative library deletes rotation.
- No lawful-basis sign-off deletes audience expansion.
Name each deleted rung and the constraint that deleted it on the verdict's `ruled out` line, then take the highest-ranked survivor. A rung merely parked at the bottom of the list is still on the list, and comes back later as scope nobody agreed to fund.
1. **Hook/thumbnail swap on the same body**
- Effort: an hour or two of editing, no brief cycle.
- Buys: the ad back on the same audience, for as long as the body holds.
- Right when: hook rate or thumbstop decayed but hold rate and CVR held - the opening is tired, not the ad (Ben Heath's "ad fatigue is usually hook fatigue").
- Caution: if negative-feedback rate is elevated, a hook swap on the same creative ID does not clear the algorithmic penalty; ship a new asset ID.
2. **Rotate from existing inventory and rest the creative**
- Effort: near-zero, configuration only.
- Buys: a window while the audience cycles, not a fix - the decay waits where you left it.
- Right when: the account keeps an always-on challenger bench; worth nothing without one.
- Note: reintroduce the rested creative after the pool has turned over, and promote a proven challenger meanwhile.
3. **Budget shift to a healthier creative**
- Effort: near-zero, reversible the same day.
- Buys: time and CPA protection while a replacement ramps.
- Note: keep a still-profitable fatigued ad running; never pause a producer with nothing staged. Pairs with any other rung rather than competing with one.
4. **Iterate the winner**
- Effort: days, plus a production brief and a test slot.
- Buys: the largest durable payoff on the ladder - the same proven concept in a new execution (new hook, opening seconds, format, aspect ratio, or copy) running for weeks.
- Evidence: practitioners report element-level iteration often recovers most of original performance without a full rebuild (Hawky's reported band is 60-80% - treat as their number, not a promise).
- Reference: brief production via `mbfinotti/advertising-skills@ad-copy-variants`, `mbfinotti/advertising-skills@ugc-ad-scripts`, or `mbfinotti/advertising-skills@ad-creative-brief`.
5. **Frequency cap or exclusion**
- Effort: configuration, plus agreement on who owns the cap.
- Buys: relief on one over-exposed warm pool, and nothing at all on a cold prospecting decay.
- Right when: exposure is concentrating on a warm pool that has seen it enough.
- Reference: design of the cap architecture itself belongs to `mbfinotti/advertising-skills@retargeting-funnel`.
6. **Audience expansion or fresh seed**
- Effort: configuration plus list work and, for customer-list or ABM uploads, a lawful-basis sign-off - days to a week, and hard to walk back once the list is uploaded.
- Buys: a new pool, the only thing that moves a `saturating` verdict.
- Caution: applying it to true fatigue wastes a good audience; under `saturating` it is rung 1.
7. **New concept**
- Effort: weeks, full production and a test plan.
- Buys: a fresh line of assets when the concept, not the execution, is exhausted - the lowest hit rate per attempt on the ladder, and the highest ceiling.
- Reference frame: a commonly cited portfolio split is 70% proven / 20% tests / 10% experiments; treat it as an operator heuristic to adapt, not a rule.
- Reference: test-plan the launch with `mbfinotti/advertising-skills@ad-creative-test-plan`.
8. **Pause**
- Effort: near-zero and instantly reversible.
- Buys: a stopped loss and nothing else - the creative's remaining contribution goes with it.
- Caution: last resort, and only with a replacement live or the budget re-homed. Pausing does not reset platform learning; editing a live ad does.
## Per-Platform Notes
Pointers only - metric mechanics and platform quirks live in [references/signal-reference.md](references/signal-reference.md).
- **Meta**: native "Creative fatigue" and "Creative limited" delivery statuses exist but lag badly (reported to fire around a doubling of cost per result, per Meta docs as relayed by industry sources) - if they are firing, detection was already too slow. First-time impression ratio is native. Advantage+ shifts budget away from fatigued assets silently, masking the signal; read per-asset breakdowns. Any significant edit resets the learning phase.
- **Google Ads**: on RSAs, asset performance labels (Learning/Low/Good/Best) are the performance-based signal. Ad Strength is not - it measures diversity and completeness, and Optmyzr's analysis found "Average"-strength ads with the best CPA/CVR; never read Ad Strength as fatigue. Performance Max is opaque at combination level - work from the per-asset report. YouTube: view-rate decay and per-user frequency are the levers.
- **LinkedIn**: official guidance is to rotate the lowest-engagement ad every 1-2 weeks and keep 4-5 ads per campaign. Delivery mechanics tie exposure to creative count (AJ Wilcox's observed "7 in 48" pattern), so a small creative set throttles itself. Small B2B audiences accumulate frequency slowly but relentlessly against expensive CPMs.
- **TikTok**: fastest wear-out of the major platforms; TikTok's own guidance suggests roughly 7-day refresh cycles, framed around a consistently declining delivery trend. Hook rate moves first; daily new-user reach flags saturation directly; automated creative optimization masks per-asset decay - read asset breakdowns.
## B2B vs B2C
The method (own-baseline, confounder screen, two-signal rule, verdict ladder) is identical in both. What differs is data volume, exposure mechanics, and which signals carry the call.
**B2B:**
- Addressable audiences are small and sometimes platform-throttled, so frequency climbs structurally rather than as a decay symptom.
- Conversion volume is usually too low for the conversion floor; lean on leading engagement signals (link CTR, engagement-rate trend, CPL drift over 3-6 week windows) and say so in the confidence line.
- The long sales cycle means the conversion signal lags the creative signal by weeks. Never read a flat pipeline week as creative failure.
- ABM and list-based audiences saturate by design; expect the `saturating` verdict more often than `fatigued`.
**B2C/ecommerce:**
- Large pools and high conversion volume make the full statistical gate usable.
- Wear-out runs fastest on narrow retargeting segments.
**Both:** warm/retargeting audiences tolerate far higher exposure than cold. Ben Heath reports Meta results commonly holding at frequency 10+ on warm audiences versus a drop-off he sees around 2.0-2.5 on cold. Always split cold from warm before reading any exposure number.
## Measuring Whether This Worked
The skill's own KPI is decision quality, tracked on a rolling log of every verdict:
- **Refresh win rate**: share of fatigue-triggered replacements that beat the retired creative's pre-decline baseline (same audience, full comparison window) at the re-check date.
- **False-positive rate**: share of retired creatives that were, in retrospect, still healthy - the replacement did no better, or the "decay" reversed on its own in a holdout.
As a starting floor (this skill's practical target, not a researched constant), iterate until at least 60% of refreshes beat the retired creative and under 20% of retirements were false positives. Tighten both from the account's own history once a dozen decisions are logged.
- A low win rate with a clean gate usually means the diagnosis is right but production quality is the constraint.
- A high false-positive rate means the gate or the confounder screen is being skipped.
If refreshes stop recovering performance across several concepts at once, stop refreshing: that pattern is an offer, landing-page, or product problem no new creative will fix.
## Common Failure Modes
| Trap | Why it burns | Fix |
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Calling fatigue from one metric, one bad day | Noise and weekends mimic decay; retires healthy creatives | Two signals, two-plus periods, vs the creative's own baseline |
| Reading a learning-phase reset as decay | Any significant edit restarts volatile delivery for days | Check last-edit date first; wait a full window post-edit before judging |
| Treating a frequency number as causal | Frequency is lagging, ad-set-level, and an average - a creative at 3.0 can be spent, another at 5.0 healthy | Frequency confirms; link CTR, hook rate, and first-time-impression ratio lead |
| Campaign averages hiding per-creative decay | One fatigued ad drags the average; the rest get refreshed for nothing | Per-creative, per-concept reporting; asset breakdowns under auto-rotation |
| Blaming creative for Q4/auction CPM inflation | CPM up with CTR flat is the auction, not the ad | CPM is diagnostic only when paired with falling CTR |
| Reading CVR collapse with healthy CTR as fatigue | That pattern is landing page, offer, or tracking - no creative will fix it | CTR healthy + CVR down → funnel check before creative check |
| Trailing-window "decay" from attribution lag | Recent days always under-report conversions; every trailing window looks like decline | Compare windows of equal conversion-lag maturity |
| Fixing saturation with new creative | Opposite remedies: fatigue wants creative, saturation wants audience | Run the fatigue-vs-saturation split (step 7) before acting |
| Editing the live ad to "refresh" it | Edits reset learning and destroy the baseline mid-measurement | Launch new ads alongside; pause, never edit, a measured creative |
| Treating Google Ad Strength as a fatigue signal | It scores diversity/completeness, not performance | Use asset performance labels; ignore Ad Strength for this call |
| Hook-swapping a creative with elevated negative feedback | The algorithmic penalty rides the asset ID, not the hook | Ship the iteration as a new asset ID |
| Recommending refresh volume beyond production capacity | An unexecutable plan defaults to letting everything decay | Delete every Action Ladder rung beyond the Interview's effort ceiling, name each on the verdict's `ruled out` line, then take the highest-ranked survivor |
| Retiring a producer with nothing staged | The ad set goes dark or re-enters learning on a gap | Stage the replacement, ramp it, then retire |
## Reference
- `mbfinotti/advertising-skills@ad-copy-variants` - to produce replacement assets for a `fatigued` verdict
- `mbfinotti/advertising-skills@ugc-ad-scripts` - to produce replacement assets for a `fatigued` verdict
- `mbfinotti/advertising-skills@ad-creative-brief` - to produce replacement assets for a `fatigued` verdict
- `mbfinotti/advertising-skills@ad-creative-test-plan` - for pre-launch test design on replacement creatives
- `mbfinotti/advertising-skills@ad-account-diagnostic` - for structural/tracking root causes when the confounder screen fails
- `mbfinotti/advertising-skills@retargeting-funnel` - for frequency-cap architecture design under a `fatigued` verdict