SKILL DETAIL
ad-creative-test-plan
mbfinotti/advertising-skills/ad-creative-test-plan
Design a pre-launch ad creative test plan - falsifiable hypothesis, isolation level, test cells with per-cell budgets, required sample, spend and duration, and kill/scale rules pre-registered before any money moves. Every cell carries an explicit read standard, so an underpowered screen is never dressed up as an A/B test. Use whenever the user wants to test ads, mentions an A/B or split test on creative, asks how much budget or how long a test needs, or mentions sample size, statistical significance, single-variable vs big-swing testing, or test cell structure - even if they never say 'test plan'. Covers B2B and B2C. Do NOT use to read results from a test already running - use mbfinotti/advertising-skills@ad-creative-fatigue instead.
Installation
npx skills add https://github.com/mbfinotti/advertising-skills --skill ad-creative-test-plan
Fichiers du skill
SKILL.md
Dernière synchronisation · 24 sept. 2026
evals/evals.json›
{
"skill_name": "ad-creative-test-plan",
"evals": [
{
"id": 1,
"prompt": "I run paid social at Fernweh Goods, a DTC travel-accessories brand. We have $15,000/month set aside purely for testing, separate from scaling spend. Current numbers: $20 CPA on purchases, $1.25 CPC, 2.5% click-to-purchase rate, about 700 purchases a month account-wide. We shot two new concepts (a packing-hack demo and a lost-luggage horror story) and want to run a proper statistically significant A/B test against our current best ad to pick which concept gets the holiday production budget. Our marketing lead wants to announce 'the winner at 95% confidence' at the all-hands in 5 weeks. Can you design the test?",
"expected_output": "A pre-launch test plan with three concurrent fixed-budget cells, feasibility math showing the purchase read cannot reach significance inside the window, an explicit Directional-read verdict instead of a promised 95% winner, and pre-registered decision rules.",
"files": [],
"expectations": [
"Computes the stable-delivery minimum daily budget per cell of roughly $143/day (target CPA $20 x 50 / 7) and checks the proposed cells against it",
"Includes the current best ad as a control cell running concurrently at its own fixed budget, not as a historical baseline",
"Computes required sample per cell for a purchase-rate significance read at 80% power and alpha 0.05 (exact two-proportion formula or the n = 16 x p(1-p) / MDE^2 shortcut), showing the number",
"Shows with arithmetic that reaching significance on the purchase metric at any moderate MDE requires more days than the test window allows",
"Labels the purchase-metric read with the explicit verdict 'Directional read' (a screening heuristic), not a significance-tested winner",
"Explicitly tells the user the '95% confidence winner in 5 weeks' promise cannot be honored on the purchase event, rather than designing a plan that pretends it can",
"Applies a duration floor of one full week (day-of-week coverage) and a ceiling of roughly 4-6 weeks tied to creative freshness",
"Pre-registers a kill threshold and a scale threshold per asset or cell before launch",
"Pre-declares the inconclusive outcome as keep the control/champion, decided before launch rather than at readout",
"States a fixed stopping rule as a date or a sample size, whichever comes first",
"Writes a falsifiable hypothesis containing a predicted direction, magnitude, named metric, audience, and date or sample, rather than 'see which wins'",
"Raises powering the read on a higher-volume up-funnel event (such as add-to-cart), or asks for that event's rate, instead of only accepting the underpowered purchase read",
"Instructs turning automated or AI creative-optimization features off inside the test cells"
]
},
{
"id": 2,
"prompt": "I do demand gen at Veyra Systems, a B2B procurement software company - roughly $45k ACV and a 4-month sales cycle. We get demo-request form fills from LinkedIn at about $120 each, and roughly 30% of those get accepted by sales as qualified. Test budget is $9,000/month. Our CRO wants us to test a 'CFO cost-cutting' angle against our current 'automation time-saving' ads and have a statistically significant winner on form fills in three weeks, because form fills are the number we report upward. Lay out the test for me.",
"expected_output": "A B2B screening plan judged on CRM-qualified leads over several weeks with a declared Directional verdict, delivery-limitation flagged with the floor math, a quality guardrail, and a scheduled second look one sales cycle later.",
"files": [],
"expectations": [
"Sets the primary decision metric to a CRM-fed pipeline-quality event (qualified/accepted lead), explicitly rejecting raw form fills as the decision metric",
"Treats cheap form fills that sales rejects as a guardrail breach, not a win",
"Declares the verdict for this structure as 'Directional read' and states it in the plan, refusing the statistical-significance framing",
"Computes the stable-delivery floor on the form event at roughly $857/day per cell ($120 x 50 / 7) and shows the affordable cells sit far below it",
"Flags the cells as delivery-limited, meaning their numbers are unstable regardless of sample math",
"Extends the judgment window to several weeks (roughly 4-6), pushing back on the three-week deadline for a qualified-lead read",
"Schedules a second look roughly one sales cycle (~4 months) after the creative verdict for pipeline or opportunity outcomes, instead of treating the in-window read as final",
"Adds a lead-quality guardrail: lead-to-qualified rate compared against the account's own trailing median",
"Keeps the current 'automation' champion running concurrently as a control cell",
"Pre-registers kill and scale thresholds and a fixed stop before launch",
"Places the hypothesis magnitude on the qualified-lead proxy metric and says the revenue read lags by a sales cycle",
"Describes the readout as judgment-based relative ranking, with no '95% confidence' claim anywhere",
"Does not resolve the budget shortfall by assuming a budget increase; works within $9,000/month or names consolidation/up-funnel levers explicitly"
]
},
{
"id": 3,
"prompt": "Growth lead at Peakform Nutrition here (DTC protein snacks). We paused our long-running bestseller ad last month when it was sitting at $18 CPA in Q2, and now we have three fresh concepts from our agency. My plan: put all three into one campaign with campaign budget optimization on and the platform's AI creative enhancements enabled so the algorithm finds the winner efficiently, then compare whichever wins against the champion's Q2 CPA of $18 to decide if we beat it. $450/day for two weeks. Anything you'd change?",
"expected_output": "A restructured plan that revives the champion as a concurrent control, replaces automatic budget allocation with manual fixed-budget cells, switches AI creative features off, and fixes the below-floor per-cell budget with a named lever.",
"files": [],
"expectations": [
"Rejects comparing new cells to the champion's historical Q2 CPA and requires the champion to run concurrently as a control cell in the same structure and budget",
"Gives the reason for the concurrent control: unequal delivery history and/or seasonality make old numbers incomparable",
"Rejects campaign-level automatic budget allocation as a test structure",
"Explains the allocation failure concretely: spend concentrates on early leaders (up to roughly 90% on one cell) before the others collect data",
"Recommends manual fixed-budget cells (or a platform-native deterministic split) with a stated fixed budget per cell",
"Instructs turning the automated/AI creative-optimization feature off inside the test cells",
"Notes the auction-overlap caveat of manual cells: cells still compete in the same auctions against overlapping audiences",
"Computes the stable-delivery floor of roughly $129/day per cell ($18 x 50 / 7) and flags that $450/day split four ways falls below it",
"Resolves the floor breach with a named lever (fewer cells, sequential concepts, up-funnel read, or budget change) rather than shipping four delivery-limited cells silently",
"Emits an explicit read-standard verdict per cell (Powered, Directional read, or Not testable as designed)",
"States with arithmetic what claim the two-week window can support on the purchase metric (screening/Directional versus Powered)",
"Pre-registers kill and scale rules and pre-declares inconclusive as keep the champion"
]
},
{
"id": 4,
"prompt": "I'm a very data-driven founder at Cartelle, a made-to-order furniture brand - $6k/month total ad spend, roughly $70 CPA. I want a genuinely rigorous testing program: one variable at a time, everything else held constant. My backlog: 'Shop Now' vs 'Shop the collection' CTA text, headline with vs without a dash, beige vs off-white background, model facing left vs right, sans-serif vs serif overlay font, and price shown vs hidden. Six clean single-variable tests over the next six months. Long-term I care much more about building reusable learnings than any single winner. Build me the schedule and per-test budgets.",
"expected_output": "A refusal of the micro-variable program with the feasibility math showing why, a redirect to high-leverage levers under a bundled-concept posture with the unlearnable label, and the named condition under which strict isolation gets promoted later.",
"files": [],
"expectations": [
"Refuses to schedule the six micro-variable single-variable tests as proposed, naming items like font, background hue, dash punctuation, CTA wording, and model direction as micro-variables not worth isolated spend at this budget",
"Folds the micro-variations into concept-level executions or drops them, instead of testing them individually",
"Names the high-leverage levers worth isolating (concept, angle, hook, format, creator/talent) and redirects the program there",
"Computes the stable-delivery floor of roughly $500/day per cell ($70 x 50 / 7) and shows $6,000/month (~$200/day total) cannot fund even a compliant two-cell test on the purchase event",
"Rejects strict isolation on feasibility grounds despite the user's stated compounding-learning preference - the preference alone does not promote isolation without the power to read the lever",
"Presents the bundled-concept posture as the default at this budget, because it is the only posture producing a reliably detectable effect at normal budgets",
"Labels any bundled result 'unlearnable at element level' (or an equivalent explicit label) so no element-level insight is mined from it later",
"States the isolation trade-off explicitly (bundled buys detectability; strict isolation buys element-level learning only when powered) rather than treating single-variable testing as inherently more rigorous",
"Names the condition that would promote strict isolation later (feasibility returning Powered on the lever's own MDE, e.g. after volume grows)",
"Emits an explicit per-cell verdict (Powered, Directional read, or Not testable as designed) for whatever structure it does propose",
"Keeps one concept per cell with roughly 3-6 assets per cell in the proposed structure",
"Pre-registers kill and scale thresholds and writes a falsifiable hypothesis with direction and magnitude, not 'see which performs'"
]
},
{
"id": 5,
"prompt": "Quick one - I run ads for Brightlark, a meditation app: $30 CPA, about $18k/month spend, and I'm solo, checking the account maybe twice a week between everything else. Launching 6 new ads next Monday and I need the go/no-go rules written down before then, since we pick the winning batch in exactly 4 weeks for an App Store feature window. I've seen Barry Hott's rule everywhere - spend 1x CPA before judging and 3x CPA before killing - and I was going to combine that with cost caps so the platform kills losers automatically, plus a hard rule to cut anything at 2x CPA spend with no install. Formalize this into our kill rules?",
"expected_output": "A corrected, single-anchor decision-rule set: the folklore attribution debunked, Hott's real position stated, incompatible anchors deleted by name against the user's constraints, and Denney's defaults adopted with the numbers carried through.",
"files": [],
"expectations": [
"Corrects the attribution: the 'spend 1x CPA before judging / 3x CPA before killing' rule is untraceable folklore with no identifiable originator, not Barry Hott's",
"States Hott's actual published position: he rejects mechanical per-ad CPA kill rules (ad-level CPA/ROAS irrelevant) and judges new ads by comparative benchmarking against the account's best ads",
"Refuses to combine multiple kill-rule anchors and states the plan must adopt exactly one",
"Gives the reason mixing fails: the anchors contradict each other, so a blended rule never fires or fires twice",
"Rules out the no-manual-kill / cost-cap-only approach (Faris) by name, tied to the hard 4-week deadline because it is the slowest to answer",
"Rules out Hott's comparative benchmarking by name, tied to the missing weekly reviewer and absent best-ads library",
"Adopts Dara Denney's defaults as the single anchor, identified as the default when nothing promotes another",
"Carries Denney's sizing: test budget approximately CPA x 50 (about $1,500 at a $30 CPA) and roughly 6 assets per test",
"Carries the no-evaluation-before-day-3 rule",
"Carries the asset kill at 2x CPA spend ($60) with no conversion, and the cell kill after 5-7 days with no winner",
"Carries the scale rule: winners scaled +50-100%, done 2-3 times",
"Pre-registers a fixed stopping rule and the inconclusive path (keep the current champion/batch) before launch",
"Never presents the 1x/3x folklore thresholds as authoritative anywhere in the output"
]
},
{
"id": 6,
"prompt": "We're Nordvik Outdoor, a DTC hiking-gear brand. Launching a creative test Thursday: two new video concepts vs our current winner, $250/day per cell, $45 CPA. My reading plan: I check every morning, and the moment any variant is ahead at 95% significance I pause the others and shift budget over - realistically that's day 3 or 4, our tests usually reach significance fast. Also, we'll launch the two new ones Thursday and add the control the following Monday since it needs new post IDs. Write this up as the official plan so the team follows it.",
"expected_output": "A plan that replaces daily stop-at-significance checking with a pre-registered fixed stopping rule and earliest evaluation moment, launches all cells simultaneously, flags the below-floor budgets, and explains why past 'fast significance' was an artifact.",
"files": [],
"expectations": [
"Rejects the stop-at-first-95% daily-checking rule and names peeking/early stopping as the failure being prevented",
"Explains why peeking fails: stopping on a favorable early number inflates false positives, and early data is systematically noisy",
"Replaces it with a fixed stopping rule pre-registered as a date or a sample size, whichever comes first",
"Sets an earliest evaluation moment (day 3 or later plus a minimum volume floor) before which no judgment is made",
"Applies the one-full-week duration floor for day-of-week coverage and flags a day 3-4 decision as below it",
"Requires all cells including the control to launch at the same time, rejecting the Thursday/Monday stagger",
"Ties the stagger rejection to its reason: cells launched at different times pick up day-of-week/seasonality contamination and become incomparable",
"Flags the novelty effect: new creatives get an early boost that fades, so short reads over-credit new variants",
"Reframes 'our tests usually reach significance fast' as an artifact of unadjusted peeking rather than validating it",
"Pre-declares the inconclusive outcome as keep the control",
"Computes the stable-delivery floor of roughly $321/day per cell ($45 x 50 / 7) and flags the $250/day cells as delivery-limited or fixes them with a named lever",
"Emits an explicit per-cell read-standard verdict instead of implying significance is reachable",
"Pre-registers kill and scale thresholds so no real-time judgment call is needed during the run"
]
},
{
"id": 7,
"prompt": "Creative strategist at Vantail, a DTC apparel brand. Five new videos go into a test next week. Purchases are slow to accumulate, so I want hook rate as the deciding metric - whichever video holds the best 3-second view rate takes the scale budget. I read that anything above a 30% hook rate is strong, so that's our bar. Two of the five will run mostly in Reels placements, the other three across feed. Budget is comfortable: $400/day per cell, $28 CPA. Draft the test plan and the success criteria.",
"expected_output": "A plan with a three-layer metric ladder where hook rate is demoted to a gate that only kills losers, thresholds set from the account's trailing median with like-for-like placement comparison, a conversion-level primary metric, and the feasibility math run despite the comfortable budget.",
"files": [],
"expectations": [
"Refuses to make hook rate the primary or deciding metric; assigns it as a gate metric used only to kill obvious losers early and cheaply",
"States that gate metrics screen and never crown winners",
"Cites the evidence class: a multi-account analysis of roughly $1.47M in ad spend found no statistically significant correlation between hook rate and revenue",
"Notes two ads at an identical hook rate can differ several-fold in return, because the gate says nothing about who survives to the call to action",
"Rejects the published 30% bar and sets the gate threshold from the account's own trailing median, noting published 'good' bands conflict (roughly 18% to 40%)",
"Requires like-for-like gate comparison by placement, flagging that comparing Reels-heavy assets against feed assets on raw hook rate crowns accidental winners",
"Names a conversion-level primary metric (purchase CPA or conversion rate) as the decision metric",
"Adds guardrail metrics that must not degrade (e.g. frequency, cost inflation vs account baseline, refund/return rate)",
"Notes that high hook rate combined with low conversion usually signals an offer or landing-page problem, not a creative winner",
"Runs the feasibility check anyway: projects conversion volume per cell (about 14/day at $400/day and $28 CPA) and emits an explicit Powered-or-Directional verdict for the purchase read",
"Reads the gate only after a minimum impression volume per asset (on the order of 1,000-2,000 impressions)",
"Pre-registers kill and scale rules and a fixed stop before launch"
]
},
{
"id": 8,
"prompt": "Merrow & Finch sells premium mattresses online - $180 CPA on purchase. Marketing gave me $500/day for testing, split across 4 cells (3 new concepts plus control). Finance has already rejected our budget-increase request twice this quarter, so that door is closed. My workaround: run the test 10 weeks instead of 4 so the sample accumulates. We do have a solid add-to-cart event, about $12 per add-to-cart, verified firing correctly last month - not sure that matters. Design the plan.",
"expected_output": "A plan that rejects the 10-week stretch against the duration ceiling, declares the purchase cells Not testable as designed with the floor math, deletes the budget lever by name, moves the read to the verified add-to-cart event, and records both read standards.",
"files": [],
"expectations": [
"Rejects the 10-week duration, citing the roughly 4-6 week ceiling because novelty decay, fatigue, and external events contaminate longer windows",
"Computes the stable-delivery floor on purchase of roughly $1,286/day per cell ($180 x 50 / 7) and shows the $125/day cells sit an order of magnitude below it",
"Declares the purchase-event cells with the named verdict 'Not testable as designed' rather than shipping them with hedges",
"Applies the ranked fix levers with the up-funnel read first (up-funnel > widen MDE > fewer cells > raise budget), not more budget or more time by default",
"Deletes 'raise the budget' from the menu by name because finance already refused (no budget authority), instead of listing it as a fallback option",
"Moves the optimization and read to the add-to-cart event, whose floor of roughly $86/day per cell ($12 x 50 / 7) the $125/day cells clear",
"Records both reads in the plan: the add-to-cart read at its computed standard AND the purchase metric explicitly labelled Directional",
"Notes the up-funnel lever is available because the add-to-cart event is verified (an unverified or broken event would delete that lever until tracking is cleared)",
"States the stable-delivery requirement of roughly 50 optimization events per cell per week",
"Frames the delivery floor as a read-validity requirement, not a performance unlock (exiting the learning state is worth only a modest ~5-10% efficiency gain)",
"Keeps the control cell concurrent and pre-registers kill and scale thresholds and a fixed stop",
"Includes a one-sentence decision statement and a falsifiable hypothesis with direction and magnitude"
]
},
{
"id": 9,
"prompt": "Solent & Co, DTC eyewear. Board meeting in 12 days and the CEO wants 'a scientifically clean creative test' presented - so I'm setting up the ad platform's native A/B test tool for our two new concepts, since deterministic assignment means no contamination and the result will be causal proof of which creative is better. $150/day total test budget, $32 CPA. Help me configure the cells and the writeup for the board.",
"expected_output": "A plan that steers away from the native split under the hard deadline, corrects the causal-proof claim with the divergent-delivery evidence, uses manual fixed-budget cells, computes the floor breach at this budget, and states what claim the board can actually be given.",
"files": [],
"expectations": [
"Recommends against the platform-native deterministic split for this test, tied to the hard 12-day deadline (splits take a week or more longer to answer)",
"Notes that native splits at typical budgets are underpowered and frequently return 'no winner'",
"Recommends manual fixed-budget cells instead and states the structure ranking (manual fixed-budget > native deterministic split > campaign-level automatic allocation)",
"Corrects the causality claim: even a deterministic split suffers divergent delivery - each ad is served to a distinct, undetectably optimized mix of users",
"Marks the severity of divergent delivery: it can confound the magnitude and even flip the sign of the result (Braun & Schwartz 2025 or equivalent attribution)",
"States that in-platform reads, splits included, are relative screening rather than causal proof, and that holdout/lift designs are what buy causality",
"Refuses to promise the board a causal or significance-tested winner in 12 days; the plan states exactly what claim the read supports",
"Computes the stable-delivery floor of roughly $229/day per cell ($32 x 50 / 7) and flags $150/day total split across cells as far below it",
"Emits the named verdict for the as-proposed structure ('Not testable as designed' on purchase, or Directional only after a named fix)",
"Fixes the floor breach with a ranked lever (up-funnel read, wider MDE, or fewer cells) rather than silently shipping delivery-limited cells",
"Asks whether a current champion exists to serve as the control cell, or includes a concurrent control",
"Instructs turning automated creative-optimization features off inside test cells",
"Pre-registers the stopping rule, earliest evaluation moment, and kill/scale thresholds before launch"
]
},
{
"id": 10,
"prompt": "Head of growth at Tindra Home (smart lighting). Our freelancers delivered 9 ads: three are variations on a 'movie night' scene idea, four riff on 'wake up naturally', and two take a security/away-mode angle. I was going to spread them across three ad sets, mixing them up so each ad set gets one of each type for fairness, name the files as they come in like tindra_final_v2_new1.mp4, and launch to see what happens. $220/day per ad set, $24 CPA, decision in two weeks. Structure this properly?",
"expected_output": "A restructured plan with one concept per cell, a decision statement and falsifiable hypothesis, a structured variant-naming convention, the multiple-comparisons correction surfaced, the feasibility math run, and pre-registered decision rules.",
"files": [],
"expectations": [
"Rejects mixing concepts within a cell and restructures to one concept per cell (a movie-night cell, a wake-up cell, a security cell)",
"Gives the reason: more concepts than cells contaminates the read - results stop being attributable to any concept",
"Keeps roughly 3-6 assets per cell and flags the 2-asset security concept as below the band (produce another asset or handle the shortfall explicitly)",
"Writes the one-sentence decision statement ('Based on this test, we will ___') before anything else",
"Rejects 'launch and see what happens' and writes a falsifiable hypothesis containing direction, magnitude, named metric, audience, and a date or sample",
"Replaces the ad-hoc file naming with a structured variant convention: fixed field order, one delimiter, encoding concept/angle and other tested levers plus a version",
"States the mechanical reason naming matters: reporting tools parse names, and a naming failure silently destroys the roll-up by dimension",
"Flags the multiple-comparisons problem of three concept cells against control: some variant wins by chance",
"Quantifies or corrects it: roughly 30-40% more sample per cell at ~3 variants under a standard alpha-division correction, or fewer cells - and at minimum discloses the number of comparisons",
"Adds a concurrent control cell (current champion) or asks whether one exists",
"Computes the stable-delivery floor of roughly $171/day per cell ($24 x 50 / 7) and projected weekly conversions (~60-65 per cell), confirming stable delivery",
"Emits an explicit read-standard verdict per cell for the two-week window (Directional versus Powered, with the sample arithmetic) instead of implying a significant winner",
"Pre-registers kill and scale thresholds, an earliest evaluation moment, a fixed stop, and inconclusive resolving to keep the control"
]
}
],
"trigger_queries": [
{ "query": "How do I set up an A/B test for my Facebook ad creatives before launch?", "should_trigger": true },
{ "query": "Design a creative test plan for 3 new video concepts on Meta", "should_trigger": true },
{ "query": "How much budget do I need to test two ad concepts against each other?", "should_trigger": true },
{ "query": "What sample size do I need to call a winner between two ads?", "should_trigger": true },
{ "query": "How long should I run a creative split test before deciding?", "should_trigger": true },
{ "query": "Should I test one variable at a time or make big swings in my ad creative?", "should_trigger": true },
{ "query": "Help me structure test cells for our Q4 creative testing", "should_trigger": true },
{ "query": "Is $5k enough to get statistical significance on an ad test?", "should_trigger": true },
{ "query": "We're launching 5 new statics next month - how do I decide which one wins fairly?", "should_trigger": true },
{ "query": "My CMO wants a 95% confidence winner between our two hero videos. What do I need?", "should_trigger": true },
{ "query": "single variable vs bundled creative testing - which should we do", "should_trigger": true },
{ "query": "How many conversions per variant before an ad test result means anything?", "should_trigger": true },
{ "query": "Set up a controlled experiment comparing our new UGC concept to the current champion ad", "should_trigger": true },
{ "query": "What kill rules should I pre-register before launching new ad creatives?", "should_trigger": true },
{ "query": "Plan a TikTok creative test - 4 concepts, $300 a day", "should_trigger": true },
{ "query": "How do I know if my ad test is underpowered before I spend the money?", "should_trigger": true },
{ "query": "We keep arguing about which ad concept to scale - design a fair test", "should_trigger": true },
{ "query": "creative testing framework for a B2B lead gen account", "should_trigger": true },
{ "query": "how big should my test budget be per ad set", "should_trigger": true },
{ "query": "Is a week long enough to test new ad creative?", "should_trigger": true },
{ "query": "I have 6 new ads and $2,000. What's the smartest way to find the best one?", "should_trigger": true },
{ "query": "Write me a test matrix for hook variations on our best performing video ad", "should_trigger": true },
{ "query": "MDE and sample size for an ad test with a 2% conversion rate", "should_trigger": true },
{ "query": "before we spend on these new concepts, how should we structure the experiment", "should_trigger": true },
{ "query": "What's the minimum spend before judging a new creative?", "should_trigger": true },
{ "query": "Should I use Meta's A/B test tool or split budgets manually for creative testing?", "should_trigger": true },
{ "query": "How many ad variants can I test at once with $150/day?", "should_trigger": true },
{ "query": "statistical significance for ad creative tests - how do I calculate it", "should_trigger": true },
{ "query": "Our agency proposed testing 12 creatives simultaneously. Sanity check the plan?", "should_trigger": true },
{ "query": "Design a test to find out whether the founder-story angle beats the discount angle", "should_trigger": true },
{ "query": "I want to test a new ad format without tanking performance - how do I structure it?", "should_trigger": true },
{ "query": "how do I set up ongoing creative experiments with a protected testing budget", "should_trigger": true },
{ "query": "success criteria to define before launching new ad variations", "should_trigger": true },
{ "query": "how do I avoid calling an ad winner too early", "should_trigger": true },
{ "query": "We're a startup with 40 conversions a month. Can we even A/B test our ads?", "should_trigger": true },
{ "query": "test plan for comparing 3 messaging angles in paid social", "should_trigger": true },
{ "query": "big swing testing vs isolating the hook - what's right for us?", "should_trigger": true },
{ "query": "champion vs challenger setup for testing new ads", "should_trigger": true },
{ "query": "how much of my budget should go to each cell in a creative test", "should_trigger": true },
{ "query": "Plan a creative experiment where we'll actually learn which element drove the lift", "should_trigger": true },
{ "query": "boss asked for a 'proper scientific test' of our new ads - where do I start", "should_trigger": true },
{ "query": "when is a creative test result trustworthy enough to act on - planning next quarter's tests", "should_trigger": true },
{ "query": "pre-launch checklist for an ad creative experiment", "should_trigger": true },
{ "query": "How do I test two YouTube ad cuts against each other properly?", "should_trigger": true },
{ "query": "we've got 3 concepts from the agency - decide the test order and budgets", "should_trigger": true },
{ "query": "what metrics should decide the winner in an ad creative test", "should_trigger": true },
{ "query": "our last creative test came back 'inconclusive' - design the next one so that can't happen", "should_trigger": true },
{ "query": "help me size an ad experiment: $80 CPA, need an answer in 3 weeks", "should_trigger": true },
{ "query": "do I need a control group when testing new ad creatives?", "should_trigger": true },
{ "query": "LinkedIn ads - want to test a practitioner pain angle vs an ROI angle, what's the plan", "should_trigger": true },
{ "query": "how do I keep the algorithm from ruining my creative test", "should_trigger": true },
{ "query": "How many assets per concept when testing new ad angles?", "should_trigger": true },
{ "query": "set up a fair bake-off between our in-house creative and the agency's", "should_trigger": true },
{ "query": "should the creative test have its own budget or share the campaign's?", "should_trigger": true },
{ "query": "I read most ad tests never reach significance - how should we test then?", "should_trigger": true },
{ "query": "hypothesis format for ad creative experiments", "should_trigger": true },
{ "query": "test roadmap for new ad concepts this quarter on a $25k/month account", "should_trigger": true },
{ "query": "can I trust a 4-day read on a new ad? planning our testing rules now", "should_trigger": true },
{ "query": "what does a good ad testing process look like before any money is spent", "should_trigger": true },
{ "query": "My winning ad's CPA doubled over three weeks - is it fatigued or something else?", "should_trigger": false },
{ "query": "How do I read the results of the creative test we launched last Monday?", "should_trigger": false },
{ "query": "Write 10 ad copy variants for our project management tool", "should_trigger": false },
{ "query": "Draft a creative brief for the UGC creator we hired", "should_trigger": false },
{ "query": "Check whether our Meta pixel purchase event is double-firing", "should_trigger": false },
{ "query": "Our whole ad account has been bleeding for two months - diagnose what's wrong", "should_trigger": false },
{ "query": "Which ad platforms should a $10k/month DTC brand be on?", "should_trigger": false },
{ "query": "How should I split $50k across Google, Meta and TikTok this quarter?", "should_trigger": false },
{ "query": "We're pacing 40% over budget mid-month - what do I do?", "should_trigger": false },
{ "query": "Should I switch this campaign from max conversions to target CPA bidding?", "should_trigger": false },
{ "query": "Build an audience targeting plan for our meal-kit launch", "should_trigger": false },
{ "query": "Which customers should seed our lookalike audience?", "should_trigger": false },
{ "query": "Design a retargeting sequence for cart abandoners", "should_trigger": false },
{ "query": "Audit the landing page our ads send traffic to", "should_trigger": false },
{ "query": "Set up an A/B test for our landing page headline", "should_trigger": false },
{ "query": "How do I A/B test email subject lines for our newsletter?", "should_trigger": false },
{ "query": "Sample size for an A/B test on our checkout flow redesign", "should_trigger": false },
{ "query": "We want to test two pricing tiers against each other - how long and how many users?", "should_trigger": false },
{ "query": "Score these five video hooks and tell me which deserve budget", "should_trigger": false },
{ "query": "Build a swipe file of competitor ads organized by hook type", "should_trigger": false },
{ "query": "Our campaign is profitable - how fast can we scale the budget?", "should_trigger": false },
{ "query": "Too many overlapping ad sets - which should we merge?", "should_trigger": false },
{ "query": "Mine our search terms report for negative keywords", "should_trigger": false },
{ "query": "Is a 3.2 ROAS healthy given our margins?", "should_trigger": false },
{ "query": "Meta says 120 purchases, Shopify says 74 - explain the gap", "should_trigger": false },
{ "query": "Write UGC scripts with hook variants for our skincare brand", "should_trigger": false },
{ "query": "What interview loop should we run for a media buyer hire?", "should_trigger": false },
{ "query": "How do I move from PPC specialist to creative strategist?", "should_trigger": false },
{ "query": "Set a maximum allowable CAC policy for the whole company", "should_trigger": false },
{ "query": "Is a carousel the right format for our top-of-funnel campaign?", "should_trigger": false },
{ "query": "Map the buying committee for our $80k ACV security product", "should_trigger": false },
{ "query": "Should we promote our CEO's LinkedIn posts as paid ads?", "should_trigger": false },
{ "query": "Adapt our ad copy for AI assistant answer placements", "should_trigger": false },
{ "query": "Build me a watch list of paid media newsletters and podcasts", "should_trigger": false },
{ "query": "A/B test my onboarding flow - control vs the new tooltip tour", "should_trigger": false },
{ "query": "How do I run a feature-flag experiment for our new checkout service?", "should_trigger": false },
{ "query": "Design a survey experiment to measure brand recall", "should_trigger": false },
{ "query": "statistical significance calculator for my NPS survey results", "should_trigger": false },
{ "query": "Which of these two logo designs should we pick? Set up a test", "should_trigger": false },
{ "query": "Test two versions of our app store screenshots against each other", "should_trigger": false },
{ "query": "Help me plan a geo holdout to measure whether our Meta ads are incremental at all", "should_trigger": false },
{ "query": "Can we split test our SEO title tags?", "should_trigger": false },
{ "query": "Set up a multivariate test on the homepage in our CRO tool", "should_trigger": false },
{ "query": "Compare our two email nurture sequences and tell me which converts better", "should_trigger": false },
{ "query": "We printed two direct mail postcard versions - how do we measure which one works?", "should_trigger": false },
{ "query": "How many user interviews do I need before the findings are reliable?", "should_trigger": false },
{ "query": "Run a power analysis in R for our product experiment", "should_trigger": false },
{ "query": "Why did my ad get rejected by Meta's ad review?", "should_trigger": false },
{ "query": "The test ended - write up the results deck for leadership", "should_trigger": false },
{ "query": "How often should we rotate and refresh our ad creatives?", "should_trigger": false },
{ "query": "Which thumbnail should I use for our YouTube video?", "should_trigger": false },
{ "query": "Pick the best subject line for tomorrow's product launch email", "should_trigger": false },
{ "query": "How do I test whether our new sales pitch deck lands better with prospects?", "should_trigger": false },
{ "query": "Analyze last quarter's creative tests and tell me what patterns won", "should_trigger": false },
{ "query": "Set up experiments on our pricing page to lift signups", "should_trigger": false },
{ "query": "What's a good CTR benchmark for Facebook ads in fashion?", "should_trigger": false },
{ "query": "Brainstorm 10 new ad concepts for our spring campaign", "should_trigger": false },
{ "query": "Our Meta A/B test finished with 'no clear winner' - what does that mean?", "should_trigger": false },
{ "query": "Split test two versions of our Shopify product page", "should_trigger": false }
]
}
references/example-test-plan.md›
# Worked Example Test Plans
Contents: 1. B2C ecommerce plan (full math) · 2. B2B lead-gen plan (Directional default) · 3. Negative example
All numbers below are illustrative scenario inputs; the formulas they run through are the ones in SKILL.md section 4.
## Table of Contents
- [1. B2C ecommerce - three angle concepts vs champion](#1-b2c-ecommerce---three-angle-concepts-vs-champion)
- [2. B2B lead gen - Directional by default](#2-b2b-lead-gen---directional-by-default)
- [3. Negative example - what this skill must never emit](#3-negative-example---what-this-skill-must-never-emit)
## 1. B2C ecommerce - three angle concepts vs champion
Scenario inputs:
- DTC skincare brand.
- Protected test budget $27,000/month ($900/day).
- Target CPA $30; observed CPC $1.50.
- Click→purchase rate 2.0%; click→add-to-cart rate 8%.
- Decision: which of three new messaging angles earns next quarter's production budget.
Feasibility math, shown so the plan can be audited:
- 4 cells (control + 3 concepts) at $225/day each. Stable-delivery floor = $30 × 50 ÷ 7 ≈ $214/day → passes. Projected conversions ≈ 225 ÷ 30 = 7.5/day ≈ 52/week per cell → clears ~50/week.
- Projected clicks ≈ 225 ÷ 1.50 = 150/day per cell.
- Powered check on purchase rate at a 25% relative MDE (2.0%→2.5%): required n ≈ 13,800 clicks/cell → ~92 days at 150 clicks/day → far beyond the 4-6 week ceiling → **not Powered on purchase**.
- Move the read up-funnel: add-to-cart at 8% baseline, 25% relative MDE (8%→10%): required n ≈ 3,200 clicks/cell → ~22 days → within bounds → **Powered on add-to-cart**.
- Purchases accumulated by day 22: 7.5/day × 22 ≈ 165/cell → enough for a relative CPA ranking, only detects very large gaps → **Directional on cost per purchase**.
```
CREATIVE TEST PLAN - Q3 angle test, 2026-08-26
decision : winning angle gets Q4 production budget; losers' angles retired
hypothesis : because reviews mention "routine takes too long" 3x more than price,
a time-saved angle will lower cost per purchase ~20% vs the
ingredient-story champion, for cold prospecting, by day 22
isolation : single variable: angle (concept-level execution held to same
format mix per cell) - high-leverage lever
structure : 3 test cells + control champion, concurrent; manual fixed-budget
cells $225/day each; automated creative-optimization: off
metrics : gate = hook rate vs account trailing median, per placement,
read only at >=2,000 impressions per asset
primary = cost per purchase (decision), add-to-cart rate (powered read)
guardrails = frequency, CPM vs account baseline, refund rate
cells : C01_ANG-champion_V01 (control) | $225/day | 4 assets
projected 52 conv/wk | VERDICT: reference cell, same rules
C02_ANG-timesaved | $225/day | 4 assets
C03_ANG-sensitive-skin | $225/day | 4 assets
C04_ANG-social-proof | $225/day | 4 assets
each test cell: projected ~52 conv/wk, ~150 clicks/day
required (purchase, 25% MDE): n=13,800 clicks, ~$20,700, ~92d
-> VERDICT: Directional read on cost per purchase
required (add-to-cart, 25% MDE): n=3,200 clicks, ~$4,800, ~22d
-> VERDICT: Powered on add-to-cart by day 22
kill: asset at $60 spend (2x CPA) with zero purchases;
cell at $1,000 spend with no activated asset
scale: best cell +50-100% budget; first scale step is its own read
iterate: cell that wins add-to-cart but loses CPA -> landing-page
check before any creative iteration
schedule : launch Mon | earliest evaluation day 3 | hard stop day 22 or
3,200 clicks/cell, whichever first | inconclusive -> keep control
naming : C##_ANG-<angle>_HOOK-<type>_FMT-<format>_TAL-<creator>_V##
caveats : manual cells share auctions - overlap noted; divergent delivery
means even a "Powered" read is relative, not causal; CPA verdict
is directional and will be reported as a ranking, not a winner
at 95% confidence
```
Note the double read standard, stated per metric: the cell is Powered on add-to-cart and Directional on purchase, and the plan says which claim each metric may make.
## 2. B2B lead gen - Directional by default
Scenario inputs:
- B2B security SaaS.
- Test budget $12,000/month ($400/day); CPL $80 on the form event.
- CRM-fed qualified-lead rate ~35% (cost per qualified ≈ $229).
- 90-day sales cycle.
- Decision: does a practitioner-fear angle beat the compliance-checklist champion.
Feasibility math:
- Stable-delivery floor on the lead event = $80 × 50 ÷ 7 ≈ $571/day per cell. Even a single test cell + control at $200/day each sits far below it → every cell is delivery-limited on the lead event. Raising budget is not available; consolidating below 2 cells is impossible (a test needs a control).
- Projected leads at $200/day ≈ 2.5/day ≈ 105 per cell over 6 weeks; ~37 qualified per cell. No MDE reachable at significance.
- Verdict, declared: **no cell can be Powered - this plan is a screening plan**, judged on relative ranking over a 6-week window with lagging quality guardrails.
```
CREATIVE TEST PLAN - practitioner-fear angle screen, 2026-08-26
decision : challenger replaces champion in always-on prospecting if it ranks
better on cost per qualified lead without breaching quality
hypothesis : because sales calls open with breach anecdotes, a practitioner-fear
angle will lower cost per CRM-qualified lead ~20% vs the
compliance champion over 6 weeks (magnitude is a screening target,
not a significance claim)
isolation : bundled - unlearnable at element level (angle + new visuals + new
copy change together; labelled so no element-level insight is
mined from the result)
structure : 1 test cell + control, $200/day each, 3 assets per cell; manual
fixed budgets; automated creative-optimization: off
metrics : gate = CTR vs account trailing median, per placement
primary = cost per CRM-qualified lead (fed back from CRM,
never raw form fills)
guardrails = lead-to-qualified rate >= account trailing median;
frequency (small audience saturates by design); CPL drift
cells : both cells: projected ~17 leads/wk - below the ~50/wk stable-
delivery floor ($571/day needed at $80 CPL)
required for Powered: not reachable at any acceptable MDE
-> VERDICT: Directional read, declared; delivery-limited flagged
kill: asset at $320 spend (4x CPL) with zero leads; cell decision
deferred to week 6 - no mid-window kills on a 17/wk event
scale: replace champion only if challenger ranks better on cost
per qualified AND quality guardrail holds; +50% budget max
schedule : launch Mon | earliest evaluation day 14 | hard stop day 42
| inconclusive -> keep champion | second look at day ~130:
opportunity creation per cell after one sales cycle
naming : C##_ANG-<angle>_FMT-<format>_V##
caveats : sample cannot support significance - all reads are judgment plus
ranking; revenue verdict arrives one sales cycle after the
creative verdict and is scheduled, not skipped
```
## 3. Negative example - what this skill must never emit
> "Variant B won with 40 conversions vs 28 at 97% confidence after 4 days. Scaling it 3x."
Everything wrong with it, in order:
1. The sample is an order of magnitude below the required n for any realistic MDE, so the "97% confidence" is a peeked artifact.
2. Day 4 is before the pre-registered evaluation moment and inside the novelty window.
3. The cells ran under campaign-level automatic budget allocation, so "B" had 3× the spend and a different delivered audience (divergent delivery) - the comparison was never even.
4. The 3× scale step ignores regression to the mean on a low-spend winner.
The honest version of the same data: "Directional: B ranks ahead of A on CPA at day 4; no verdict until the pre-registered stop; allocation was uneven, so the ranking is provisional."
references/sizing-reference.md›
# Sizing Reference: Formulas, Anchors, and Sources
Contents: 1. Sample-size formula · 2. Quick-reference tables · 3. Duration math · 4. Multiple comparisons · 5. Stable-delivery floor · 6. Practitioner kill/judge anchors (attributed) · 7. Proxy-metric evidence · 8. Academic reality check · 9. Spend-tier guidance · 10. Naming-convention source
## Table of Contents
- [1. Sample-size formula](#1-sample-size-formula)
- [2. Quick-reference tables (sample per cell)](#2-quick-reference-tables-sample-per-cell)
- [3. Duration math and bounds](#3-duration-math-and-bounds)
- [4. Multiple comparisons](#4-multiple-comparisons)
- [5. Stable-delivery floor](#5-stable-delivery-floor)
- [6. Practitioner kill/judge anchors - attributed](#6-practitioner-killjudge-anchors---attributed)
- [7. Proxy-metric evidence](#7-proxy-metric-evidence)
- [8. Academic reality check](#8-academic-reality-check)
- [9. Spend-tier guidance](#9-spend-tier-guidance)
- [10. Naming-convention source](#10-naming-convention-source)
## 1. Sample-size formula
Two-proportion test, per cell, alpha 0.05 two-tailed (Z = 1.96), 80% power (Z = 0.84):
```
n = ( 1.96 × sqrt(2 × p̄ × (1−p̄)) + 0.84 × sqrt(p1(1−p1) + p2(1−p2)) )² / (p2 − p1)²
```
p1 = baseline rate; p2 = p1 × (1 + relative MDE); p̄ = (p1 + p2) / 2.
Shortcut (within a few percent, using p̄): `n ≈ 16 × p̄(1−p̄) / MDE²`, MDE in absolute terms. Worked: 5% baseline, detecting a 1-point absolute lift (5%→6%): 16 × 0.055 × 0.945 / 0.0001 ≈ **8,300 per cell** (the exact formula gives ~8,100).
Worked anchors, computed from the exact formula above:
- 2% baseline, 50% relative lift (2%→3%): ~3,800 per cell.
- 5% baseline, 20% relative lift (5%→6%): ~8,100 per cell.
- 5% baseline, 5% relative lift (5%→5.25%): ~122,000 per cell.
- Halving the MDE roughly quadruples n; a 1% baseline needs ~5× the sample of a 5% baseline for the same relative MDE.
Published calculators disagree with these by a few percent because they differ on pooled vs unpooled variance. That spread is irrelevant next to the decision the number drives - do not chase it.
The "~100 conversions per variant" and "25-50 conversions minimum" rules that circulate correspond roughly to detecting only large (>30-50%) relative effects; the "100-400 conversions per variant" band is community convention with no single traceable authority. Use them as what they are: thresholds for a Directional read, not for a Powered one.
## 2. Quick-reference tables (sample per cell)
| Baseline | 10% rel. lift | 20% rel. lift | 50% rel. lift |
| -------- | ------------- | ------------- | ------------- |
| 1% | 163,000 | 43,000 | 7,700 |
| 3% | 53,000 | 14,000 | 2,500 |
| 5% | 31,000 | 8,100 | 1,500 |
| 10% | 15,000 | 3,800 | 700 |
Per cell - double each figure for the two-cell total, and multiply by the number of cells for the whole test. Units are observations at the metric's denominator (impressions for CTR, clicks for post-click CVR, and so on). Read the table before promising anyone a significance-tested creative winner on a conversion metric.
## 3. Duration math and bounds
```
required duration (days) = required sample per cell / projected daily events per cell
projected daily events = daily cell budget / cost per event
```
- Floor: one full week minimum, always - day-of-week composition; two business cycles for B2B.
- Ceiling: ~4-6 weeks - beyond it, novelty decay, fatigue, and external events contaminate the read, and the opportunity cost of blocked test slots compounds.
- If required duration exceeds the ceiling, those are the only four levers, ranked by power recovered per unit of effort and learning given up: **up-funnel read > wider MDE > fewer cells > more budget**. Per-axis breakdown and the re-ranking conditions are in SKILL.md section 4, where the choice is made.
## 4. Multiple comparisons
Testing k variants against control inflates family-wise error. Standard correction: divide alpha by the number of comparisons (Bonferroni).
Practical effect: ~3 variants raises required sample per cell roughly 30-40%. Media buyers rarely apply it; the plan should either correct alpha or reduce the cell count - and must at minimum disclose how many comparisons the test runs.
## 5. Stable-delivery floor
- Delivery algorithms need roughly **50 optimization events per cell per rolling 7 days** for stable delivery; below that the cell stays delivery-limited and its numbers are unstable regardless of sample math. The count is per cell (ad-set level), not per asset.
- Minimum daily budget per cell ≈ **target CPA × 50 ÷ 7**. Example: $35 CPA → ~$250/day/cell.
- Fixes when the floor cannot be met, ranked by events recovered per unit of effort: **consolidate cells** (near-zero - a restructure before launch) **> optimize on a higher-funnel event with more volume** (a decision plus a tracking check) **> improve signal capture** (server-side event feeds - an engineering project, highest ceiling, slowest, and the only one carrying a consent-scope review). Re-rank when server-side capture already exists, which makes the third fix near-free and moves it first.
- Perspective: crossing from delivery-limited to stable is worth a modest ~5-10% efficiency gain. The floor matters for read validity, not as a performance unlock - "spend more to escape learning" is bad advice when unit economics are already poor.
- Roughly half of low-budget campaigns never stabilize; over-segmentation (too many thin cells) is the most common self-inflicted cause.
## 6. Practitioner kill/judge anchors - attributed
Treat these as starting defaults to adapt against account history, not laws. Adopt exactly one - they contradict each other. Rows are in efficiency order (decision quality per unit of test spend burned): **Hott > Denney > CTC > Bachman > Faris**; the per-axis breakdown, the default rung, and the conditions that reorder it are in SKILL.md section 6, where the choice is made.
| Source | Rule |
| ----------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Barry Hott | Rejects mechanical ad-level CPA kill rules: "Ad-level CPA and ROAS is irrelevant! The system isn't trying to get you the best CPA on every individual ad, it's trying to get you the most possible conversions for your overall budget." Method: comparative benchmarking of new ads against the best ads' spend/CPA range in a controlled fixed-budget environment. Also the named "champion vs challenger" framing |
| Dara Denney (Thesis; Motion) | Budget per test ≈ avg CPA × 50 conversions (e.g. $35 CPA → $1,750 total, ~$250/day for a week); 6 ads per test; do not evaluate before day 3; kill an ad at 2× average CPA with no purchase; kill the ad set if no winner after 5-7 days; scale winners by +50-100%, done 2-3 times |
| Common Thread Collective (Taylor Holiday) | Kill on spend thresholds, not time: no activation by $500-1,000 spend (~72 hours to 7 days) → kill; allow 10-15% of testing budget for longer-runway ads |
| Jess Bachman (FireTeam) | "Until you at least spend enough to get a purchase or spend 3-4 times your CPA at least, you haven't given that new creative a chance to prove itself" |
| Andrew Faris (AJF Growth) | No manual kill rule; launch new creative into the evergreen structure under bid caps so the platform enforces the CPA ceiling; on when to pause: "never" |
| Motion (vendor blog) | Minimum before evaluating: ≥2,000 impressions, 50-100 clicks, or 3-5 purchases per creative, and ≥3 days runtime; kill at CTR <50% of control or CPA >25% worse than target sustained 48-72h; ~10,000 impressions or 1,000 conversions for directional reads. Explicitly decides winners on comparative thresholds, not statistical significance |
**Folklore correction - do not launder this.** The ubiquitous "spend 1× CPA before judging / 3× CPA before killing" rules are **untraceable folklore**: repeated across agency blogs, attributed to no identifiable originator, and sometimes mis-attributed to Barry Hott, who publishes the opposite position (above). The closest properly attributed anchors are Bachman's "3-4× CPA before judging" and Denney's "2× CPA no-purchase kill". Cite those, with their names, or cite nothing.
The impression-threshold convention (~1,000-2,000 impressions before judging a hook-rate gate) is common practice, not an authoritative rule.
## 7. Proxy-metric evidence
- Funnel Insiders: an analysis of $1.47M in ad spend across client accounts found **no statistically significant correlation between thumbstop (hook) rate alone and revenue**.
- Two ads at an identical hook rate can differ several-fold in return, because the gate metric says nothing about how many of those viewers were still watching at the call to action.
- Published hook-rate "good" bands conflict from ~18% to ~40% across vendors - set gate thresholds from the account's own trailing median, never from a published band.
- Practitioner consensus: hook rate is most predictive in cold prospecting, least at bottom-funnel; high hook rate + low conversion usually signals an offer/landing-page problem. Gates must be compared like-for-like by placement, or a placement-mix shift crowns an accidental winner.
## 8. Academic reality check
- **Lewis & Rao, "The Unfavorable Economics of Measuring the Returns to Advertising"** (Quarterly Journal of Economics 130(4), 2015): across 25 large field experiments, "informative advertising experiments can easily require more than 10 million person-weeks"; "the median confidence interval on return on investment is over 100 percentage points wide."
- **Gordon, Zettelmeyer, Bhargava & Chapsky** (Marketing Science 38(2), 2019): 15 large ad experiments; observational methods (the pixel/last-click reads most kill decisions rely on) "often fail to produce the same effects as the randomized experiments," typically overestimating effectiveness.
- **Braun & Schwartz** (Journal of Marketing, 2025, DOI 10.1177/00222429241275886): even platform-native A/B tests are confounded by **divergent delivery** - platforms "deliver different ads to distinct and undetectably optimized mixes of users that vary across ads, even during the test," which can "confound the magnitude, and even the sign, of ad A/B test results." Only holdout/lift designs isolate the creative effect.
Implication for the plan: in-platform results - including deterministic splits - are relative screening reads. Reserve geo-holdout / conversion-lift designs for channel-level and big-swing decisions where causality is worth its cost.
## 9. Spend-tier guidance
| Monthly spend | What testing can honestly be | Practitioner guidance |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| Under ~$20k | No significance. Proxy-gated screening only: gate metric kills obvious losers cheaply, then relative CPA picks scalers. 5-15 concept-level ads/month | Testing viability generally starts around $20k/month (Denney) |
| $20k-100k | Dedicated fixed-budget test structure, one concept per cell, 4-6 assets each; Denney's numeric defaults apply | ~2 concepts/week × 2 hooks × 2 visuals at $30-100k (Savannah Sanchez); ~10-15 ads per $50k (Manson Chen) |
| $100k-1M | Full program: dedicated test cells plus cost-cap throttles; 40-70 new ads/month at 8-figure scale (CTC) | Layer a rolling incrementality program to validate that in-platform winners are causal |
| $1M+ | Continuous: 80-150 ads/month (CTC); winners recycled; incrementality on top | Testing-budget share: Denney recommends ≥20% of spend on testing; CTC 40-50% early-stage, 20-30% mature |
Outlier economics (CTC, single-agency data across 170+ brands - strong prior, not a universal constant):
- The top 3.5% of ads generate 66% of spend.
- ~79% of ads never reach $1,000 spend before being killed.
- Discovery of an outlier takes ~18-81 days.
Creative testing is a power-law portfolio game - plan volume accordingly (if outliers are ~3.5% of tests and you need 3/month, plan ~86 tests/month).
## 10. Naming-convention source
The structured variant-name pattern in SKILL.md follows documented practice: Growth-Rocket's format `HOOK-Question_FORMAT-UGC_OFFER-FreeTrial_V03`, and Motion's naming glossary encoding source, format/concept, messaging, hook, product, talent, asset length, offer, landing-page type, and test parameters.
Three-level discipline:
- Campaign encodes objective/funnel/geo.
- Cell (ad set) encodes audience/optimization.
- Asset encodes concept/hook/format/version.
Creative-reporting tools are syntactic parsers - garbage in, garbage out.
SKILL.md›
---
name: ad-creative-test-plan
description: "Design a pre-launch ad creative test plan - falsifiable hypothesis, isolation level, test cells with per-cell budgets, required sample, spend and duration, and kill/scale rules pre-registered before any money moves. Every cell carries an explicit read standard, so an underpowered screen is never dressed up as an A/B test. Use whenever the user wants to test ads, mentions an A/B or split test on creative, asks how much budget or how long a test needs, or mentions sample size, statistical significance, single-variable vs big-swing testing, or test cell structure - even if they never say 'test plan'. Covers B2B and B2C. Do NOT use to read results from a test already running - use mbfinotti/advertising-skills@ad-creative-fatigue instead."
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.2.2"
---
# Creative Test Plan
You are a paid-media experimentation lead. Design the creative test _before_ launch, then emit a plan document a media buyer can execute without you:
- Isolate the variable deliberately.
- Structure the cells.
- Size the spend and sample.
- Pre-register the success criteria and decision rules.
The core discipline is honesty about power: most creative "A/B tests" at normal budgets cannot reach statistical significance, and the plan must say so explicitly rather than let a screening heuristic masquerade as a controlled experiment. Peer-reviewed work shows informative ad experiments can require millions of person-weeks (Lewis & Rao, 2015 - full citations in [references/sizing-reference.md](references/sizing-reference.md)), which is exactly why practitioners run spend-threshold heuristics instead. Both frames are legitimate, provided the plan states which one each cell is using.
This skill ends when the plan document ships. Handoffs beyond that boundary:
- Reading results after launch and monitoring wear-out: `mbfinotti/advertising-skills@ad-creative-fatigue`.
- Writing the ad copy variants: `mbfinotti/advertising-skills@ad-copy-variants`.
- The production brief for designers or video: `mbfinotti/advertising-skills@ad-creative-brief`.
- Verifying that conversion events actually fire: `mbfinotti/advertising-skills@ad-conversion-tracking` (run it before launch - a test on a broken event measures nothing).
- Diagnosing an underperforming account: `mbfinotti/advertising-skills@ad-account-diagnostic`.
## Interview
Ask before designing anything. One question per message; offer multiple-choice options where possible; skip anything already answered or visible in supplied data. Stop and ask when an answer is missing - never assume.
- What decision will this test inform? (Which concept gets next quarter's production budget / which angle to scale / whether a new format earns a slot / something else.) A test that informs no decision is spend, not learning.
- Platform class and campaign objective? (Paid social / search / video / native; conversions, leads, or traffic objective.)
- Monthly or planned test budget, and is it protected from the scaling budget?
- Current CPA or CPL, and the conversion rate at the step you would judge on?
- Monthly conversion volume on that event - roughly how many per month, account-wide?
- B2B or B2C, and how long is the sales cycle?
- How many creative assets already exist or can be produced for this test, and at what lead time?
- What has already been tested, and what happened? (Prevents re-testing settled questions and new-vs-old comparisons.)
- Is an automated or AI creative-optimization feature currently on for these campaigns? (It recombines and selects assets itself, which breaks controlled cells.)
- Is there a current champion creative to serve as the control cell?
- By what date must this test have produced its answer, and is that date hard? (Test structures differ by weeks in time-to-answer, so the ordering in sections 2, 3 and 6 cannot be chosen without it.)
- Do you want a one-off winner to scale now, or a transferable element-level learning that compounds across future tests? (The first favours bundled concepts, the second favours strict isolation.)
- What is the effort ceiling - who watches the account, how often, and how much budget authority do they have without asking anyone?
- Any consent, disclosure, claim-substantiation or regulated-category constraint on these creatives? (Health, finance, employment, housing, minors; testimonial or AI-generated-likeness disclosure.)
## 1. Define the decision, then the hypothesis
1. Write the decision in one sentence: "Based on this test, we will ___." If the blank cannot be filled, stop and redesign.
2. Write a falsifiable hypothesis with this template - all three parts required:
```
Because [observation or evidence],
changing [the one thing this test varies]
will [raise/lower] [named metric] by roughly [magnitude]
for [audience], and we will know by [date/sample].
```
3. Reject hypotheses with no predicted direction and magnitude ("let's see what happens") and hypotheses whose metric is not measurable within the test window. In B2B, the honest magnitude claim often lives on a proxy (qualified-lead rate) because the revenue event lags the test by a sales cycle - say so in the hypothesis.
## 2. Choose the isolation level deliberately
Default ordering, stated out loud so the choice is not made by row order:
- efficiency (transferable learning bought per unit of spend and calendar burned): **bundled concept > tiered > strict single-variable**
- value if the read lands: **strict single-variable > tiered > bundled** - only isolation yields element-level learning
- spend and calendar burned before any read: **strict single-variable > tiered > bundled**
There is a real, live debate here - present it honestly, but as an efficiency question, not a philosophical one. The three postures, in that default order:
1. **Concept-level "big swings"** (dominant modern camp) - change messaging and execution together in one bundled test.
- Buys a larger, detectable effect at normal budgets, and feeds delivery algorithms the creative diversity they reward.
- Costs the element-level learning: a bundled win says the concept won, never why. Label it `bundled - unlearnable at element level` in the plan so nobody later mines a fake element-level insight out of it.
2. **Tiered** - validate the angle with bundled cells first, then isolate hook or format inside the winning angle. Buys both reads, in sequence; costs two test windows instead of one, so it needs volume and calendar the other two postures do not.
3. **Strict single-variable isolation** (classical camp) - change one element and hold everything else constant.
- Only clean isolation produces transferable, element-level learnings.
- At normal budgets single-element effects are usually too small to detect, so the test burns spend proving nothing: run it only when the feasibility check (section 4) shows the cell can reach at least a Directional read on that lever's natural metric. Below that, the isolation is theater.
Bundled leads because it is the only posture that reliably produces a detectable effect at normal budgets - strict isolation returns more per test, but only when the test can resolve at all.
What this order starves is strict single-variable isolation, permanently: it ranks first on value and last on efficiency, so the ratio never selects it and the account accumulates no transferable element-level learning at all. Promote it on named conditions rather than waiting for the ratio to turn - the section 4 feasibility check returns Powered on that lever's own MDE, _and_ the interview answered "compounding learning" rather than "a winner to scale now". A settled angle is the third trigger: once bundled cells keep crowning the same angle, the only question left is which element carries it, and no bundled test can answer that.
This is a default, not a law: it shifts with conversion volume and with who runs the account. Re-rank against the interview answers and name which answer moved which posture - a hard near-term date promotes bundled, a wide effort ceiling with calendar to spare promotes tiered. Re-rank again against the account: volume that can power a single-element read, or a creative team that ships matched variants cheaply, moves isolation up.
Delete a posture the constraints rule out instead of leaving it at the bottom of the list - a posture parked there returns as scope the week before launch.
- No creative capacity to produce variants differing in exactly one element deletes strict isolation from the menu: state that it is deleted, and design between bundled and tiered.
- A single test window before the decision date deletes tiered the same way.
Whichever posture wins, only isolate **high-leverage levers**: concept, angle, hook, format, creator/talent. Refuse to burn spend isolating micro-variables (button color, font, minor copy) at normal budgets - fold them into a concept or drop them.
Those five levers are deliberately left unranked against each other: their effect sizes are account-specific and mostly a function of what the account has already settled, so a general order would be false precision. Rank them against this account's own test log instead - the lever with the widest untested spread goes first.
## 3. Build the cell matrix
1. **Control cell**: the current champion creative, running concurrently in the same structure at the same budget. Never compare new cells to the champion's historical numbers - unequal delivery history and seasonality make old numbers incomparable.
2. **One concept per cell**, 3-6 assets per cell (variations executing the same concept).
- More cells than concepts fragments budget.
- More concepts than cells contaminates the read.
3. **Pick the cell structure**, ranked by comparable read bought per unit of setup and delivery efficiency given up: **manual fixed-budget cells > platform-native deterministic split > campaign-level automatic allocation**.
- _Manual fixed-budget cells_ - the default. Each cell holds its own budget, so cells stay comparable; the trade is a little overall delivery efficiency, which is the point of a test.
- Known limitation: manual cells still compete in the same auctions against overlapping audiences, inflating costs and blurring attribution. Note that contamination caveat in the plan.
- _Platform-native deterministic split test_ - the only in-platform structure that removes overlap: users are deterministically assigned to exactly one cell and never see the other.
- Buys the cleanest available read, but takes a week or more longer to answer and, at typical budgets, is underpowered and frequently returns "no winner".
- Use it when the decision demands that read _and_ the feasibility check says the cells can be Powered. A hard near-term date pushes it back below manual.
- _Campaign-level automatic budget allocation_ - last on every axis. Near-zero effort and it corrupts the read: spend shifts to early leaders and can concentrate up to ~90% of budget on one cell before the others collect data.
- Cheap is not efficient. This is never a test structure.
Re-rank if the account already runs a trusted split-test workflow, or cannot hold fixed budgets against a performance team's objections. Whichever structure wins, turn the automated creative-optimization feature **off** inside test cells - otherwise the platform, not the plan, decides which asset combinations run. Even the deterministic split does not remove divergent delivery (see Failure modes).
4. **Variant naming convention**: encode the decision fields in every asset name so results roll up by dimension. Fixed field order, one delimiter, versioned:
```
C04_ANG-timesaved_HOOK-question_FMT-ugc-video_TAL-creator-jm_V01
```
Concept ID, angle, hook type, format, talent, version - adapt fields to the levers this account tests, then never deviate. Reporting tools parse names; a naming failure silently destroys the roll-up.
## 4. Feasibility check - compute before launch
This is the heart of the plan. For **every cell**, before any money moves:
1. **Projected volume**: daily events per cell = daily cell budget ÷ cost per event (conversions: budget ÷ CPA; clicks: budget ÷ CPC; impressions: budget ÷ CPM × 1,000). Multiply by planned duration for projected sample.
2. **Required sample** for the primary metric, per cell, at 80% power and alpha 0.05, two-proportion formula:
```
n per cell = ( 1.96 × sqrt(2 × p̄ × (1−p̄)) + 0.84 × sqrt(p1(1−p1) + p2(1−p2)) )² / (p2 − p1)²
```
where p1 = baseline rate, p2 = baseline × (1 + relative MDE), p̄ = (p1+p2)/2. Usable shortcut: `n ≈ 16 × p̄(1−p̄) / MDE²` with MDE in absolute terms. Worked anchor: 2% baseline, detecting a 50% relative lift (2%→3%) needs ~3,800 per cell. Halving the MDE roughly quadruples n. Full tables and worked math in [references/sizing-reference.md](references/sizing-reference.md).
3. **Required spend and duration**:
- Required spend = required sample × cost per event.
- Required duration = required sample ÷ projected daily events.
- Duration floor: one full week, always, for day-of-week coverage.
- Duration ceiling: the creative's fresh window (~4-6 weeks before novelty decay and fatigue contaminate the read) and the decision deadline.
4. **Stable-delivery constraint**: a cell that cannot reach roughly **50 optimization events per week** stays delivery-limited - delivery never stabilizes, so the read is unreliable regardless of sample math.
- Minimum daily budget per cell ≈ **target CPA × 50 ÷ 7**.
- A cell below that floor is `Not testable as designed`; fix it with the ranked levers below.
- Exiting the learning state is worth a modest ~5-10% efficiency gain - the floor is about read validity, not a performance unlock.
5. Compare required vs available and emit **exactly one verdict per cell**:
| Verdict | Condition | What the cell is allowed to claim |
| ---------------------------- | -------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Powered** | Projected events per cell ≥ required sample at 80% power, alpha 0.05, within duration bounds | A significance-tested winner |
| **Directional read** | Cannot reach significance in bounds | A screening heuristic only: judged on the gate metric and relative ranking, declared as directional, **never** reported as a winner "at 95% confidence" |
| **Not testable as designed** | Fails stable delivery, or duration exceeds bounds even for a wide MDE | Nothing - fix it with the ranked levers below, then recompute |
Fixing a `Not testable as designed` cell - default ordering, since the four levers are not interchangeable:
- efficiency (power recovered per unit of effort and learning given up): **move the read up-funnel > widen the MDE > fewer cells > raise the budget**
- effort: **raise the budget** (approval, political capital, reversibility) **> move the read up-funnel** (a tracking check and one decision) **> fewer cells > widen the MDE** (near-zero - a declaration)
- learning given up: **widen the MDE** (only large effects remain visible) **> fewer cells** (a question drops out) **> move the read up-funnel** (the business event goes Directional) **> raise the budget** (none)
- compliance cost: **move the read up-funnel > widen the MDE == fewer cells == raise the budget** (none) - only the up-funnel lever can require a new event or a server-side feed, which pulls in consent scope and a privacy review before it can ship
That three-way tie is a real equality, not a dodge: widening the MDE, dropping a cell and raising a budget each change only the test's structure or its declared threshold, so none of them collects anything new and none triggers a review.
Up-funnel leads because it multiplies events without asking anyone for money and without dropping a question from the test. Re-rank against the interview answers: a protected test budget with headroom and real budget authority moves "raise the budget" to the front, and a hard date moves "widen the MDE" up.
Delete a lever the account cannot pull rather than ranking it last - no budget authority and no approver deletes "raise the budget" from the menu by name, and an unverified or broken up-funnel event deletes that lever until `mbfinotti/advertising-skills@ad-conversion-tracking` clears it. A lever left at the bottom gets assumed into the plan's math anyway.
Never silently present an underpowered test as an A/B test. A Directional cell is a legitimate, common outcome - most creative testing at normal budgets is screening - but the plan must say the word. If moving the read up-funnel (a higher-volume event earlier in the funnel) turns a cell Powered, record both reads: Powered on the up-funnel event, Directional on the business event.
## 5. Fix the metric ladder in advance
Three layers, all named in the plan before launch:
1. **Gate metric** - a cheap upper-funnel signal (hook rate, CTR) used _only_ to kill obvious losers early and cheaply. Compare gates like-for-like by placement: placement mix shifts these metrics enough to crown accidental winners. Gates screen; they never crown.
2. **Primary metric** - the single decision metric. B2C: purchase CPA or conversion rate. B2B: a pipeline-quality event (qualified lead, opportunity) fed back from the CRM - never raw form fills.
3. **Guardrails** - metrics that must not degrade while the primary improves: frequency, cost inflation vs account baseline, refund/return rate, lead-quality rate, blended efficiency.
Warning, load-bearing: proxy metrics do not reliably predict conversion. A multi-account analysis of $1.47M in spend found **no statistically significant correlation between hook rate and revenue** (attributed in [references/sizing-reference.md](references/sizing-reference.md)), and two ads at an identical hook rate can differ several-fold in return depending on how many viewers survived to the call to action. High gate + low conversion usually signals an offer or landing-page problem, not a creative winner.
## 6. Pre-register the decision rules
Write these into the plan before launch; changing them after seeing data is the failure the plan exists to prevent.
1. **Kill threshold** per asset and per cell - a spend or performance level at which it dies.
2. **Scale threshold** - what a winner must show, and what happens next (budget increase size, promotion path).
3. **Iterate path** - what qualifies for a re-test instead of a kill or scale.
4. **Earliest evaluation moment** - no judgment before it (day 3 and a minimum volume floor are the common anchors; delayed attribution makes earlier reads systematically pessimistic).
5. **Fixed stopping rule** - a date or a sample size, whichever comes first, plus the pre-declared inconclusive path: no winner resolves to _keep the control_, decided now, not at readout.
6. **Peeking is the named failure**: checking results early and stopping on a favorable number inflates false positives. The stopping rule exists so nobody has to resist temptation in real time.
Adopt **exactly one** of the anchors below. They contradict each other by design - Faris forbids the manual kill Denney prescribes, Hott rejects the per-ad CPA rule both of them use - and mixing two produces a rule that never fires or fires twice. Each also burns a very different amount of test spend before it reaches a decision, which makes the choice an efficiency decision, not a taste one:
- efficiency (decision quality per unit of test spend burned): **Hott comparative benchmarking > Denney numeric defaults > CTC spend thresholds > Bachman 3-4x > Faris no-manual-kill**
- test spend burned before a decision: **Faris > Bachman > Hott > CTC > Denney**
- standing effort to run: **Hott** (a standing job - maintaining a best-ads benchmark and judging against it) **> Faris** (setup of a trusted cost-cap structure, then near-zero) **> Denney == CTC == Bachman** (a threshold checked once)
- time-to-answer: **Faris > Hott > CTC > Bachman > Denney**
The three-way tie on standing effort is genuine: Denney, CTC and Bachman each reduce to one number checked once per asset, with no library to maintain, no comparative benchmark set to keep current, and no structure to earn trust in first.
Hott leads the efficiency axis conditionally, so the default rung is **Denney's set**; what moves you up to Hott is the library plus a weekly reviewer, not a bigger budget. This is a default, not a law - it shifts with account maturity and with who watches the account. Re-rank against the interview answers and name which answer moved which anchor.
Then delete what the answers rule out, by name, instead of listing it as an option for later: a hard date or a compounding-learning mandate deletes Faris, and an effort ceiling with no weekly reviewer deletes Hott. An anchor left on the page gets adopted halfway, which is exactly the mixing this section forbids.
The anchors, in that efficiency order, with what each one needs to work (full sourcing in [references/sizing-reference.md](references/sizing-reference.md)):
- **Barry Hott** - explicitly _rejects_ mechanical per-ad CPA kill rules ("ad-level CPA and ROAS is irrelevant"); judges new ads by comparative benchmarking against the account's best ads inside a controlled fixed-budget structure. Needs a tagged library of best ads and someone judging weekly; without both, it is unavailable, not merely expensive.
- **Dara Denney** - needs nothing the account does not already have, which is what makes it the default:
- Test budget ≈ average CPA × 50; ~6 assets per test.
- No evaluation before day 3.
- Kill an asset at 2× CPA spend with no conversion.
- Kill the cell after 5-7 days with no winner.
- Scale winners +50-100%, done 2-3 times.
- **Common Thread Collective** - kill on spend thresholds, not time: no activation by $500-1,000 spend; reserve 10-15% of test budget for longer-runway ads. Spend-based beats time-based on a variable-delivery account; costs a wider spend band per decision than Denney's.
- **Jess Bachman** - spend at least 3-4× CPA before judging a new creative at all. Buys fewer false kills on high-variance creative; pays for it directly in test spend per decision.
- **Andrew Faris** - no manual kill rule: launch into the evergreen structure under cost controls and let the platform enforce the CPA ceiling; "never" pause manually. Near-zero standing labor once the structure is trusted, but it answers last and returns a portfolio outcome with no element-level read.
- **Flag as folklore** - the widely repeated "spend 1× CPA before judging / 3× CPA before killing" rules have **no traceable originator**; they circulate in agency blogs attributed to no one. Never present them as authoritative, and never let one stand in for an anchor above.
## 7. Emit the plan document
One block per plan; one line per cell. Iterate until the Completion bar passes.
```
CREATIVE TEST PLAN - <name>, <date>
decision : <what changes based on the result>
hypothesis : <because X, changing Y will move METRIC by ~Z% for AUDIENCE by DATE>
isolation : <single variable: which lever | bundled - unlearnable at element level>
structure : <N test cells + control | manual fixed-budget cells or native split test;
automated creative-optimization: off>
metrics : gate <metric + like-for-like rule> | primary <metric> | guardrails <list>
cells : <name per convention> | <concept> | $<x>/day | <assets> assets
projected <n>/wk | required n=<n>, $<spend>, <days>d
VERDICT: Powered | Directional read | Not testable as designed
kill: <threshold> | scale: <threshold> | iterate: <path>
schedule : launch <date> | earliest evaluation <date> | hard stop <date or sample>
| inconclusive -> keep control
naming : <convention string>
caveats : <auction overlap noted | divergent-delivery limits of the read | B2B lag>
```
Full worked examples - a B2C plan with the math computed, a B2B plan defaulting to Directional, and a negative example - live in [references/example-test-plan.md](references/example-test-plan.md).
If your harness has persistent memory, memorize the shipped plan - cells, verdicts, pre-registered thresholds, stop date - so the post-launch read can be checked against the pre-registration instead of a remembered version of it.
## Completion bar
The plan is complete only when **every cell** has, explicitly stated rather than implied:
- a declared read standard: Powered, Directional read, or Not testable as designed;
- a computed required sample, required spend, and required duration;
- a pre-registered kill threshold and scale threshold;
- a named primary metric with guardrails.
Any cell still failing the bar gets removed, merged into another cell, or re-scoped - and the feasibility check re-run - before the plan ships. Iterate until the bar passes; do not ship a plan with an implicit verdict.
## B2B vs B2C
The method - decision, hypothesis, isolation choice, cell matrix, feasibility check, pre-registration - is identical for both. What differs is what the feasibility check concludes and where the primary metric lives:
- **B2C ecommerce**:
- High volume can sometimes produce Powered cells.
- Reads land in 3-7 days on the purchase event.
- The full statistical machinery is usable when the math clears.
- **B2B lead gen**: conversion volume rarely supports significance at all - default the verdict to **Directional read** and say so.
- Reads run over 4-6 weeks, not days.
- Optimize and judge on pipeline-quality events fed back from the CRM (qualified lead, opportunity), never raw form fills. Cheap leads that sales rejects are a guardrail breach, not a win.
- Use judgment-based relative reads plus lagging quality guardrails (lead-to-qualified rate against the account's own trailing median).
- The sales cycle means the revenue verdict arrives one cycle after the creative verdict; the plan must schedule that second look rather than pretend the day-30 read is final.
**Optional integration note** (the only place vendor names belong):
- Meta: the native A/B Test tool provides the deterministic split (7-30 days recommended; shows estimated power at setup); Advantage+ Creative / dynamic creative are the automated features to switch off in test cells.
- TikTok: Split Test runs 7-30 days and declares winners at 90% confidence.
- Google: Experiments run ~4-6 weeks with the first week typically excluded.
Verify against current platform docs if you can browse the web - this layer changes fast. If you cannot, rely on the category-level mechanics above, which are stable.
## Failure modes
| Trap | Why it burns | Mitigation in the plan |
| ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Divergent delivery** - the deepest structural flaw | The delivery algorithm shows each ad to a differently-responsive, undetectably optimized user mix - even inside a deterministic split. Braun & Schwartz (2025, Journal of Marketing) show this can confound the magnitude _and flip the sign_ of a result | State in the plan that in-platform reads are relative screening, not causal proof; reserve holdout/lift designs for decisions that need causality |
| Peeking / early stopping | Stopping on a favorable early number inflates false positives; early data is systematically noisy | Pre-registered stopping rule and earliest evaluation moment (section 6) |
| Multiple comparisons across many cells | Testing many variants makes some "win" by chance; ~3 variants needs roughly 30-40% more sample under a standard correction | Fewer cells, or correct alpha (divide by number of comparisons); details in sizing reference |
| Novelty effect | New creatives get an early boost that fades; short tests over-read it | One-week floor; compare late-window to early-window before calling |
| Day-of-week / seasonality contamination | Weekend buyers differ from weekday; promos and holidays skew a window | Full-week multiples; never launch cells at different times; note calendar events in the plan |
| Mid-test edits | Any significant edit resets delivery learning and invalidates the read | Freeze cells at launch; fix errors by relaunching the cell, not editing it |
| New-vs-old comparison | The incumbent's delivery history and accumulated optimization make old numbers incomparable | Control cell runs concurrently, always (section 3) |
| Regression to the mean when scaling | A low-spend winner's numbers degrade at higher spend; the win was partly selection | Pre-register the scale step size; treat the first scale step as its own read |
| Survivorship bias | Winner libraries and swipe files over-represent survivors; concepts get credit their losers would refute | Log every cell's outcome, including kills, in the test log |
## References
- Read [references/sizing-reference.md](references/sizing-reference.md) for the full sample-size tables, worked formula math, multiple-comparison adjustments, spend-tier guidance, attributed practitioner anchors with sources, and the academic citations.
- Read [references/example-test-plan.md](references/example-test-plan.md) for the worked B2C and B2B plan documents and a negative example.