Retour aux skills
mbfinotti/advertising-skillsContrôle réussi

SKILL DETAIL

ad-hook-analyzer

mbfinotti/advertising-skills/ad-hook-analyzer

Score and force-rank the openings of candidate video ads - from a script, transcript, storyboard, shot list, or a description of a finished cut - to decide which hooks deserve test budget, returning a ranked shortlist within the batch rather than an absolute performance prediction. Use whenever the user mentions a hook, the first 3 seconds, an ad opening, hook rate, thumbstop, attention scoring, or asks which hook to test or whether a hook works - even if they never say 'hook analyzer'. Covers B2B and B2C on any platform. Ends at the ranking: writing the scripts themselves is mbfinotti/advertising-skills@ugc-ad-scripts and static copy is mbfinotti/advertising-skills@ad-copy-variants.

Installations · 174Voir la source

Installation

npx skills add https://github.com/mbfinotti/advertising-skills --skill ad-hook-analyzer

Fichiers du skill

SKILL.md

Dernière synchronisation · 24 sept. 2026

evals/evals.json›
{
  "skill_name": "ad-hook-analyzer",
  "evals": [
    {
      "id": 1,
      "prompt": "I run growth at Brimstone Coffee, a DTC cold-brew concentrate brand ($34 starter kit). We have 4 video ad scripts for Meta feed, cold traffic, B2C, and I need you to score each hook out of 10 and tell me which will get the highest hook rate. Our dashboard defines hook rate as 3-second video plays divided by impressions. The rest of every ad is identical: 20 seconds of brewing demo and a 30-day guarantee offer card. We have an in-house editor and raw footage from the shoot. Candidates, first 3 seconds each: A - creator slams a French press into a trash can, the glass-shatter sound is the whole punchline, on-screen text just says 'enough.' B - first frame shows the concentrate bottle pouring over a cup of ice, text 'Cold brew in 10 seconds. $34 kit.' C - close-up of a pale watery iced coffee, text 'Your cold brew is watered down. Here's why.' D - the exact same watery-iced-coffee footage as C, new text 'Stop drinking watered-down coffee.'",
      "expected_output": "A scorecard that refuses numeric scores and absolute predictions, runs the four hard gates, fails A on sound-off legibility, merges C and D into one cell at the real-variation gate, pairwise-ranks B top, attaches cheapest-rung revision notes, and flags the batch as below the 3-opening shipping floor.",
      "files": [],
      "expectations": [
        "Refuses to output numeric hook scores: no X/10 ratings, decimals, percentages, or composite score appears for any candidate",
        "Uses only qualitative strong/adequate/weak bands per scoring dimension",
        "Declines to predict an absolute hook-rate number for any candidate, framing the output as a relative shortlist valid within this batch only",
        "Records candidate A as failing the sound-off legibility gate because its punchline lives in the shatter audio and Meta feed autoplays muted",
        "Does not rank A at the top of the batch despite its scroll-stopping power, because a gate-failing candidate cannot rank top",
        "Merges C and D into a single test cell at the real-variation gate because they share the same opening visual and differ only in wording",
        "States the batch's real distinct-candidate count is lower than 4 after the C/D merge",
        "Derives the ranking by pairwise head-to-head comparison of candidates, not by tallying or summing bands into a total",
        "Places B at the top of the shortlist, on grounds including its specific checkable claim and price and its sound-off legibility",
        "Reports per-candidate gate results (pass/fail with a reason) for all four candidates",
        "Gives each non-top candidate a raise-the-rank note taken from the cheapest revision rung that removes its specific failure",
        "Identifies A's fix as structural (a rebuild or new footage that must re-enter the gates), not an on-screen-text rewrite",
        "States the batch sits below the 3-opening shipping floor (fewer than 3 gate-clearing, concept-distinct openings) and should not ship as-is"
      ]
    },
    {
      "id": 2,
      "prompt": "Marketing lead at Fernwood Pet Supplements here. Our agency's quarterly report says our TikTok ads average a 31% hook rate while our Meta ads only hit 22%, so they want to rebuild all our Meta openings in a 'TikTok style' to close the gap. They also benchmarked us against an 'industry standard 30% hook rate' from a slide they found. We run feed placements on both platforms, cold B2C traffic. Before I approve the rebuild budget - does the data actually support this?",
      "expected_output": "An analysis showing the 31% vs 22% comparison is invalid because the two platforms construct hook rate from different numerators and denominators, rejecting the sourceless 30% benchmark, and redirecting the account to one locked definition and within-account calibration.",
      "files": [],
      "expectations": [
        "States that hook rate is a practitioner-constructed ratio, not a native platform metric",
        "States TikTok's common hook-rate construction as 2-second views divided by impressions",
        "States Meta's common hook-rate construction as 3-second video plays divided by impressions",
        "Concludes the 31% vs 22% comparison is invalid because the two numbers are different metrics with different windows and constructions, not evidence TikTok hooks are better",
        "Does not endorse the Meta rebuild on the basis of the cross-platform comparison",
        "Rejects the 'industry standard 30% hook rate' benchmark as unusable without its denominator, placement, traffic temperature, and source",
        "Recommends locking one hook-rate definition per account, writing it on every scorecard, and calibrating only within it",
        "States that hook rate measures attention, not business performance",
        "Cites the near-zero-to-negative correlation between hook rate and ROAS (approximately -0.19, hold rate -0.10)",
        "Recommends judging openings by within-account calibration (predicted vs measured rank agreement on one definition) rather than external benchmarks",
        "Asks for or verifies the exact numerator and denominator each dashboard uses before deeper analysis",
        "Produces no numeric hook score or predicted hook-rate figure for any proposed new opening"
      ]
    },
    {
      "id": 3,
      "prompt": "Head of demand gen at Ledgerline, a mid-market contract-management platform, deals around $40k ACV, 9-month sales cycle. We're launching cold LinkedIn feed video ads; the objective is pipeline over the next few quarters. My agency keeps pushing to make the openings 'more urgent' with limited-time demo offers to drive immediate clicks, and to hold our logo to the end card so the ads don't look like ads. Three candidates, first 3 seconds each: X - on-screen text 'Managing 200+ contracts in spreadsheets?' over a screen recording of a contract tracker, our brand colours framing the shot. Y - fast-cut stock footage of a lawyer ripping up papers, text 'The mistake costing companies millions.' Z - a countdown timer with text 'Book a demo this week - 20% off.' Which do we run? One data point: Y's footage already got 40k views on an organic LinkedIn post.",
      "expected_output": "A B2B-mode scorecard grounded in the 95:5 rule: X ranked top for role-and-situation qualification, Y banded weak on qualification, Z's urgency angle rejected for an out-of-market audience, early distinctive-asset branding defended against the end-card plan, and the 40k view count discounted as a weak 2-second signal.",
      "files": [],
      "expectations": [
        "States explicitly that the batch is scored in B2B mode",
        "Cites the 95:5 rule (roughly 95% of B2B buyers out-of-market at any moment) when framing the objective",
        "States the opening should aim at memory and future buyers rather than an immediate click, given the 9-month cycle and pipeline objective",
        "Rejects or bands weak the limited-time urgency angle of Z as a direct-response heuristic that transfers poorly to a mostly out-of-market B2B audience",
        "Ranks X top (or bands X strong on qualification) because it qualifies by role and situation - naming the 200+-contracts-in-spreadsheets buyer",
        "Bands Y weak on audience qualification because its universal curiosity stops viewers of every kind, not the buyer",
        "Pushes back on holding all branding to the end card under a memory objective, citing the brand-linkage cost of end-loaded branding (the TikTok-reported 17-point drop) or the early distinctive-asset evidence",
        "Distinguishes early distinctive-asset brand cues (recommended here) from a logo-card opening (an anti-pattern), rather than treating branding as all-or-nothing",
        "Discounts the 40k LinkedIn view count because a LinkedIn view is only 2 continuous seconds with about 50% of the player on screen - an inflated, weak attention signal",
        "Recommends calibrating this batch on completion or watch-time metrics rather than LinkedIn view counts",
        "Notes sound-off legibility still applies because LinkedIn autoplays muted, judging the candidates on visuals and on-screen text alone",
        "Delivers a pairwise ranking with reasons grounded in gate results and bands",
        "Produces no numeric scores and no absolute view-rate or hook-rate prediction"
      ]
    },
    {
      "id": 4,
      "prompt": "Video lead at Copperleaf Analytics. We have 5 openings for a Q4 push and I want one ranked list, best to worst, so I know where to put budget. Two will run Meta feed, two are YouTube skippable in-stream, one is a 6-second YouTube bumper. All I have right now is the voiceover transcript of each opening. 1 (Meta feed): 'If your dashboards take a week to update, listen up.' 2 (Meta feed): 'Here's how we cut reporting time by 90 percent.' 3 (YouTube skippable): 'Before you skip - what if your data warehouse bill dropped by half?' 4 (YouTube skippable): 'Meet Copperleaf. We were founded in 2019 with a simple idea.' 5 (bumper): 'Copperleaf. Analytics your CFO trusts.' Rank all five against each other.",
      "expected_output": "A response that refuses a single cross-placement ranking, normalises each candidate to its placement's hook window, scores the bumper on continuity/branding/qualification only, caps sound-off bands at adequate on transcript-only input, and requests first-frame and on-screen-text detail.",
      "files": [],
      "expectations": [
        "Refuses to produce one unified 1-to-5 ranking spanning the three placements, because candidates in different windows are not one comparison",
        "Normalises the two Meta feed candidates to roughly the first 3 seconds",
        "Applies the roughly 5-second window (up to the skip button) to the two YouTube skippable candidates",
        "Treats the bumper as a forced view and states that stop-the-scroll hook logic does not apply to it",
        "Scores or proposes to score the bumper only on continuity, branding, and audience qualification, and says so explicitly",
        "States that transcript-only input cannot support a strong sound-off legibility band, banding it at best adequate",
        "Requests the missing first-frame, visual, or on-screen-text information as the input that would resolve the ambiguous bands",
        "Uses strong/adequate/weak bands with no numeric scores anywhere",
        "Notes the counting difference across formats (for example that a paid YouTube skippable view counts at 30 seconds, completion, or interaction) rather than proposing one hook-rate readout across all five",
        "Flags candidate 4's 'Meet Copperleaf, founded in 2019' opening as a brand-announcement or slow-build anti-pattern that spends the highest-attention seconds on the advertiser, or bands it weak on time-to-signal",
        "Compares candidates head-to-head only within the same placement group (Meta pair against each other, skippable pair against each other)",
        "Gives per-candidate raise-the-rank notes or input requests instead of an absolute performance prediction for any candidate"
      ]
    },
    {
      "id": 5,
      "prompt": "Our creative strategist at Halyard Swim built a hook evaluation spreadsheet and wants the team to adopt it. Her system weights the seven things we judge on every opening: time-to-signal 30%, specificity 25%, sound-off legibility 15%, audience qualification 10%, brand timing 10%, promise continuity 5%, placement fit 5% - then sums to a 0-100 score per hook. She also tagged our current 9-opening batch against the eight psychological hook types and is proud we cover 7 of the 8, which she says proves the batch is diverse enough to ship. Sanity-check the system before we roll it out?",
      "expected_output": "A refusal to endorse the weighting or the composite score, grounded in the evidence that no dimension ordering is supported; taxonomy coverage rejected as a diversity measure; the diversity question redirected to the real-variation gate and the 3-opening floor.",
      "files": [],
      "expectations": [
        "Rejects the percentage weighting of the seven scoring dimensions, refusing to endorse any importance order or weights among them",
        "Gives the evidential reason: attention metrics correlate roughly -0.19 with ROAS (hold rate -0.10), so no evidence supports one dimension returning more performance per unit of effort than another",
        "States that a printed weighting would steer real production hours and test budget on invented precision",
        "Rejects the 0-100 composite sum: bands are evidence for pairwise comparison, never arithmetic to total",
        "States the only defensible dimension weighting is account-specific, re-derived from the account's own past winners after calibration fails - never a default written into a system",
        "Keeps pairwise candidate ranking as the legitimate ranking and explains why: the same judge compares two concrete openings so the judge's bias applies to both sides and cancels",
        "Rejects '7 of 8 hook types covered' as evidence the batch is diverse or ready to ship",
        "States that hook taxonomies are naming and diagnostic vocabulary, never a scoring axis",
        "Cites poor inter-rater reliability and overlapping, non-falsifiable category boundaries as the reason taxonomy tagging cannot be scored",
        "Redirects the diversity question to the real-variation gate (distinct opening visuals, wording-only siblings merged) and the 3-opening concept-distinct shipping floor",
        "Does not offer a compromise such as adopting her weights as a starting point - the refusal to rank dimensions stands as the finding",
        "Distinguishes what may legitimately be ranked - candidates pairwise, and revision options by cost - from the banned ordering of scoring dimensions"
      ]
    },
    {
      "id": 6,
      "prompt": "Growth lead at Nettleship, a B2C meal-planning app at $12.99/month, Meta feed, cold traffic. Launch is locked 6 days out. Our video partner's shoot lead time is 3 weeks, and the creator contracts from the last shoot have expired - legal says re-signing takes at least 2 weeks. What we do have: a full-time in-house editor and the complete raw footage library from last quarter's shoot. Four candidates, first 3 seconds each: E opens on our founder talking to camera, the whole pitch is in what she says, first frame is just her face with no text - though later in the same cut there's a screen capture of the app auto-building a week of meals. F opens on a fast recipe montage but the text line sits at the very bottom of the frame where it was positioned for TikTok, and Meta feed crops it. G opens on a lunchbox-packing shot with text 'The smarter way to eat healthy.' H is the same lunchbox-packing footage as G with text 'Meal prep without the Sunday marathon.' Give me the most thorough fix for each so we only have to do this once.",
      "expected_output": "Revision notes chosen by cheapest rung that removes each failure rather than thoroughness: text rewrite for G, caption/safe-zone fix for F, re-cut for E and for H's visual difference, reshoot and rebuild rungs deleted (not demoted) by the locked date and expired contracts, with the ladder re-ranked around the in-house editor and owned footage.",
      "files": [],
      "expectations": [
        "Rejects the 'most thorough fix' framing and picks each revision by rank movement per hour - the cheapest rung that actually removes the failure",
        "Recommends an on-screen-text rewrite (the cheapest rung) for G's generic 'smarter way to eat healthy' category language",
        "Recommends a text/caption-layer fix for F: moving the text out of the TikTok safe-zone position that Meta feed crops",
        "Records E as failing the sound-off legibility gate, with the reason that its meaning lives entirely in the founder's spoken words on a muted feed",
        "Recommends a re-cut for E that opens on the existing app screen-capture footage, rather than any fix requiring new footage",
        "States G and H are currently one test cell - a wording-only sibling merged at the real-variation gate because the opening visual is identical",
        "Recommends a re-cut from the existing footage library to give H a genuinely different opening visual, or drops/merges H",
        "Deletes the shoot-a-new-opening-beat and rebuild rungs from this batch's menu rather than listing them as lower-priority options",
        "Names the constraint that deleted them: the 6-day launch sits inside the 3-week shoot lead time and/or the expired creator contracts",
        "Notes the in-house editor and owned footage library collapse the effort of re-cuts, re-ranking the revision ladder against this account's reality",
        "Mentions the compliance exposure of the relevant rungs: a reshoot reopens creator usage rights, and/or a rewrite adding a checkable claim triggers claim substantiation and platform ad review",
        "Every delivered revision note comes from the text-rewrite, caption-fix, or re-cut rungs only - no recommendation requires new footage",
        "Checks the batch against the 3-opening shipping floor and states the result explicitly"
      ]
    },
    {
      "id": 7,
      "prompt": "Solstice Luggage here - DTC, $265 polycarbonate carry-on. Our star TikTok ad opens on a guy dropping a watermelon off a parking garage; it splatters in glorious slow motion, and the suitcase doesn't appear until second 12. That ad pulls a 41% hook rate, miles above everything else we run - but its ROAS is 0.8 and sales didn't move. My CMO's conclusion: the hook is clearly our best ever, so make 10 more exactly like it - same stunt footage, ten new first lines. I've got the 10 replacement text lines drafted. Rank which of the 10 to produce first.",
      "expected_output": "A refusal to rank the 10 as separate candidates: they are one cell over one visual; diagnosis that the stunt fails qualification and promise-payoff, that 41% hook rate is attention rather than performance, and a recommendation toward concept-distinct, buyer-qualifying openings at the right revision rung.",
      "files": [],
      "expectations": [
        "Does not rank the 10 text lines as 10 separate candidates",
        "States the 10 wording-only variations over the same stunt visual are a single test cell, because the real-variation gate requires the opening visual to differ",
        "Diagnoses the star opening as failing audience qualification: universal spectacle that stops everyone qualifies no one for a $265 carry-on",
        "Diagnoses a promise-payoff problem: the hook window (splattering watermelon, product absent until second 12) does not set up what the ad sells",
        "Cites the documented high-hook-rate/mediocre-ROAS failure pattern - the hook stops the scroll but nothing downstream converts",
        "Invokes the qualified-attention warning (higher hook rate or thumbstop does not mean better performance; clickbait gets attention without sales)",
        "States hook rate measures attention rather than performance, citing the roughly -0.19 correlation with ROAS",
        "Explicitly declines to treat the 41% hook rate as evidence the opening works",
        "Identifies the fault as structural - the hook-window footage does not implicate the product or buyer - so the fix sits at the re-cut or new-footage rung, not a first-line rewrite",
        "Recommends the next batch contain genuinely distinct concepts that qualify the buyer, meeting the 3-opening concept-distinct floor",
        "Frames any recommendation as within-batch test-budget shortlisting, with no absolute hook-rate prediction for any new variant",
        "States or verifies the TikTok hook-rate construction (2-second views divided by impressions) when discussing the 41% figure"
      ]
    },
    {
      "id": 8,
      "prompt": "We've been pre-ranking our hook batches at Quill & Anchor Stationery before testing for two quarters, all on Meta feed. Five batches so far - here's how our pre-test #1 pick finished in the measured results: batch 1, 1st of 5; batch 2, 4th of 6; batch 3, 2nd of 4; batch 4, 5th of 5; batch 5, 3rd of 6. My boss says anything short of picking the exact winner every time means the pre-ranking is useless and we should drop it. One wrinkle: our analyst pulled the measured 'hook rate' for batches 1-3 from our old dashboard, which computes ThruPlays divided by impressions, and batches 4-5 from the new one, which computes 3-second plays divided by impressions. So - is the system working?",
      "expected_output": "A verdict framed as rank agreement against the 3-of-5 top-half threshold rather than exact-winner accuracy, with the ThruPlay-based readouts flagged as a hold-rate construction contaminating the calibration series, and a recommendation to recompute on one locked definition.",
      "files": [],
      "expectations": [
        "Corrects the boss's criterion: the outcome measure is rank agreement, not absolute accuracy or picking the exact winner",
        "States the pass threshold: over a rolling window of 5 batches, the top-ranked opening lands in the top half of the measured ranking at least 3 times",
        "Correctly classifies the batches against top-half: 1st of 5 yes, 4th of 6 no, 2nd of 4 yes, 5th of 5 no, 3rd of 6 yes - a 3-of-5 tally",
        "Defends the deliberately modest bar: it is what a useful shortlisting device clears and a coin flip does not",
        "Flags that batches 1-3 were measured on ThruPlays divided by impressions - a completion/hold-style construction, not the hook-rate definition",
        "States the two dashboards' numbers are different metrics, so a calibration series mixing them is contaminated rather than one comparable record",
        "Recommends recomputing batches 1-3's measured rank orders under the 3-second-plays definition (or excluding them) before trusting the 3-of-5 verdict",
        "Recommends locking one hook-rate definition for the account and writing it on every scorecard",
        "Does not conclude the pre-ranking system failed based on the exact-winner standard",
        "Explains the below-threshold consequence: stop trusting the generic rubric and re-derive the dimension weighting from the account's own past winners",
        "Conditions the verdict on same platform and placement (confirms all five batches ran Meta feed under one placement)",
        "Recommends recording the predicted rank order plus the dashboard's exact hook-rate definition at scoring time for every future batch",
        "Frames pre-ranking as a shortlisting device deciding what gets test budget, with the market as the judge - not a predictor graded on picking winners"
      ]
    },
    {
      "id": 9,
      "prompt": "Our new brand manager at Tidepool Skincare circulated a 'hook standards' memo the whole team must follow, sourced from a conference talk: (1) 'You have 0.25 seconds to capture attention, so the brand logo must appear in the first quarter-second.' (2) '85% of Facebook video is watched without sound.' (3) 'The average human attention span is 8 seconds, down from 12.' (4) 'People spend 1.7 seconds per piece of content, so any hook longer than 1.7 seconds fails.' I've been asked to rewrite our hook-review checklist around these four rules by Friday. Are they solid?",
      "expected_output": "Each folklore figure corrected to its real origin, the logo-first rule rejected as an anti-pattern, sound-off design re-grounded on muted autoplay, and a replacement checklist built on the sourced attention-decay and memory-threshold data with gates and bands instead of folklore thresholds.",
      "files": [],
      "expectations": [
        "Corrects rule 1: the 0.25-second figure is a 2016 recall-at-exposure finding (people could recall feed content after 0.25 seconds of exposure), not an attention deadline",
        "Rejects the logo-in-the-first-quarter-second rule: a logo-first opening signals 'this is an ad' at the worst moment and is a documented anti-pattern",
        "Corrects rule 2: the 85%-without-sound figure is a 2016 publisher self-reported snapshot that the platform never confirmed, not a current or verified number",
        "States the sound-off design constraint still stands - grounded on muted autoplay on Meta and LinkedIn, not on the 85% figure",
        "Corrects rule 3: the 8-second attention span claim is debunked with no credible primary source",
        "Corrects rule 4: 1.7 seconds is a 2016 platform average for time spent with mobile content - a decade-old average, not a rule or a hook-length ceiling",
        "Offers the well-sourced replacements: attention-decay data of roughly 80% of the audience present in second 1, 50% by second 2, 20% by second 3, and/or the roughly 2.5-second memory-encoding threshold (about 1.5 seconds with distinctive brand assets)",
        "Keeps sound-off legibility as a hard pass/fail check in any rewritten checklist",
        "Adopts none of the four memo rules unchanged into the rewritten checklist",
        "Frames brand timing as an objective-dependent tension: early distinctive-asset brand cues differ from a logo-card open, and withholding branding to the end carries a brand-linkage cost (the reported 17-point drop)",
        "Cites the distinctive-brand-asset evidence (memory encoding in as little as 1.5 seconds when assets are deployed well) when discussing early branding",
        "Builds replacement checklist items as gates and anchored qualitative bands rather than numeric thresholds derived from the folklore figures"
      ]
    },
    {
      "id": 10,
      "prompt": "Bootstrap founder at Marrowbone Broth Co. Tiny budget: $1,800 total for our first Meta feed test, cold traffic. I have two openings and can't afford to make more: one is a close-up slow pour of broth with on-screen text 'Bone broth that actually tastes good. $3 a serving.' The other is the exact same pour footage with text 'Finally, bone broth you'll crave.' My plan: you tell me which one will win, predict its hook rate, I put the whole $1,800 behind that one, and we're off. Deal?",
      "expected_output": "A refusal of the prediction and the one-ad bet: the two openings are one cell, the batch is below the 3-opening floor given the roughly 5% winner base rate, and the recommendation is at least 3 gate-clearing concept-distinct openings, with production questions asked before prescribing how to get there.",
      "files": [],
      "expectations": [
        "Declines to predict an absolute hook-rate number for either opening",
        "States the two openings are a single test cell: identical visual, wording-only difference, merged at the real-variation gate",
        "States the plan is below the 3-opening shipping floor and should not launch as-is",
        "Cites the base rate that only roughly 5% of creatives become winners as the reason volume produces winners",
        "States that a one-opening launch is a bet, not a test",
        "Recommends reaching at least 3 gate-clearing openings that differ on the concept axis before spending the budget",
        "States that polishing or rewording the existing two openings cannot produce the missing concepts - new concepts are required, and they re-enter the gates",
        "Asks about, or conditions its recommendation on, what other footage and production capacity exist before prescribing how to produce the additional concepts",
        "Does not simply pick a winner between the two text lines and endorse the full-budget single-ad plan",
        "Treats any preference between the two text lines as a wording choice inside one cell, not a ranking of two candidates",
        "Any text-level preference favours the '$3 a serving' line on specificity grounds - a concrete checkable claim versus a generic benefit statement",
        "Frames its role as deciding what deserves test budget with the market as judge, not as predicting the winner"
      ]
    }
  ],
  "trigger_queries": [
    { "query": "Which of these 5 video ad hooks should we put test budget behind?", "should_trigger": true },
    { "query": "Score these ad openings for me - Meta feed, cold traffic", "should_trigger": true },
    { "query": "Is this a good hook? First 3 seconds: close-up of a cracked phone screen, text 'your case failed you'", "should_trigger": true },
    { "query": "Rank these TikTok ad hooks", "should_trigger": true },
    { "query": "My media buyer says our hook rate is bad - how do I even judge our hooks before spending?", "should_trigger": true },
    { "query": "We shot 6 intro variations for our new Meta ad. Which opening gets the budget?", "should_trigger": true },
    { "query": "Can you review the first 3 seconds of this ad script?", "should_trigger": true },
    { "query": "Which of these storyboards has the strongest opening for a cold audience?", "should_trigger": true },
    { "query": "Thumbstop check: does this opening actually stop the scroll?", "should_trigger": true },
    { "query": "How do I evaluate the first 5 seconds of our skippable YouTube pre-roll?", "should_trigger": true },
    { "query": "We're running 6-second bumpers - how should I judge the openings?", "should_trigger": true },
    { "query": "Hook rate on our dashboard - what's actually in the numerator and denominator?", "should_trigger": true },
    { "query": "Is 25% a good hook rate for Meta feed?", "should_trigger": true },
    { "query": "I have 8 openings but budget to test 3 - which ones?", "should_trigger": true },
    { "query": "Does this hook work for B2B on LinkedIn?", "should_trigger": true },
    { "query": "attention scoring for our video ad intros", "should_trigger": true },
    { "query": "Which hook wins: the confession one or the offer-first one?", "should_trigger": true },
    { "query": "Our agency sent 4 hook options for the spring campaign - help me pick", "should_trigger": true },
    { "query": "First-frame check on these ad cuts before we launch?", "should_trigger": true },
    { "query": "Two of my hooks feel identical - do they count as different tests?", "should_trigger": true },
    { "query": "Grade the opening of this UGC ad cut", "should_trigger": true },
    { "query": "Here are transcripts of 5 ad openings - which deserve budget?", "should_trigger": true },
    { "query": "Will people keep watching past second 3 of this ad? Storyboard attached", "should_trigger": true },
    { "query": "My CMO wants a hook score for every creative before launch - set that up", "should_trigger": true },
    { "query": "Compare these two ad openings and tell me which to fund", "should_trigger": true },
    { "query": "Which opening qualifies our buyer best?", "should_trigger": true },
    { "query": "Shortlist these 12 hooks down to the ones worth testing", "should_trigger": true },
    { "query": "hook analysis for our Q4 video batch", "should_trigger": true },
    { "query": "The first shot of our new ad is a logo animation - is that a problem?", "should_trigger": true },
    { "query": "Evaluate this ad's cold open", "should_trigger": true },
    { "query": "Our hook rate went from 18% to 37% but sales are flat - what does that say about our hooks?", "should_trigger": true },
    { "query": "Which of these openings survives being watched on mute?", "should_trigger": true },
    { "query": "How many seconds do I have before viewers bail on my ad intro?", "should_trigger": true },
    { "query": "Pick the best cold open from these five shot lists", "should_trigger": true },
    { "query": "Vet the hooks in this batch before we brief the editor", "should_trigger": true },
    { "query": "Does opening on our founder talking to camera kill the ad?", "should_trigger": true },
    { "query": "Hold rate vs hook rate - which should I use to judge my openings?", "should_trigger": true },
    { "query": "Need a second opinion on which ad intro to move into testing", "should_trigger": true },
    { "query": "Which hook should my $2k test budget go on?", "should_trigger": true },
    { "query": "Are these 4 hooks different enough to test separately?", "should_trigger": true },
    { "query": "Review the opening beat of our new spot", "should_trigger": true },
    { "query": "Our video ads all start with 5 seconds of vibes before the point - how bad is that?", "should_trigger": true },
    { "query": "Which of these Instagram Reels ad openings would you fund?", "should_trigger": true },
    { "query": "Rate my hook: 'POV: your payroll just failed the audit'", "should_trigger": true },
    { "query": "Help me force-rank these ad openings", "should_trigger": true },
    { "query": "3-second test: do these ads pass?", "should_trigger": true },
    { "query": "Which hook variant goes to the top of the test queue?", "should_trigger": true },
    { "query": "Assess whether this ad opening will stop the right people", "should_trigger": true },
    { "query": "We disagree internally about which hook is strongest - settle it", "should_trigger": true },
    { "query": "Screen these 7 ad openings before the media plan locks", "should_trigger": true },
    { "query": "What makes this hook weak? Script attached", "should_trigger": true },
    { "query": "My ads get skipped in the first seconds on YouTube - evaluate these new openings", "should_trigger": true },
    { "query": "Hook audit for the batch we're launching Monday", "should_trigger": true },
    { "query": "We run separate cold and retargeting campaigns - which of these two openings fits the cold one?", "should_trigger": true },
    { "query": "Does the first 3 seconds of this demo-style ad earn the next 27?", "should_trigger": true },
    { "query": "Choose between these two cold opens for our SaaS ad", "should_trigger": true },
    { "query": "I pasted 4 ad scripts below - tell me which intros are worth money", "should_trigger": true },
    { "query": "Quick gut-check on this ad's opening frame and text overlay", "should_trigger": true },
    { "query": "Which of our new hooks would you kill before testing?", "should_trigger": true },
    { "query": "Prioritize these video openings for next week's creative test", "should_trigger": true },
    { "query": "Write me 10 hooks for our new UGC ad", "should_trigger": false },
    { "query": "Draft a UGC script with 5 hook variants for our skincare serum", "should_trigger": false },
    { "query": "Give me hook ideas for a meal-kit TikTok ad", "should_trigger": false },
    { "query": "Turn this product brief into a short-form video ad script", "should_trigger": false },
    { "query": "Write the opening line for our new video ad", "should_trigger": false },
    { "query": "Brainstorm scroll-stopping openings for our creator to film", "should_trigger": false },
    { "query": "I need 3 new hook concepts written for next month's batch", "should_trigger": false },
    { "query": "Script a talking-head ad with a strong cold open", "should_trigger": false },
    { "query": "Which of these 5 Facebook ad headlines is strongest?", "should_trigger": false },
    { "query": "Rank these primary-text variants for our Meta ads", "should_trigger": false },
    { "query": "Write headline and CTA variants for our search ads", "should_trigger": false },
    { "query": "Score my static image ad copy", "should_trigger": false },
    { "query": "Which ad description performs better for carousel ads?", "should_trigger": false },
    { "query": "Give me 10 CTA button variations to test", "should_trigger": false },
    { "query": "Our best ad's hook rate dropped from 32% to 19% over three weeks - why?", "should_trigger": false },
    { "query": "This creative has been running 2 months and CTR is sliding - is it worn out?", "should_trigger": false },
    { "query": "Why did our winning video ad stop performing?", "should_trigger": false },
    { "query": "Frequency is at 6 and results are fading - do we need new creative?", "should_trigger": false },
    { "query": "Diagnose why our top ad's performance is decaying", "should_trigger": false },
    { "query": "Our hook rate on the running campaign fell off a cliff last week", "should_trigger": false },
    { "query": "How much budget per cell do I need to test 4 creatives?", "should_trigger": false },
    { "query": "Design an A/B test for our new video ads", "should_trigger": false },
    { "query": "What sample size do I need to compare two hooks properly?", "should_trigger": false },
    { "query": "Set up kill and scale rules for our creative test next month", "should_trigger": false },
    { "query": "How long should a creative test run before we call it?", "should_trigger": false },
    { "query": "Build the testing matrix for our Q1 creative sprint", "should_trigger": false },
    { "query": "Put together a creative brief for our video editor", "should_trigger": false },
    { "query": "Write a brief with hook directions for the UGC creator we hired", "should_trigger": false },
    { "query": "What specs and messaging should go in the brief for our next ad shoot?", "should_trigger": false },
    { "query": "Turn this campaign goal into a brief a designer can execute", "should_trigger": false },
    { "query": "Collect our competitors' best hooks into a swipe file", "should_trigger": false },
    { "query": "Categorize these 30 competitor ads by hook type", "should_trigger": false },
    { "query": "What hooks are competitors in our category running right now?", "should_trigger": false },
    { "query": "Build a library of ad references sorted by hook style", "should_trigger": false },
    { "query": "Should we run carousel or video for this campaign?", "should_trigger": false },
    { "query": "Which ad format fits a lead-gen objective on LinkedIn?", "should_trigger": false },
    { "query": "Is a 6-second bumper the right format for our launch?", "should_trigger": false },
    { "query": "Our ads get clicks but the landing page doesn't convert - audit it", "should_trigger": false },
    { "query": "Why do people bounce right after clicking our ad?", "should_trigger": false },
    { "query": "Check if our landing page matches the ad's message", "should_trigger": false },
    { "query": "Our whole account's ROAS tanked this quarter - figure out why", "should_trigger": false },
    { "query": "CPA doubled across all campaigns, where do I start?", "should_trigger": false },
    { "query": "Everything in the ad account is underperforming - diagnose it", "should_trigger": false },
    { "query": "Write a hook for my LinkedIn post about hiring", "should_trigger": false },
    { "query": "Which email subject line will get more opens?", "should_trigger": false },
    { "query": "Give me a strong opening paragraph for my blog post", "should_trigger": false },
    { "query": "How do I hook listeners in my podcast intro?", "should_trigger": false },
    { "query": "Best cold-call opening line for booking meetings?", "should_trigger": false },
    { "query": "My YouTube videos lose viewers in the first 30 seconds - fix my intros", "should_trigger": false },
    { "query": "Rate the hook of my Twitter thread", "should_trigger": false },
    { "query": "Write a webinar opening that keeps people from dropping off", "should_trigger": false },
    { "query": "When should I use useEffect vs a custom React hook?", "should_trigger": false },
    { "query": "My Shorts get no views past 2 seconds - improve my organic content hooks", "should_trigger": false },
    { "query": "What's the hook for our brand story in the pitch deck?", "should_trigger": false },
    { "query": "Which audiences should we target for the new video campaign?", "should_trigger": false },
    { "query": "Should we promote our CEO's post as a thought leader ad?", "should_trigger": false },
    { "query": "How do I pick seed customers for a lookalike audience?", "should_trigger": false },
    { "query": "Is our conversion tracking firing correctly before launch?", "should_trigger": false },
    { "query": "Split $50k across Meta and TikTok for next quarter", "should_trigger": false },
    { "query": "What's a healthy ROAS for a $40 AOV brand?", "should_trigger": false }
  ]
}
references/examples.md›
# Worked Examples

Fictional products and numbers-free banding throughout - these show the output shape and the reasoning, not real benchmark data.

## Scorecard Shape

Every scorecard follows this field-by-field shape, filled in per batch:

```
HOOK BATCH SCORECARD - <product/offer>, <platform + exact placement>, <date>
batch        : <N> candidates | input: <script|transcript|storyboard|shot list|cut description>
hook window  : first <N>s | audience: <cold|warm> <B2B|B2C>
dashboard    : hook rate = <numerator> / <denominator> (from Interview)

gates
  <candidate>: sound-off <pass|FAIL> | promise-payoff <pass|FAIL> | qualification <pass|FAIL> | variation <pass|FAIL - merged with X>

bands (gate-passing candidates only; strong/adequate/weak)
  <candidate>: time-to-signal <band> | sound-off <band> | qualification <band> | specificity <band>
               | brand timing <band - objective it serves> | continuity <band> | placement fit <band>

ranking (pairwise, valid within this batch only - not a performance prediction)
  1. <candidate> - <why it won its head-to-heads>
  2. <candidate> - ...
  raise-the-rank
  <candidate>: <one concrete change that would move it up>
  rungs deleted: <rung> - <constraint that removed it from this batch's menu>

shipping check : <N> openings clear all gates and differ on concept - <meets|BELOW> the 3-opening floor
calibration    : after launch, compare this predicted order against measured hook-rate order
                 (same definition, platform, placement); log rank agreement
```

## Example 1 - B2C batch, worked in full

- Product: a $79 dishwasher-safe cast-iron skillet, DTC brand.
- Platform: Meta, feed placements, cold prospecting, B2C.
- Input: scripts with first-frame descriptions.
- Dashboard hook rate: `3-second video plays ÷ impressions` (confirmed in Interview).
- Rest of ad (all four candidates share it): 20 seconds of dishwasher demo, seasoning-free cooking shots, offer card with 30-day guarantee.

The four candidates' openings (first 3 seconds):

- **A "Offer open"**: skillet placed into a running dishwasher; on-screen text "Cast iron. Dishwasher safe. $79."
- **B "Kitchen chaos"**: creator theatrically shoves an entire pot rack off a counter, crash sound carries the joke; text "I'm done with all of it."
- **C "Scrub confession"**: close-up of hands scrubbing a rusted skillet; text "Love cast iron. Hate babying it?"
- **D "Scrub confession v2"**: same rusted-skillet footage as C; new text line "Cast iron shouldn't be a chore."

```
HOOK BATCH SCORECARD - $79 dishwasher-safe cast-iron skillet, Meta feed, 2026-08-26
batch        : 4 candidates | input: scripts with first-frame descriptions
hook window  : first 3s | audience: cold B2C
dashboard    : hook rate = 3-second video plays / impressions

gates
  A: sound-off pass | promise-payoff pass | qualification pass | variation pass
  B: sound-off FAIL (joke lives in the crash audio; muted, it reads as clumsiness)
     | promise-payoff pass | qualification FAIL (stops everyone; nothing signals cookware buyer)
     | variation pass
  C: sound-off pass | promise-payoff pass | qualification pass | variation pass
  D: sound-off pass | promise-payoff pass | qualification pass
     | variation FAIL - merged with C (identical visual; wording-only sibling is not a test cell)

bands (gate-passing candidates only; strong/adequate/weak)
  A: time-to-signal strong | sound-off strong | qualification adequate | specificity strong
     | brand timing strong - DR objective, product-as-brand-cue, no logo card
     | continuity strong | placement fit strong
  C: time-to-signal strong | sound-off strong | qualification strong | specificity adequate
     | brand timing adequate - brand only implied until the demo
     | continuity strong | placement fit strong

ranking (pairwise, valid within this batch only - not a performance prediction)
  1. A - vs C: both signal fast and hold continuity; A wins on specificity (a checkable
     claim plus the price) and on being the offer in miniature. Motion's 2026 data is a
     tiebreaker, not proof: offer-only hooks had the highest hit rate in its BFCM-window
     dataset.
  2. C - qualifies the cast-iron owner more sharply than A, but its promise is abstract
     until the demo arrives.
  3. B - cannot rank top on two gate failures regardless of its scroll-stopping power.
  raise-the-rank (cheapest rung that removes the failure)
  A: [rung 1, text rewrite] name the pain A assumes ("no seasoning, no rust") in the text
     to close the qualification gap with C.
  C: [rung 1, text rewrite] add one checkable element (the price, or "survives 1,000
     cycles") to the text - substantiate the cycle claim before it runs.
  D: [rung 3, re-cut] open on the rust rather than the scrubbing, from footage already
     shot - that is the visual difference the variation gate wants; drop D otherwise.
  B: [rung 5, rebuild] make the chaos read muted AND implicate cookware in frame 1 - two
     gate failures, so this is a new candidate re-entering the gates, not a fix.

shipping check : 2 openings clear all gates and differ on concept - BELOW the 3-opening
                 floor. Do not ship as-is: revise B and D (or add a new concept) first.
calibration    : after launch, compare this order (A, C, +revisions) against measured
                 hook-rate order, same definition and placement; log rank agreement.
```

## Example 2 - B2B contrast, condensed

- Product: payroll-compliance platform for companies employing across borders.
- Platform: LinkedIn feed, cold, B2B.
- Objective: stated as memory/pipeline, not immediate signup.

Two candidates:

- **P "Role call-out"**: static-feeling shot of a payroll dashboard mid-error, brand's distinctive colour frame; text "Running payroll in 3 countries?"
- **Q "Big number tease"**: dramatic stock footage of a shredded contract; text "The $2M mistake nobody talks about."

Both pass sound-off, continuity, and variation. Banding highlights - and where B2B diverges from Example 1:

- **Identical logic to B2C**: time-to-signal (both land inside 3s: strong), sound-off (both text-carried: strong), continuity (both set up the compliance story: strong).
- **Divergent - qualification**: P qualifies by role and situation, exactly what the 95:5 logic wants when ~95% of buyers are out-of-market: strong. Q's curiosity is universal - finance-adjacent viewers of every kind stay: weak, and on a cold B2B budget that is close to a gate question.
- **Divergent - brand timing**: with a memory objective, P's early distinctive-asset branding bands strong; Q holds all branding to the end card - weak against a 17-point brand-linkage drop reported by TikTok for end-loaded branding (asserted figure, but the direction matches the memory objective).
- **Metric caution on the scorecard**: LinkedIn's 2-second/50% view definition is a weak signal; calibrate this batch on completion and watch time, not view rate.

Pairwise: P over Q. Raise-the-rank for Q, at rung 3 (re-cut) rather than rung 1: naming the role in the text is a minutes-long fix, but Q's stock footage is the structural fault, so the rung-1 rewrite alone would not clear it. Replacing that footage with the branded dashboard converges Q toward P, so it must then pass the real-variation gate against P.

## Example 3 - negative example: the analysis done wrong

The same Example 1 batch, analysed badly:

> "Candidate B is the clear winner - Hook Score: 9.1/10. Its pattern interrupt is the
> strongest scroll-stopper in the batch and should achieve a 40%+ hook rate, well above
> the 30% benchmark. A: 7.4/10. C: 7.1/10. D: 7.0/10. Recommend putting most of the
> test budget behind B."

Why every part of this is wrong:

1. **Fabricated precision.** "9.1/10" and "7.4 vs 7.1" claim resolution no pre-launch judgment has. The evidence behind this skill (near-zero-to-negative hook-rate/ROAS correlation, poor inter-rater reliability) is precisely why decimals are banned - two competent analysts would not reproduce these numbers.
2. **Gate override by cleverness.** B failed sound-off and qualification. A gate-failer cannot rank top no matter how arresting it is; ranking it first funds the documented failure mode where the scroll stops and nothing converts.
3. **Attention treated as performance.** "Should achieve a 40%+ hook rate" predicts an absolute in-market number from a script, then treats that number as success. Hook rate measures attention, not sales - a high one on an unqualified audience is how a 43%-hook-rate ad ends up with mediocre ROAS.
4. **A benchmark imported blind.** "The 30% benchmark" arrives with no denominator, placement, traffic temperature, or source - exactly the kind of directional folklore the calibration step exists to replace with the account's own history.
5. **D scored as a fourth cell.** It shares C's visual; the platform will not treat it as a distinct variation, so its "7.0" describes a test cell that does not exist.

## Common Failure Modes

| Trap                                                                      | Why it burns                                                                                  | Fix                                                                                                        |
| ------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| Emitting a numeric or composite hook score                                | False precision the evidence contradicts; invites cross-batch comparison                      | Bands per dimension, pairwise rank within batch only                                                       |
| Ranking or weighting the seven scoring dimensions                         | Nothing in the evidence orders them; a printed weighting directs budget on invented precision | Leave them unranked by design; rank candidates pairwise, and re-weight only from the account's own winners |
| Picking the most thorough revision instead of the cheapest one that works | A reshoot spends a production cycle to fix what a text rewrite fixes in minutes               | Take the lowest rung of Raising a Rank that removes the failure                                            |
| Ranking a clever gate-failer top "because it will stop everyone"          | Attention without qualification is the documented high-hook-rate/no-ROAS failure              | Gates are absolute; a gate-failer cannot rank top                                                          |
| Scoring candidates against different time windows                         | A 3s feed open vs an 8s in-stream open is not one comparison                                  | Normalise to the placement's window first (workflow step 2)                                                |
| Treating wording-only siblings as separate cells                          | The platform collapses them; the "test" tests nothing                                         | Merge at the variation gate; demand a visual difference                                                    |
| Comparing hook rates across platforms or dashboards                       | Different denominators produce different numbers for the same ad                              | Lock one definition per account; calibrate only within it                                                  |
| Treating the ranking as a performance prediction                          | Hook metrics correlate near zero (or negative) with ROAS                                      | Frame every output as budget shortlisting; the market decides                                              |
| Keeping the generic rubric after calibration fails                        | The account's winners are telling you the weighting is wrong                                  | Re-derive weights from the account's own past winners                                                      |
| Scoring a bumper/non-skippable like a feed hook                           | The view is forced; there is no scroll to stop                                                | Score continuity, branding, qualification only; say so                                                     |
references/hook-patterns.md›
# Hook Patterns - Vocabulary and Anti-Patterns

**How to use this file.** The taxonomies below are a generative and diagnostic vocabulary: use them to name what a candidate opening is doing, to spot that a batch is five flavours of one mechanism, or to suggest what a revision could try.

**Never use them as a scoring axis.** They overlap, they mix executional formats with psychological mechanisms and narrative structures, and none has falsifiable category boundaries - the same ad can be tagged under several types depending on who tags it, which is exactly why human hook-taxonomy tagging has poor inter-rater reliability. A "this batch covers 6 of 8 hook types" claim is not evidence of anything.

## Motion "Hook Tools" - eight psychological types

Source: Motion (motionapp.com, current 2026), framed as "the psychological triggers that top creative strategists use to stop the scroll." Named mechanisms, not measured categories.

1. **Contrarian** - challenges something the viewer believes is true.
2. **Identity Call-Out** - names the viewer's role, situation, or tribe directly.
3. **Confession** - the speaker admits something costly or embarrassing.
4. **Pain Agitation** - opens on the problem, sharpened.
5. **Curiosity Gap** - opens a question with a defined shape that the ad resolves.
6. **Loss Aversion** - frames what the viewer stands to lose by not acting.
7. **Unspoken Truth** - says what the audience thinks but nobody in the category says.
8. **Pattern Interrupt** - an unexpected visual or framing that breaks the feed's texture.

## Motion 2026 empirical top hook/headline categories

Source: Motion 2026 Creative Benchmarks - 578,750 creatives, 6,015 accounts, $1.29B in Meta spend, 1 September 2025 - 1 January 2026. **Caveat Motion itself states**: the window spans BFCM and the holidays, so offer/urgency/newness framings are seasonally inflated and the ranking would differ in another season.

Top-performing categories by hit rate: **Offer only** (highest, ~9.29%), Storytelling, Question, POV, Listicle, How-to, Explainer, Curiosity, **Confession** (~8.74%), Bold claim. Note what the top of the list is: the plain offer, not cleverness - a useful corrective when a batch is all conceptual and nothing states the deal.

## Mark Jung - structural B2B hook types

Source: Mark Jung (Authority), B2B LinkedIn taxonomy via Goldcast. Structural rather than psychological - these describe _which layer of the opening does the work_, useful for B2B where UGC/unboxing executional hooks rarely fit:

- **Composition hook** - the framing/visual arrangement of the shot stops the eye.
- **Physical hook** - a physical action or object in motion.
- **Audio hook** - a sound or line that arrests (remember: fails muted feeds - pair with a subtitle layer).
- **Subtitle hook** - the on-screen text line carries the stop.

Jung's stated rule - "You need to have 80% of your time go to your hook" - is a practitioner's emphasis claim, not a measured allocation.

## Anti-pattern list

Name the pattern when diagnosing; a named pattern ("this is a generic category open") is more actionable than "this is weak."

| Anti-pattern                              | What it looks like                                                                 | Why it fails                                                                                                                                                          |
| ----------------------------------------- | ---------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Logo-first**                            | Logo card, brand jingle, or "we're excited to announce" as the opening beat        | Signals "this is an ad" at the exact moment the viewer decides whether to leave; spends the highest-attention seconds on the advertiser, not the viewer               |
| **Audio-dependent**                       | The promise lives in the voiceover or sound design; first frame + text say nothing | Feed video autoplays muted on Meta and LinkedIn - the meaning never arrives for muted viewers                                                                         |
| **Slow build**                            | Setup, greeting, or context first; the real hook sits in sentence three            | Attention decays to roughly half the audience by second 2 - the "real" hook plays to a fraction of the viewers who saw second 1                                       |
| **Generic category open**                 | "The smarter way to X," "revolutionize your workflow"                              | Interchangeable with any competitor; gives the viewer's brain nothing checkable to grab                                                                               |
| **Hook/body mismatch**                    | Arresting promise, unrelated body and offer                                        | The documented high-hook-rate/mediocre-ROAS failure: the scroll stops, nothing downstream converts, and the metric looks great while the ad loses money               |
| **Cosmetic-only variation**               | Same visual, new first line, shipped as a "new hook"                               | The platform does not treat it as a different variation (Savannah Sanchez) - the test cell is an illusion and the batch's real concept count is lower than it looks   |
| **Clickbait that stops the wrong viewer** | Universal-curiosity spectacle unrelated to the buyer                               | Maximises raw hook rate while attracting unqualified attention - the pattern behind Barry Hott's warning that a higher hook rate does not mean the ad performs better |

Diagnosis procedure:

1. Check every candidate against this table _after_ the gates.
2. Name any pattern found in the scorecard's band rationale.
3. Pair each named pattern with the **cheapest rung of the skill's Raising a Rank ladder that actually removes it** - a text rewrite where the wording carries the fault, a re-cut or a reshoot where the footage itself does.

That revision is the candidate's "raise-the-rank" note.
references/platform-notes.md›
# Platform Notes - Hook Windows, View Counting, Denominators

This file is the platform-differences reference: hook windows for normalising a batch (workflow step 2) and counting rules for the calibration step. Everything here is platform-defined and changes over time - when a number matters, verify it against the platform's current help documentation before relying on it.

**The core warning**: no platform exposes "hook rate" as a native metric. It is always a constructed ratio, and constructions differ. Two dashboards can report different "hold rates" for the same ad because one divides completion-style plays by impressions and the other by 3-second plays.

Never compare a hook rate, view rate, or hold rate across platforms, across dashboards with different constructions, or across a platform's own counting changes. Within one account: lock one definition, write it on every scorecard, calibrate only against it.

## Meta (Facebook/Instagram)

- **Hook window**: feed placements, first ~3 seconds.
- **Autoplay**: muted by default - sound-off legibility is a hard constraint. Meta's creative guidance recommends carrying the opening with text, graphics, and captions.
- **View counting**: 2-second continuous play is the shortest view unit (2+ continuous seconds, ~50% of video in view); 3-second video play is the standard short metric (3+ seconds, or ~97% of a shorter video; replays excluded); ThruPlay = completion or at least 15 seconds.
- **Hook-rate construction**: `3-second video plays ÷ impressions` - the dominant practitioner definition. Because Meta autoplays in feed, video plays ≈ impressions, so this reads roughly as "share of impressions retained past 3 seconds."
- **Hold-rate constructions in circulation**: `ThruPlays ÷ impressions` _and_ `ThruPlays ÷ 3-second plays` - both are in active use; this is the two-dashboards trap. State which one the account uses.

## TikTok

- **Hook window**: first ~3 seconds - TikTok's own Creative Codes emphasise hooking within the first 3 seconds (platform-echoed heuristic, partially evidence-backed via TikTok Marketing Science / Kantar and Ipsos studies).
- **View counting**: 2-second video view (played at least 2 seconds, replays excluded) is the shortest unit; 6-second view (or an engagement in the first 6 seconds) is the standard short metric; quartile completions at 25/50/75/100%.
- **Hook-rate construction**: `2-second views ÷ impressions` (some practitioners call this "thumb-stop rate"). Not comparable with Meta's 3-second construction - a "better hook rate on TikTok than Meta" claim is comparing different metrics.
- Sound-on culture is stronger than on Meta/LinkedIn, but muted and captioned viewing is common enough that sound-off legibility still gates.

## YouTube / Google Ads

Windows and counting differ _per format_ - normalise each candidate to the format actually booked:

- **Skippable in-stream**: the skip button appears at 5 seconds, so the hook window is the first ~5 seconds. A paid "view" counts at 30 seconds, completion, or interaction - whichever comes first - so view rate here measures something very different from a feed hook rate.
- **Bumper (6s) and non-skippable**: forced views - no view is counted in the skippable sense, and there is no scroll to stop. "Stop the scroll" logic does not apply at all; judge continuity, branding, and qualification instead.
- **In-feed**: a view = a click through to the watch page - closer to a thumbnail/headline test than a video hook test.
- **Counting-change dates that break comparisons**:
  - On **31 March 2025**, YouTube switched Shorts to play-based counting (any start or replay counts; the older stricter metric was renamed _engaged views_).
  - On **24 August 2026**, play-based counting extended to long-form, live, and podcasts.

  Public view counts before and after each date are different metrics - never trend or calibrate across them. Monetization still runs on engaged views.

## LinkedIn

- **Hook window**: feed, first ~3 seconds - same feed logic as Meta.
- **Autoplay**: muted by default - captions and on-screen text are effectively mandatory.
- **View counting**: a view = 2+ continuous seconds with at least 50% of the player on screen (MRC-aligned). This is a weak "pause" signal, not attention - treat LinkedIn view counts as inflated relative to real attention, and prefer completion and watch-time metrics for calibration. Practitioner-reported median watch time sits near ~6 seconds (ZenABM's figure - directional, not official).
- **Counting-change date**: in **October 2022** LinkedIn changed video completion rate from `completions ÷ views` to `completions ÷ plays`. Completion-rate data spanning that date is not comparable.
- Mostly B2B inventory: the 95:5 out-of-market logic applies, so calibrating a LinkedIn batch on short-view rates alone is doubly weak - the metric is soft _and_ the objective is usually memory, not the click.

## Folklore ledger - famous numbers, real origins

If any of these appear (in the user's brief, an agency deck, a brief you are reviewing), correct the record in the same breath or leave them out entirely:

- **"You have 0.25 seconds to capture attention"** - a 2016 Fors Marsh finding that people can _recall_ feed content at statistically significant rates after 0.25 seconds of exposure. A recall-at-exposure statistic, not an attention deadline.
- **"1.7 seconds per piece of content"** - Facebook's own 2016 average for time spent with mobile content. A decade-old average, not a rule; researchers (Field, Nelson-Field) argue it is below the threshold for reliable brand effect anyway.
- **"85% of video is watched without sound"** - a 2016 Digiday report of a few publishers' self-reported numbers; Facebook never confirmed it. Design for sound-off because of muted autoplay, not because of this number.
- **"The average attention span is 8 seconds"** - debunked; no credible primary source. Do not use.

Prefer the well-sourced material instead:

- Nelson-Field/Amplified attention-decay and memory-threshold data.
- Motion's 2026 disclosed-methodology benchmarks.
- Google's ABCD framework: the only opening framework with disclosed large-sample validation, Kantar/Ipsos-measured, associated with up to a 30% lift in short-term sales likelihood.
- The 95:5 rule for B2B.
references/scoring-rubric.md›
# Scoring Rubric - Seven Dimensions

Band every gate-passing candidate **strong / adequate / weak** on each dimension. **Never invent numeric sub-scores, percentages, or a composite total** - the bands exist precisely because a number would claim precision the evidence does not support, and because human hook-scoring has poor inter-rater reliability even among professionals. When two bands both seem defensible from the input available, band `adequate`, and flag the missing input (e.g. "storyboard doesn't show the first frame - a frame grab would resolve time-to-signal").

Bands feed the pairwise ranking as evidence, not arithmetic. Never rank by counting strongs.

**The seven dimensions are unranked on purpose** - no weights, no importance order, no fix-first
sequence, and nothing below should be read as one. With hook rate correlating **-0.19 with
ROAS**, no evidence supports claiming that any one dimension returns more performance per unit
of effort than another, so an ordering here would be invented precision dressed as guidance. The
ranking this skill does emit is between candidate openings, never between the axes used to judge
them; the ordered ladder in the skill's Raising a Rank section ranks _revisions_, whose effort
differs by orders of magnitude and is observable before launch.

## 1. Time-to-signal

**Measures**: how many seconds pass before the viewer knows what is in it for them - not when something _happens_, but when the personal relevance lands.

**Evidence**: feed attention decays steeply - roughly 80% of the audience still present in second 1, 50% by second 2, 20% by second 3 (Karen Nelson-Field / Amplified Intelligence, human-attention datasets across Facebook, YouTube, Instagram, TV). Memory reliably encodes only after roughly 2.5 seconds of active (eyes-on) attention; Amplified/VCCP (May 2025, 20,000+ views of 72 digital video ads) found 1.5 seconds can suffice when distinctive brand assets are deployed well. The practical read: the value signal must land while the majority is still there - by around second 2.

- **Strong**: the "what's in it for me" is legible in the first frame or first line - a viewer who saw only second 1-2 could state the promise.
- **Adequate**: the signal lands within the placement's hook window, but seconds 1-2 are spent on setup that earns nothing on its own.
- **Weak**: throat-clearing - greeting, logo card, scene-setting, or a tease with no defined shape - and the actual signal arrives after the window, when most of the audience has gone.

**Ambiguous case**: scripts often under-specify the first visual frame. If the verbal line signals fast but the visual is unknown, band `adequate` and request the first-frame plan.

## 2. Sound-off legibility

**Measures**: how well the opening delivers its meaning with audio muted - beyond the pass/fail gate, which only asks whether it survives at all.

**Evidence**: feed video autoplays muted on Meta and LinkedIn (platform behaviour, verified); Meta's own creative guidance recommends carrying the opening with text, graphics, and captions. Note on folklore: the famous "85% of Facebook video is watched without sound" is a 2016 publisher-reported snapshot (Digiday, citing LittleThings and Mic) that Facebook never confirmed - do not cite it as a current or platform-verified figure. The design constraint stands on muted autoplay alone.

- **Strong**: first frame plus on-screen text carry the full promise; audio only adds texture. A muted viewer loses nothing essential.
- **Adequate**: the gist survives muting, but a meaningful layer (the specific claim, the punchline) lives only in the audio.
- **Weak** (usually a gate FAIL instead): meaning depends on voiceover or sound design; muted, the opening is wallpaper.

**Ambiguous case**: transcripts contain no visual/text information at all. Band `adequate` at best and say the input cannot support `strong`.

## 3. Audience qualification / self-selection

**Measures**: how precisely the opening lets the _right_ viewer recognise "this is for me" - and, equally, lets the wrong viewer scroll past. The goal is not maximum attention; it is qualified attention.

**Evidence**: Barry Hott (growth/creative consultant, ~$1B managed spend): "Higher CTR or hook rate (thumbstop) doesn't mean an ad will perform better or worse... Clickbait can get tons of clicks without generating any sales." The same logic drives the documented high-hook-rate/mediocre-ROAS failure accounts.

- **Strong**: the opening names or unmistakably implies the buyer's role, situation, or problem, and the disqualified viewer (per the Interview) has no reason to stay.
- **Adequate**: relevant to the buyer but broad - the right viewer might self-select, plenty of wrong viewers will too.
- **Weak**: pure spectacle or universal curiosity - it stops everyone, which is qualification for no one - or so narrow/inside that even the buyer doesn't recognise themselves.

**Ambiguous case**: when the Interview's "what disqualifies a viewer" answer is thin, band `adequate` and push the question again - this dimension cannot be banded without it.

## 4. Specificity

**Measures**: whether the opening makes a concrete, checkable claim or shows a concrete situation, versus generic category language.

**Evidence**: convergent practitioner consensus rather than a single controlled study - specific numbers, named situations, and checkable claims recur across practitioner hook analyses as separating working openings from generic ones, and generic category openings recur in every anti-pattern list. Treat as well-supported practice, not measured law.

- **Strong**: a claim you could check or a situation you could film - a number, a named context, a visible demonstration ("this skillet went through 1,000 dishwasher cycles").
- **Adequate**: a real benefit stated abstractly - true, relevant, but interchangeable with a competitor saying it.
- **Weak**: category wallpaper - "revolutionize your workflow," "the smarter way to X" - language any brand in the category could run unchanged.

**Ambiguous case**: a bold specific claim that the product cannot back is not `strong` specificity - it is a promise-payoff gate problem. Check the gate before banding.

## 5. Brand legibility timing

**Measures**: whether the choice of when and how the brand becomes legible in the opening fits the stated objective and platform. This is a genuine, unresolved tension - band the _fit of the choice_, not adherence to a rule.

**Evidence, deliberately in tension**: opening on a logo card signals "this is an ad" at the worst moment (documented anti-pattern); yet withholding branding costs brand linkage - TikTok reports a 17-point drop in brand linkage when branding is left to the end (TikTok-cited figure, asserted rather than a controlled finding). Amplified/VCCP (2025) found distinctive brand assets, used well, let memory encode in as little as 1.5 seconds - which argues for early _distinctive-asset_ branding without a logo card. The evidence does not settle the question; objective and platform decide it.

- **Strong**: a deliberate choice matching the objective - direct-response cold traffic: brand cues woven in (product in hand, distinctive colour/asset) without an ad-announcing logo open; brand/memory objective (most B2B): distinctive assets legible early, per the linkage evidence.
- **Adequate**: branding present and findable, but by default rather than by design - neither hurting nor working.
- **Weak**: the accidental extremes - a logo card or brand announcement as the opening beat with a DR objective, or a brand-building objective where nothing identifies the brand until the end card.

**Ambiguous case**: if the user has not stated an objective, this dimension cannot be banded honestly - ask, don't assume DR.

## 6. Promise-payoff continuity

**Measures**: beyond the pass/fail gate - how tightly the opening's promise _is_ the ad's payoff and the offer's substance, versus merely not contradicting them.

**Evidence**: the documented failure this guards against - "I've run campaigns with 43% hook rates and mediocre ROAS. The hook stopped the scroll... but it didn't qualify the right viewer or set up the offer" (named practitioner account, Imagine.art). The same source describes teams lifting hook rate from 18% to 37% across 40 variations with no ROAS change.

- **Strong**: the opening is the offer in miniature - the promise made in the window is exactly what the ad demonstrates and the offer sells; a viewer who converts got what second 1 promised.
- **Adequate**: promise and payoff are related but the opening oversells or angles away - the ad delivers a cousin of the promise.
- **Weak** (usually a gate FAIL): bait - the promise exists to stop the scroll and the body pivots to something else.

**Ambiguous case**: when only the opening was supplied, the gate and this band cannot be judged - require at least a summary of the full ad and the offer before scoring the batch.

## 7. Placement-native fit

**Measures**: whether the opening is built for the stated placement's hook window and viewing context, or transplanted from another format.

**Evidence**: the windows differ materially (platform-defined, verified):

- Feed placements: ≈ first 3 seconds.
- YouTube skippable in-stream: ≈ first 5 seconds, before the skip button.
- Bumpers and non-skippable formats: forced views, so scroll-stopping logic does not apply at all. Only continuity, branding, and qualification carry meaning.

Sound norms, text-safe zones, and pacing also differ by placement.

- **Strong**: window, first-frame composition, text placement, and pacing all built for the stated placement.
- **Adequate**: transferable with light edits - right length, but text sits in another platform's safe zone or pacing follows another format's rhythm.
- **Weak**: a transplant - a 15-30 second cold-open structure dropped into a 3-second feed window, a watermarked cross-post, or a feed-style scroll-stopper submitted for a forced-view slot.

**Ambiguous case**: if the user plans multiple placements for one batch, score against the primary placement and note that the ranking holds only for it - a re-rank per placement is a separate pass.
SKILL.md›
---
name: ad-hook-analyzer
description: "Score and force-rank the openings of candidate video ads - from a script, transcript, storyboard, shot list, or a description of a finished cut - to decide which hooks deserve test budget, returning a ranked shortlist within the batch rather than an absolute performance prediction. Use whenever the user mentions a hook, the first 3 seconds, an ad opening, hook rate, thumbstop, attention scoring, or asks which hook to test or whether a hook works - even if they never say 'hook analyzer'. Covers B2B and B2C on any platform. Ends at the ranking: writing the scripts themselves is mbfinotti/advertising-skills@ugc-ad-scripts and static copy is mbfinotti/advertising-skills@ad-copy-variants."
license: MIT
metadata:
  author: Maya-Beth Finotti
  version: "1.3.2"
---

# Hook Analyzer

You are a performance-creative analyst. Read a batch of candidate video-ad openings and decide
which ones are worth spending test budget on. The output is a relative ranking within one batch

- a shortlisting device.

The market is the judge; this skill only decides which candidates get to face it.

The method is built on honesty about what a pre-launch hook score can and cannot do:

- "Hook rate" is a practitioner-invented derived ratio, not a native platform metric, and its
  denominator differs across dashboards.
- One agency's analysis of the same 11 brands (drawn from 3,859 ads) found hook rate correlated
  **-0.19 with ROAS** and hold rate -0.10 - slightly the _wrong_ direction (Sweat Pants Agency;
  single-agency data, not peer-reviewed, but directionally matched by independent practitioner
  accounts).
- Roughly 5% of creatives become winners at all (Motion 2026, 578,750 creatives).
- Human hook-taxonomy tagging has poor inter-rater reliability, because the taxonomies overlap
  and lack falsifiable definitions.

So this skill never produces an absolute predicted-performance number. It produces:

- a **relative ranking of candidate openings within one batch** - never a score comparable
  across batches, accounts, or platforms;
- **anchored qualitative bands** (strong / adequate / weak) per dimension, with concrete anchors
  - never decimals, percentages, or a composite score. A "7.4/10 hook score" is false precision
    the evidence cannot support;
- a decision about **what to spend test budget on**, explicitly framed as such.

**Scope ends at the ranking.** Editing, pacing, or structure beyond the opening is out of scope.

- Writing new scripts or hooks: `mbfinotti/advertising-skills@ugc-ad-scripts`.
- Diagnosing an already-running ad's decline: `mbfinotti/advertising-skills@ad-creative-fatigue`.
- Sizing and designing the test itself: `mbfinotti/advertising-skills@ad-creative-test-plan`.
- Ad copy variants: `mbfinotti/advertising-skills@ad-copy-variants`.
- Briefs: `mbfinotti/advertising-skills@ad-creative-brief`.
- Collecting competitor hooks: `mbfinotti/advertising-skills@ad-swipe-file`.
- Format/placement fit: `mbfinotti/advertising-skills@ad-format-fit`.
- Post-click problems: `mbfinotti/advertising-skills@paid-landing-page-audit`.
- Account-level root causes: `mbfinotti/advertising-skills@ad-account-diagnostic`.

## Interview

Ask before scoring anything.

- One question per message.
- Offer multiple-choice answers where possible.
- Skip anything already answered or visible in the supplied material.

- Which platform, and which exact placement? (Feed / short-form vertical / skippable in-stream /
  bumper or non-skippable / other. The hook window depends on this.)
- B2B or B2C?
- Cold prospecting or warm/retargeting?
- What input exists per candidate: script, transcript, storyboard, shot list, or a description
  of a finished cut?
- How many candidate openings are in the batch?
- What is the offer, and what does the rest of the ad actually deliver? (Needed for the
  promise-payoff gate.)
- Who is the buyer, and - just as important - what disqualifies a viewer? (Needed for the
  qualification gate.)
- What openings are already running or previously tested? (Needed to judge real variation.)
- What production capacity exists for revised openings, how fast, and by what date must this
  batch launch? (Decides which rungs of Raising a Rank are available at all: an editor only, or
  a shoot slot. A locked date deletes the expensive rungs before any ranking starts.)
- How will this batch actually be tested and measured after launch?
- Which hook-rate definition does your dashboard use - what exactly is the numerator and
  denominator? (Two dashboards can report different numbers for the same ad; the calibration
  step depends on this answer.)

## Workflow

1. Run the Interview; collect every answer before scoring.
2. **Normalise the batch to a comparable unit.** Every candidate is scored on its first N
   seconds, with N set by the stated placement:
   - Feed and short-form vertical: first ≈3 seconds.
   - Skippable in-stream: first ≈5 seconds, up to the skip button.
   - Bumper and non-skippable: forced views, so "stop the scroll" logic does not apply at all.
     Score only continuity, branding, and audience qualification, and say so.

   Details in [references/platform-notes.md](references/platform-notes.md). Never compare a
   candidate's first 3 seconds against a sibling's first 8.

3. **Run the four Hard Gates** (below) on every candidate. Record pass/fail per gate with a
   one-line reason.
4. **Band each gate-passing candidate** on the seven scoring dimensions, using the anchors in
   [references/scoring-rubric.md](references/scoring-rubric.md). Bands only - no numbers. When a
   dimension is genuinely ambiguous from the input available, band it `adequate` and flag what
   input would resolve it.
5. **Force-rank within the batch by pairwise comparison.** Compare candidates head to head - A
   vs B, winner vs C - asking for each pair "which of these two would I fund first, and why,"
   grounded in the gate results and bands. Paired same-judge comparison is materially more
   reliable than comparing absolute scores, because the judge's bias applies equally to both
   sides and cancels out. Never derive the ranking by tallying bands into a number.
6. **Emit the scorecard**: gates, bands, ranked shortlist, and one concrete "what would raise
   this rank" note per candidate - each note taken from the cheapest rung of Raising a Rank
   that actually removes the failure. See The Scorecard below, and the shape and worked
   versions in [references/examples.md](references/examples.md).
7. **Set the calibration check**: record the predicted rank order and the dashboard's hook-rate
   definition from the Interview, to be compared against measured results after launch (see
   Measuring below).
8. If your harness has persistent memory, memorize: the account's hook-rate denominator, each
   batch scored, predicted-vs-measured rank agreement per batch, and which dimensions have
   proven predictive for this account.

## Hard Gates

An opening that fails any gate cannot be ranked top of the batch, whatever else it does well.
Gates are pass/fail; they are not bands.

1. **Sound-off legibility.** The opening must land with audio muted. Feed video autoplays muted
   on Meta and LinkedIn, so an opening whose meaning lives in the voiceover or sound design
   never delivers its meaning to a large share of viewers. If the promise is not legible from
   the visuals and on-screen text alone, the gate fails.
2. **Promise-payoff continuity.** What the opening promises must be what the rest of the ad and
   the offer actually deliver. This is the documented "43% hook rate, mediocre ROAS" failure:
   the hook stopped the scroll but didn't set up the offer, and nothing downstream converted
   (named practitioner account, Imagine.art). A mismatch fails the gate no matter how arresting
   the opening is.
3. **Audience qualification.** The opening must stop the _right_ viewer, not everyone. Barry
   Hott (~$1B managed spend): "Higher CTR or hook rate (thumbstop) doesn't mean an ad will
   perform better or worse... Clickbait can get tons of clicks without generating any sales." An
   opening that maximises raw attention while attracting viewers who can't or won't buy fails
   the gate.
4. **Real variation.** An opening that differs only in wording from a sibling in the batch (or
   from an ad already running) is not a distinct test cell. Savannah Sanchez's operating rule:
   the opening _visual_ has to differ too, or the platform "doesn't really see it as a different
   variation." Wording-only siblings get merged into one cell and noted, not ranked separately.

## Scoring Dimensions

Band each gate-passing candidate strong / adequate / weak on all seven. Full definitions,
anchors, evidence, and ambiguous-case rules live in
[references/scoring-rubric.md](references/scoring-rubric.md) - read it before banding.

**These seven are deliberately unranked, and the refusal is the finding - not an omission to
fix.** They carry no importance order, no weights, and no "fix this one first" sequence. The
attention metrics they describe correlate **-0.19 with ROAS** (hold rate -0.10; Sweat Pants
Agency, 11 brands from 3,859 ads), so nothing in the evidence says one dimension returns more
performance per unit of effort than the next - an ordering printed here would be invented, and
would then steer real production hours and real test budget.

Rank candidates, not dimensions. The pairwise force-rank in workflow step 5 stays: it compares
two concrete openings under one judge, so the bias applies to both sides and cancels. The only
defensible dimension weighting is the account's own, re-derived from its past winners once
calibration fails (see Measuring) - never a default one written into this file.

1. **Time-to-signal** - how many seconds pass before the viewer knows what's in it for them.
   Feed attention decays steeply: roughly 80% of the audience still present in second 1, 50% by
   second 2, 20% by second 3 (Nelson-Field / Amplified Intelligence).
2. **Sound-off legibility** - beyond the pass/fail gate: how _well_ the opening works muted, not
   just whether it survives.
3. **Audience qualification / self-selection** - how precisely the opening lets the right viewer
   recognise "this is for me" and the wrong viewer scroll on.
4. **Specificity** - a concrete, checkable claim or situation versus generic category language.
5. **Brand legibility timing** - a genuine, unresolved tension, not a rule: opening on a logo
   signals "this is an ad," yet withholding branding costs brand linkage (TikTok reports a
   17-point drop when branding is left to the end). Band by fit to the stated objective and
   platform; the evidence does not settle this either way.
6. **Promise-payoff continuity** - beyond the gate: how tightly the opening's promise is the
   ad's actual payoff, versus merely not contradicting it.
7. **Placement-native fit** - the hook window and interface differ by placement; an opening
   built for the stated placement bands strong, a transplant from another format bands weak.

## The Scorecard

Deliver one block per batch: gates, bands, pairwise ranking, one raise-the-rank note per
candidate, the shipping check, and the calibration line. Read
[references/examples.md](references/examples.md) before writing the first one - it opens
with the field-by-field shape, then three worked versions: a full B2C batch, a condensed
B2B contrast, and a negative example of the analysis done wrong.

## Raising a Rank

Every candidate below the top of the shortlist gets one revision note. Which revision costs the
batch its production hours, so pick by rank movement per hour spent - not by which fix is the
most thorough. Ordering the fixes is legitimate where ordering the scoring dimensions is not:
these rungs differ in effort by orders of magnitude, and that difference is observable before
launch.

- efficiency (rank movement per hour): `rewrite the on-screen text == fix the text/caption
layer > re-cut from footage already shot > shoot a new opening beat > rebuild the concept`
- effort: `rebuild (a full production cycle) > new opening beat (a shoot slot, plus talent and
production coordination) > re-cut (an hour in the edit, no shoot) > text rewrite == caption
fix (minutes, one person, reversible)`
- value (gate failures the fix can actually remove): `rebuild > new opening beat > re-cut >
text rewrite == caption fix` - the cheap rungs only fix what the footage already carries,
  which is why the most valuable fix is also the least efficient
- compliance cost: `new opening beat == rebuild > text rewrite > re-cut == caption fix` - a
  reshoot reopens creator usage rights, and a rewrite that adds a checkable claim sends the ad
  into claim substantiation and platform ad review; both take longer to reverse than the edit

Each tie above is a real equality, not a shrug:

- **Text rewrite == caption fix (efficiency).** Both are minutes of one person's work in the
  same tool on the same finished cut, both reversible, and both capped at what the existing
  footage already carries - they differ in which failure they remove, never in what they cost
  or how far they can move a candidate.
- **New opening beat == rebuild (compliance cost).** The exposure comes from putting a camera
  on new material, which both do: creator usage rights reopen and the ad re-enters review
  either way.
- **Re-cut == caption fix (compliance cost).** Both only rearrange material already cleared -
  no new claim, no new footage, nothing for a reviewer to look at again.

1. **Rewrite the on-screen text.** Moves time-to-signal, specificity, and qualification on
   footage you already have.
2. **Fix the text/caption layer.** Add or re-time captions, or move text out of another
   platform's safe zone - buys sound-off legibility and placement-native fit.
3. **Re-cut from footage already shot.** Open on a different existing shot - buys time-to-signal
   and brand timing, and can supply the visual difference a wording-only sibling lacks.
4. **Shoot a new opening beat** against the same body, when no existing frame can carry the
   promise muted.
5. **Rebuild the concept.** Not a fix - a new candidate that re-enters all four gates.

**Default: rungs 1-2.** Move up one rung when the failure is structural rather than verbal - the
footage cannot deliver the promise muted, or the visual is identical to a sibling's. Go straight
to rung 5 when the batch is below the 3-opening floor: polishing two openings never produces the
third concept the shipping gate demands.

This order is a default, not a law - re-rank it against what the Interview revealed about this
account. A standing shoot cadence, an in-house editor, or owned footage the brand can re-cut
collapses the effort of rungs 3-4.

**Delete, don't demote.** A launch date inside the shoot lead time, or a creator contract that
would have to be renegotiated first, does not push rungs 4-5 down the order. It removes them
from this batch's menu, and the scorecard names them deleted with the constraint that deleted
them.

Every revision note then comes from rungs 1-3 only. When rung 5 is deleted and the batch sits
below the 3-opening floor, report that the batch cannot clear the shipping gate; a note
recommending a rebuild nobody can run reads as a plan and is not one.

## B2B vs B2C

Four dimensions are identical for both and need no adjustment: time-to-signal, sound-off
legibility, promise-payoff continuity, and the real-variation gate. Attention decay and muted
autoplay do not care what is being sold.

Where they diverge:

- **B2B:** under the 95:5 rule (LinkedIn B2B Institute with Prof. John Dawes, Ehrenberg-Bass,
  2021), roughly 95% of B2B buyers are out-of-market at any moment, so a B2B opening mostly
  reaches future buyers. It qualifies by **role and situation** ("if you run payroll across
  three countries...") and aims at **memory, not an immediate click** - hard direct-response
  hook heuristics transfer poorly, and brand-legibility timing tilts earlier.
- **B2C:** DTC/B2C openings lean the other way - executional patterns, immediate offer framing,
  self-selection by problem or desire, and qualification judged against purchase intent now.

On LinkedIn specifically, autoplay is muted and a "view" is a weak 2-second/50%-on-screen
signal, so treat platform view counts as inflated relative to attention.

State the scoring mode on the scorecard; a batch scored in the wrong mode ranks the
wrong winner.

## Measuring Whether This Worked

The outcome measure is **rank agreement, not absolute accuracy**. After the batch runs, compare
the predicted rank order against the measured hook-rate rank order for the same openings on the
same platform and placement.

- **Validate the denominator first, every time.** Hook rate is `3-second video plays ÷
impressions` on Meta and `2-second views ÷ impressions` on TikTok; other dashboards construct
  it differently. See [references/platform-notes.md](references/platform-notes.md) for the full
  per-platform list.
- **Never compare across platforms or across a counting-change date.** A hook rate from one
  platform is not comparable to another's, and YouTube's view-counting changes of 31 March 2025
  and 24 August 2026 mean the numbers before and after are not the same metric either.
- **Pass threshold**: over a rolling window of 5 scored batches, the top-ranked opening lands in
  the top half of the measured ranking at least 3 times. That is a deliberately modest bar - it
  is what a useful shortlisting device clears and a coin flip doesn't.
- **Below threshold**: the generic rubric is miscalibrated for this account. Stop trusting it:
  re-derive the dimension weighting from the account's own past winners - which dimensions did
  the openings that actually won share? - and record the re-weighting in memory if available.
- **Shipping gate**: never ship a batch with fewer than 3 openings that clear all four gates
  _and_ differ on the concept axis. With roughly 5% of creatives winning (Motion 2026), volume
  is what produces winners; a one-opening "batch" is a bet, not a test.

## Reference

- Read [references/scoring-rubric.md](references/scoring-rubric.md) before banding - full
  dimension definitions, strong/adequate/weak anchors, supporting evidence, ambiguous-case
  rules.
- Read [references/hook-patterns.md](references/hook-patterns.md) when naming what a candidate
  is doing or diagnosing why it fails - sourced hook taxonomies (as vocabulary, never as a
  scoring axis) and the anti-pattern list.
- Read [references/platform-notes.md](references/platform-notes.md) when setting the hook
  window, validating a hook-rate definition, or calibrating - per-platform windows,
  view-counting rules, denominators, and counting-change dates.
- Read [references/examples.md](references/examples.md) when writing the scorecard - the
  field-by-field shape, a worked B2C batch, a worked B2B contrast, a negative example of the
  analysis done wrong, and the common-failure-modes table.
- See `mbfinotti/advertising-skills@ad-creative-test-plan` to design the test the shortlist feeds.
- See `mbfinotti/advertising-skills@ad-copy-variants` for static copy (this skill handles video openings only).
- See `mbfinotti/advertising-skills@paid-landing-page-audit` for post-click landing page issues beyond the ad itself.