返回 Skills 目錄
mbfinotti/sales-skills已通過檢查

SKILL DETAIL

cold-email-subject-line-tester

mbfinotti/sales-skills/cold-email-subject-line-tester

Generates, scores, and split-tests email subject lines, producing angle-distinct variants, a numeric scorecard, and a test plan with sample size and decision metric. Covers B2B cold outbound and B2C lifecycle/promotional email, and flags spam-trigger words inside the subject line only. Use whenever the user mentions a subject line, open rates, email A/B test, preview text, or "nobody opens my emails", even if they don't say "subject line" explicitly. Do NOT use for body copy or inbox-placement audits (mbfinotti/sales-skills@cold-email-deliverability), personalization angles (mbfinotti/sales-skills@sales-outreach-personalization), or cadence design (mbfinotti/sales-skills@sales-outbound-sequence).

安裝量 · 189查看來源

Installation

npx skills add https://github.com/mbfinotti/sales-skills --skill cold-email-subject-line-tester

技能檔案

SKILL.md

最近同步 · 2026年9月15日

evals/evals.json
{
  "skill_name": "cold-email-subject-line-tester",
  "evals": [
    {
      "id": 1,
      "prompt": "I run outbound at Harborlight Analytics - we sell a data-quality monitoring tool to Heads of Data Engineering. Best case study we have says we cut pipeline incident triage time by 63%. I need 4 subject lines for a cold sequence going to 1,800 prospects starting next Monday. Our old ones felt too generic so I want the prospect's first name in there and the 63% stat, something like \"Priya, 63% faster triage\". Our sequencer can split test. The body is a two-line intro plus the triage benchmark chart. Give me the four and tell me which one to pick.",
      "expected_output": "Four angle-distinct B2B cold subject lines, each 1-4 words and inside the truncation budget, with no first name and no percentage figure; a scorecard at the 8/10 gate; and a pre-committed split test decided on replies rather than opens.",
      "files": [],
      "expectations": [
        "Refuses to place the prospect's first name in the B2B cold subject line and classifies it as a hard fail, not a stylistic preference",
        "States that a first name in a B2B cold subject signals a mail merge and correlates with roughly 12% fewer replies",
        "Keeps the 63% figure out of every shipped subject line and cites numbers or percentages measuring around -46% opens in B2B cold subjects",
        "Every shipped subject line is at most 4 words",
        "Every shipped subject line fits within roughly 35 characters",
        "Assigns each variant a named angle drawn from the catalog: internal-style noun, specific pain question, executive-world anchor, competitor or market reference, company initiative, or trigger event",
        "No two shipped variants carry the same angle, and none is a synonym swap of another",
        "Scores every variant on a 10-point scorecard covering length and truncation fit, angle execution, reader's-world framing, body-promise match, set distinctness, and human sound",
        "States a pass threshold of at least 8/10 with zero hard fails, and regenerates any variant below it instead of shipping it",
        "Declines to pick a single winner by judgement and instead specifies a split test with sample size pre-committed before launch",
        "States a floor of 200 or more sends per variant and notes that a 4-variant test multiplies the required total by roughly 2x",
        "Names reply rate, positive reply rate, or meetings booked as the decision metric and rules out calling the winner on opens alone"
      ]
    },
    {
      "id": 2,
      "prompt": "Open rate on our cold sequence fell from 38% to 11% over about three weeks. We migrated to a new sending domain last month. Two ideas I want to try: \"Re: quick follow up\", since apparently that lifts opens a lot, and \"Free pipeline audit - limited time\". Can you write those up properly, and also check whether our SPF record is the problem? We sell revenue forecasting software to CFOs at 200-800 person companies.",
      "expected_output": "A refusal of both proposed subject lines on hard-fail grounds, replacement angle-distinct variants inside the B2B cold budget, and the domain-migration and inbox-placement diagnosis handed to the deliverability skill rather than attempted here.",
      "files": [],
      "expectations": [
        "Refuses the fake \"Re:\" prefix and classifies it as a hard fail rather than a tactic with tradeoffs",
        "States that a deceptive subject line is prohibited under CAN-SPAM, not merely bad practice",
        "Flags \"Free\" and \"limited time\" as spam-trigger vocabulary that disqualifies the variant outright",
        "Does not audit SPF, DKIM, DMARC, domain reputation, or inbox placement",
        "Hands the domain-migration and inbox-placement review to a separate deliverability skill instead of answering it",
        "States that an open-rate collapse immediately following a sending-domain change points at the sending setup rather than at the subject line",
        "Does not write or rewrite the email body copy",
        "Provides replacement subject lines carrying no spam-trigger vocabulary and no framing implying a prior conversation",
        "Every replacement subject line is at most 4 words and fits within roughly 35 characters",
        "Each replacement variant expresses a distinct named angle rather than a reworded version of the same idea",
        "Notes that open rate in B2B cold is directional only and cannot by itself confirm the subject line caused the drop"
      ]
    },
    {
      "id": 3,
      "prompt": "Lumen & Thread here, DTC bedding. Our cart-abandonment email currently uses \"You left something behind\" with a smiley emoji - 31% opens but conversions are terrible, around 0.4%. List is about 40,000 and roughly 2,800 people abandon a cart each month. Our marketing lead is adamant that emoji goes in everything. Want 4 new options I can rotate in. The email contains the saved cart, a size-guide link, and a 10% code that expires in 48 hours.",
      "expected_output": "Four B2C lifecycle subject-plus-preheader pairs sized to the B2C budget, with emoji parked as a test variable rather than mandated or banned, opens paired with a downstream conversion read, and guardrail thresholds stated.",
      "files": [],
      "expectations": [
        "Identifies the campaign as B2C lifecycle or promotional and applies that regime's rules rather than B2B cold rules",
        "Ships a preheader alongside every subject line variant",
        "Each preheader runs roughly 90-140 characters",
        "Each preheader extends the subject rather than repeating it",
        "Treats subject plus preheader as the single tested unit",
        "Keeps each subject within roughly 50 characters, inside the 40-60 character band",
        "Does not ban numerals in the subject and states that numbers perform acceptably in B2C lifecycle unlike B2B cold",
        "Refuses to treat emoji as either a default or a blanket ban, and places it in the ranked parked-question test queue instead",
        "Pairs open rate with a downstream metric such as click or conversion rather than deciding on opens alone",
        "States guardrail thresholds of unsubscribe under 0.5% and spam complaints under 0.1%",
        "Notes that B2C should test weekend sends, unlike B2B cold which avoids weekends",
        "States that roughly 20% of the list is split across test variants with the winner sent to the remainder"
      ]
    },
    {
      "id": 4,
      "prompt": "Small agency, three of us. We have exactly 60 named accounts, all VP Engineering at Series B fintechs, and we know them well. I want to A/B test 5 subject line versions across them. Also I read that lowercase subjects beat title case, so I want to settle that at the same time. We sell a compliance-automation platform at roughly $40k ACV, and missing one of these accounts genuinely hurts.",
      "expected_output": "A plain statement that 60 accounts power no test at all, defaults shipped instead of a queue, casing refused as a settled question, and the trigger-event angle promoted on the strength of the short named-account list.",
      "files": [],
      "expectations": [
        "States plainly that 60 accounts cannot power even a two-arm test at the 200-sends-per-variant floor",
        "Declines to design the 5-way test and does not offer a shrunken substitute test as though it would decide anything",
        "Ships the defaults instead of a test queue: casing held constant, statements preferred over questions, emoji off",
        "Refuses to confirm that lowercase beats title case",
        "Describes casing as two directly contradictory vendor claims with no independent arbiter, one side teaching title case and the other lowercase",
        "Promotes the trigger-event angle despite its last place on the default efficiency order, because a 60-account named list is short enough to source and one missed reply costs more than the sourcing",
        "Names which specific fact about this user moved which angle in the ranking",
        "Moves the executive-world anchor ahead of the specific pain question for this VP-level list",
        "Still scores every variant on the 10-point scorecard and applies the 8/10 pass threshold even though no test will run",
        "Caps the shipped set below the requested 5 variants rather than producing 5 to match the ask",
        "States the ranking explicitly rather than leaving it implied by the order of the rows"
      ]
    },
    {
      "id": 5,
      "prompt": "I'm the only SDR at Coreline Freight. 120 emails a day is the quota, I get maybe 30 seconds a prospect, and we have no intent data and no enrichment tooling - just a list of VP of Logistics contacts at mid-market 3PLs. Need 5 subject lines I can reuse across the whole list. Rank them best to worst for me. The email offers a one-page benchmark on dock-to-stock time.",
      "expected_output": "A candidate set with the research-heavy angles deleted and named, fewer than five variants shipped rather than padded, the surviving efficiency order stated out loud, and a humanizer pass before the set is scored final.",
      "files": [],
      "expectations": [
        "Deletes the trigger-event, competitor or market-reference, and company-initiative angles from the candidate set outright rather than ranking them last",
        "Names each deleted angle and the specific constraint that deleted it",
        "Ships fewer than the requested 5 variants rather than padding the set with a second variant on an angle already used",
        "States that a sub-threshold variant is never shipped to fill out a set",
        "Gives the surviving efficiency order explicitly as internal-style noun, then specific pain question, then executive-world anchor",
        "Orders angles by replies bought per research minute rather than by which is cheapest",
        "Notes that an SDR against a fixed daily send quota is spending sends rather than research hours, so the angle order holds while the parked-question tests are what gets cut",
        "Every shipped subject line is at most 4 words and fits within roughly 35 characters",
        "No shipped variant contains the prospect's first name",
        "Confirms that 120 sends per day clears the 200-sends-per-variant floor within the planned test window",
        "Runs the surviving set through a humanizer pass and re-scores afterwards rather than shipping raw first-draft output"
      ]
    },
    {
      "id": 6,
      "prompt": "Skiffline - Kubernetes cost optimization, selling to Platform Engineering leads. Our current subject is \"your aws bill\" and it's pulling 44% opens, way above anything we've had. The body is a short intro to what we do plus a request for 15 minutes. Give me 3 more in the same style so we can keep the streak going.",
      "expected_output": "A flagged body-promise mismatch scored as a hard fail, the 44% open rate refused as evidence of success, and three angle-distinct replacements rather than restyled synonyms, with the body work handed off.",
      "files": [],
      "expectations": [
        "Flags the mismatch between a subject implying a specific finding about their AWS bill and a body that only gives a generic company intro and a meeting request",
        "Classifies a subject promising what the body does not deliver as a hard fail scoring the variant at zero",
        "States that a subject outrunning the body raises opens while replies and trust collapse",
        "Does not write or rewrite the email body",
        "Treats the 44% open rate as directional only and refuses to read it as evidence the subject is working",
        "Notes that 40-45% sits in the vendor-published good open band against an average around 27.7%, and labels those figures vendor-published rather than independently audited",
        "Names reply rate as the metric that would actually settle whether the subject is working",
        "Refuses to produce three synonym variations in the same style and instead gives variants each on a distinct named angle",
        "Every shipped subject line is at most 4 words and fits within roughly 35 characters",
        "Scores every shipped variant against the body-promise-match criterion specifically"
      ]
    },
    {
      "id": 7,
      "prompt": "Day 3 of our subject test at Vantera. Variant A is \"pipeline coverage\" - 41% opens on 180 sends. Variant B is \"Is Your Q3 Coverage Slipping?\" - 29% opens on 174 sends. A is clearly winning so I want to kill B and push A to the remaining 5,000 contacts today. One thing: our unsubscribe rate went from 0.3% to 0.9% this week. We sell revenue-intelligence software to sales leaders.",
      "expected_output": "A refusal to call the winner on three independent grounds - sample, duration, and metric - plus the confound between A and B named, the guardrail breach escalated over the apparent win, and the next single-variable test named from the ranked queue.",
      "files": [],
      "expectations": [
        "Refuses to call a winner on day 3",
        "Names the 180 and 174 sends as below the 200-sends-per-variant floor",
        "Names the three elapsed days as below the one-full-week minimum needed to cover day-of-week effects",
        "Identifies stopping on an apparent early winner as peeking and states that it manufactures false positives",
        "Refuses to call a B2B cold winner on open rate, citing phantom opens registered by mail-privacy proxies pre-fetching tracking pixels",
        "Identifies that A and B differ on more than one variable at once - angle, casing, and question versus statement form - so no difference can be attributed",
        "States the one-variable-per-test rule and that a confounded pair must be re-run rather than interpreted",
        "Flags the 0.9% unsubscribe rate as breaching the 0.5% guardrail threshold",
        "States that a variant winning the decision metric while tripping a guardrail loses",
        "Recommends stopping or investigating on the guardrail breach rather than scaling the send to the remaining 5,000 contacts",
        "Names the angle test as the first single-variable test to run from the ranked queue, ahead of casing"
      ]
    },
    {
      "id": 8,
      "prompt": "It's Tuesday. My VP wants a subject line test result on her desk Friday morning. We send 400 a day through our sequencer to HR Directors at mid-market companies. The angle we've settled on is that they're all hiring recruiters right now - we scrape job boards into a sheet every Monday so we always know who posted what. Give me the test plan.",
      "expected_output": "A test plan whose duration floor is stated as unmet by Friday, the monitored job-board feed named as the fact promoting the trigger-event angle, a falsifiable hypothesis in the required form, and replies rather than opens as the decision metric despite the deadline.",
      "files": [],
      "expectations": [
        "States that a Tuesday-to-Friday window cannot satisfy the one-full-week minimum test duration",
        "Refuses to promise a decided winner by Friday and says what can honestly be reported by then instead",
        "Notes that a deadline inside the week promotes the near-zero-research angles and cuts the parked-question tests entirely",
        "Recognises the weekly job-board scrape as a monitored trigger feed and promotes the trigger-event angle on that basis despite its last place on the default efficiency order",
        "Names the feed as the specific fact that moved the trigger-event angle up the ranking",
        "Writes a falsifiable hypothesis before drafting any variant",
        "The hypothesis follows the form: Because [observation], we believe [change] will [effect] for [audience]. We'll know when [metric]",
        "Confirms that 400 sends per day clears the 200-per-variant volume floor while the calendar still fails the duration floor",
        "Sets the decision metric to replies, positive replies, or meetings booked and refuses opens despite the short deadline",
        "Holds the test to one variable, the subject-line angle, with sender name, send time, body, and audience held constant",
        "Delivers a test plan naming split, sample per variant, duration, decision metric, and guardrails"
      ]
    }
  ],
  "trigger_queries": [
    { "query": "write me 5 subject lines for a cold email to a VP of Operations", "should_trigger": true },
    { "query": "test my subject line", "should_trigger": true },
    { "query": "nobody opens my emails", "should_trigger": true },
    { "query": "our open rate is 9% and I don't know why", "should_trigger": true },
    { "query": "A/B test two email subjects for next week's product launch", "should_trigger": true },
    { "query": "what should I put in the subject field for this cold outreach", "should_trigger": true },
    { "query": "need preview text for our cart abandonment email", "should_trigger": true },
    { "query": "how many sends do I need per variant before I can call a winner on an email test", "should_trigger": true },
    { "query": "variant A got 34% opens, variant B got 28%, which one do I ship", "should_trigger": true },
    { "query": "can I use Re: at the start to get more opens", "should_trigger": true },
    { "query": "should our email subjects be lowercase or title case", "should_trigger": true },
    { "query": "score these three subject lines for me", "should_trigger": true },
    { "query": "give me some options for the top line of this email so people actually click it", "should_trigger": true },
    { "query": "is it ok to put the prospect's first name up top in outbound", "should_trigger": true },
    { "query": "my emails get opened but nobody replies", "should_trigger": true },
    { "query": "draft subject line variants for our Q3 nurture campaign", "should_trigger": true },
    { "query": "how long should an email subject be for cold outbound", "should_trigger": true },
    { "query": "does emoji help or hurt in promotional email", "should_trigger": true },
    { "query": "I want to split test what people see in their inbox before they open", "should_trigger": true },
    { "query": "we need a test plan for the subject line on our renewal campaign", "should_trigger": true },
    { "query": "what's a good open rate for cold email and how do I get there", "should_trigger": true },
    { "query": "help me write the line at the top of a cart abandonment email", "should_trigger": true },
    { "query": "our subject lines all sound the same, give me genuinely different angles", "should_trigger": true },
    { "query": "do any of these subject lines contain spam trigger words", "should_trigger": true },
    { "query": "pick the winning subject line from our last test", "should_trigger": true },
    { "query": "I need 4 different hooks for the inbox preview on this promo blast", "should_trigger": true },
    { "query": "how do I know if my email subject test is statistically valid", "should_trigger": true },
    { "query": "write subject lines for a cold email to Heads of Data Engineering at Series B fintechs", "should_trigger": true },
    { "query": "should I use numbers in my email subjects", "should_trigger": true },
    { "query": "what goes in the subject box", "should_trigger": true },
    { "query": "our cold emails get deleted without ever being opened", "should_trigger": true },
    { "query": "make these subject lines shorter", "should_trigger": true },
    { "query": "rewrite our email subject so it doesn't look like a mail merge", "should_trigger": true },
    { "query": "design an experiment to find the best subject for our welcome email", "should_trigger": true },
    { "query": "how many subject variants can I test with a 900 person list", "should_trigger": true },
    { "query": "give me something short and plain for the subject, nothing salesy", "should_trigger": true },
    { "query": "what's the character count before mobile cuts off the subject", "should_trigger": true },
    { "query": "our email A/B test came back inconclusive, what now", "should_trigger": true },
    { "query": "I need the first thing they read in their inbox to land better", "should_trigger": true },
    { "query": "check our SPF and DKIM records before we start sending", "should_trigger": false },
    { "query": "our sending domain got blacklisted, what do I do", "should_trigger": false },
    { "query": "audit this cold email body for deliverability risk", "should_trigger": false },
    { "query": "should we use plain text or HTML for cold outbound", "should_trigger": false },
    { "query": "how many links can I put in a cold email before it hurts inbox placement", "should_trigger": false },
    { "query": "is our unsubscribe link compliant with GDPR", "should_trigger": false },
    { "query": "we moved to a new sending domain and everything went to spam", "should_trigger": false },
    { "query": "what's a safe daily sending volume per mailbox during warmup", "should_trigger": false },
    { "query": "find me something to personalize this outreach with for this prospect", "should_trigger": false },
    { "query": "rank the personalization angles for this account", "should_trigger": false },
    { "query": "what signals should I look for before emailing a prospect", "should_trigger": false },
    { "query": "is this trigger event recent enough to use in outreach", "should_trigger": false },
    { "query": "how many touches should my outbound cadence have", "should_trigger": false },
    { "query": "what day should the second follow up go out", "should_trigger": false },
    { "query": "design a 12 day sequence mixing email, LinkedIn and calls", "should_trigger": false },
    { "query": "when should a prospect exit the sequence", "should_trigger": false },
    { "query": "write the first 20 seconds of a cold call", "should_trigger": false },
    { "query": "how do I open a cold call without getting hung up on", "should_trigger": false },
    { "query": "do I need to scrub this list against do-not-call", "should_trigger": false },
    { "query": "score this deal against MEDDPICC", "should_trigger": false },
    { "query": "write the recap email after my discovery call", "should_trigger": false },
    { "query": "build a discovery question set for a 30 minute call", "should_trigger": false },
    { "query": "our pipeline coverage ratio is 2.1x, is that enough", "should_trigger": false },
    { "query": "handle this price objection from a CFO", "should_trigger": false },
    { "query": "who is the economic buyer in this deal", "should_trigger": false },
    { "query": "define our ICP from closed won data", "should_trigger": false },
    { "query": "set quotas for next year's AE team", "should_trigger": false },
    { "query": "which sales podcasts should I follow", "should_trigger": false },
    { "query": "design the comp plan for our first two SDRs", "should_trigger": false },
    { "query": "should we go product-led or sales-led", "should_trigger": false },
    { "query": "write the body copy for this cold email", "should_trigger": false },
    { "query": "write a nurture email sequence for our trial users", "should_trigger": false },
    { "query": "what's the best send time for our newsletter", "should_trigger": false },
    { "query": "build a landing page headline for this campaign", "should_trigger": false },
    { "query": "write an SMS blast for our Black Friday sale", "should_trigger": false },
    { "query": "set up an email drip in our marketing automation tool", "should_trigger": false },
    { "query": "how do I configure SMTP for transactional email", "should_trigger": false },
    { "query": "write the push notification copy for our mobile app", "should_trigger": false },
    { "query": "write a LinkedIn connection request message for this prospect", "should_trigger": false }
  ]
}
references/pattern-library.md
# Subject Line Pattern Library

Every statistic below is vendor-published marketing data from companies selling email tooling - email-coaching products, revenue-intelligence platforms, outbound agencies, and sales-engagement platforms. None is independently audited or peer-reviewed, and the percentages from different vendors do not reconcile with each other. Trust the direction, never the magnitude.

No named, cross-source framework exists for subject lines; what follows is the unbranded convergence across independent sources plus clearly labeled single-source claims. Company names in examples (Northwind, Acme) are fictional.

## Table of Contents

- [B2B cold outbound](#b2b-cold-outbound)
- [B2C lifecycle and promotional](#b2c-lifecycle-and-promotional)
- [Parked questions - ranked test queue](#parked-questions-ranked-test-queue)
- [Rules identical in B2B and B2C](#rules-identical-in-b2b-and-b2c)

## B2B cold outbound

Regime summary: short, plain, anchored to the recipient's world. The best-corroborated rule in the whole domain - stated independently by an email-coaching vendor, by two separate revenue-intelligence datasets (25-28M cold emails; 1M+ executive sales cycles), and by an outbound agency (5.5M emails) - is to keep it to roughly 1-4 words.

### Length

- Target 1-4 words, at most ~35 characters. Mobile clients truncate around 30-35 characters, so brevity is a hard constraint, not a style choice.
- Vendor figures for calibration: 2-word subjects ~60% more opens than 5-word (email-coaching vendor); 2-4 words ~46% open rate vs. 34% at 10 words (outbound agency, 5.5M emails); moving from 2 to 4 words reduced replies ~17.5% (same email-coaching vendor).

### Angle catalog - pick one per variant

Ranked by replies bought per research minute spent, highest ratio first - not by what is cheapest. Research cost is the only effort that counts here: minutes spent per prospect, and whether the angle is written once for a whole list or rewritten prospect by prospect.

- efficiency: `internal-style noun > specific pain question > executive-world anchor > competitor == company initiative > trigger event`
- value, replies when it lands: `trigger event > internal-style noun > executive-world anchor > competitor == company initiative > specific pain question`
- research effort: `trigger event > competitor == company initiative > executive-world anchor > internal-style noun == specific pain question`

Both ties are real:

- `internal-style noun == specific pain question` on effort: neither adds a per-prospect lookup - one comes from the body's own topic, the other from the persona pain known before opening the list.
- `competitor == company initiative` on both axes: they cost one pass over the same public account pages, survive the same sequence, and buy the same thing - proof the sender read the account - with no vendor measurement separating them.

Casing in the examples below is arbitrary. Capitalization is contested (see Parked questions), so treat casing as a test variable, not part of the pattern.

| Angle                         | Research cost                                                       | Mechanism                                                                                                                                          | Good example                                   | Bad example                            |
| ----------------------------- | ------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------- | -------------------------------------- |
| Internal-style plain noun     | near-zero; written once, reused across the whole list               | Reads like a colleague's note, not a vendor's pitch; internal-looking subjects roughly doubled opens in a revenue-intelligence study (85M+ emails) | "reply rates" · "trial delays" · "Q2 forecast" | "Boost your reply rates by 300%"       |
| Specific pain question        | near-zero; one pain per persona, reused across the whole segment    | A question naming their exact pain; generic questions fail (see Parked questions)                                                                  | "losing ramp time?"                            | "Quick question?" · "Have 15 minutes?" |
| Executive-world anchor        | an hour per segment; reused for every exec in it                    | Anchors to the exec's stated priorities, never the sender's product (revenue-intelligence executive dataset)                                       | "student dropout risk"                         | "Our platform for education leaders"   |
| Competitor / market reference | an hour per account; survives the whole sequence                    | Names a rival or market shift they track                                                                                                           | "Northwind migration"                          | "Beat your competitors today"          |
| Company initiative            | an hour per account; survives the whole sequence                    | Names a public initiative or project they own                                                                                                      | "Acme expansion plan"                          | "Partnership opportunity!"             |
| Trigger event                 | a standing job; sourced per prospect, goes stale fast, reused never | References something that just happened at their company                                                                                           | "your sdr hiring"                              | "Congrats on the news!!"               |

Default: draft the set from the top of the efficiency order downwards, and stop where the research budget runs out.

**What this order starves: the trigger event.** It has the highest ceiling of any angle here and loses every efficiency round, because a fresh trigger per prospect never amortises across a list.

Promote it anyway when:

- The list is short enough that a day of sourcing covers it.
- A trigger feed is already licensed and monitored.
- The campaign targets named accounts where one missed reply costs more than the sourcing does.

Re-rank against what you already know about the user, and name which fact moved which angle:

- A licensed intent or technographic feed makes the competitor angle near-zero - promote it alongside the internal-style noun.
- No trigger data and no budget to source it deletes the trigger-event angle from the candidate set. Delete it, say so, and draft one fewer variant; an angle parked at the bottom reappears later as work nobody scoped.
- The same deletion hits competitor and company initiative where account-by-account research is not possible.
- An executive-only list promotes the executive-world anchor above the specific pain question: the exec dataset backs the anchor, and the question form is contested.
- An SDR against a fixed daily send quota is spending sends, not research hours - the angle order stands, but the parked-question tests are what gets cut.

This order is a default, not a law. It shifts with list size, recipient seniority, and who executes it; a researcher with account-plan time and an SDR clearing a quota should not draft the same set.

### Anti-patterns with vendor-measured impact

Fix in the order listed. Avoiding any one of these costs the same near-zero effort - a word choice at drafting time, no research and no extra send - so measured harm alone orders the table, and reply damage outranks open damage because replies are what calls the winner in this regime. `empty subject == prospect's first name` on harm: both land on the same measured -12% replies, and nothing in the source set separates them further.

| Anti-pattern                              | Reported impact (vendor-published)                                                                          |
| ----------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| Fake "Re:" / "Fwd:" prefix                | Hard fail, no measurement needed: destroys trust and is the exact deceptive-subject case CAN-SPAM prohibits |
| Pitching the product in the subject       | -57% replies                                                                                                |
| Prospect's first name in the subject      | ~12% fewer replies (attributed to a sales-engagement platform's data) - signals mail-merge                  |
| Empty subject                             | +30% opens but -12% replies                                                                                 |
| Numbers and percentages                   | -46% opens                                                                                                  |
| Excessive punctuation ("!!!", "??")       | -36% opens                                                                                                  |
| Urgency words ("ASAP", "urgent")          | opens fall below 36%                                                                                        |
| Salesy verbs ("increase", "boost", "ROI") | -17.9% opens                                                                                                |

### Contested question: capitalization

- One email-coaching vendor teaches title case and claims skipping it costs roughly 30% of opens.
- Guidance derived from a revenue-intelligence dataset, and a popular paid cold-email training program, teach lowercase instead and claim the data supports that.

These are directly contradictory vendor claims with no independent arbiter. Do not enforce either.

Pick one casing for the current test and hold it constant across every variant; casing itself is ranked in Parked questions below.

### Contested question: questions in the subject

- An outbound agency reports question subjects performing well (~46% open).
- An email-coaching vendor reports questions cutting opens by ~56%.

Practitioner reconciliation: a specific pain question can work; a generic one ("Quick question?") reads as a mass template and fails. Default to statements when unsure.

### Executive recipients

- Executives process hundreds of emails a day and decide in under ~3 seconds (revenue-intelligence study, 1M+ exec cycles).
- Ultra-concise and understated wins; anything salesy is rejected on sight.
- On an executive list, the executive-world anchor moves up the efficiency order, ahead of the specific pain question: their initiative, their risk, their metric is the only framing that survives three seconds.

## B2C lifecycle and promotional

Regime summary: a genuinely different game. Clarity and stated benefit win; several things that measure as negatives in B2B cold (numbers, longer lines) are normal here.

### Length and preheader

- Subject: 40-60 characters ideal; keep to ~50 or less to survive mobile truncation.
- Preheader: ~90-140 characters. It must extend the subject - complete the thought or add intrigue - never repeat it. Subject plus preheader is the tested unit.

### Patterns that work

No value ordering here, deliberately: which pattern wins is decided by what the body contains, not by the pattern itself. A shipping confirmation cannot be a story tease, and a nurture essay cannot be a Direct - ranking these against each other across campaigns would be false precision. So one axis orders the rows, and it is stated rather than implied:

- production effort: `story tease > how-to > direct == number == question`
- Pick the cheapest pattern the body already supports. The three-way tie is genuine: direct, number and question are each one line written from content already sitting in the email, with no extra drafting pass.

| Pattern     | Production effort                                                                                                                 | Good example                               | Bad example                         |
| ----------- | --------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------ | ----------------------------------- |
| Direct      | near-zero; the event names itself                                                                                                 | "Your report is ready"                     | "We have something for you"         |
| Number      | near-zero; count what the body already lists                                                                                      | "3 ways to save on your renewal"           | "37 incredible unmissable deals!!!" |
| Question    | near-zero; name the pain the body answers                                                                                         | "Still struggling with meal planning?"     | "Want to know a secret?"            |
| How-to      | an hour, and only if the body genuinely teaches something                                                                         | "How to cut your grocery bill in one week" | "How to change your life"           |
| Story tease | an hour; needs a real story, and the only pattern that fails silently - intrigue that never pays off costs trust on the next send | "The pricing mistake I almost made"        | "You won't believe what happened!"  |

- Clear beats clever; specific beats vague. If a reader must decode the line, ship the clearer variant.
- Numbers are fine in this regime, unlike B2B cold.
- Emoji is polarizing: never a default, and ranked as a test candidate in Parked questions below.

## Parked questions - ranked test queue

Three questions in this file have no answer worth asserting:

- B2B casing.
- B2B question-versus-statement.
- B2C emoji.

Each is resolvable, and each is paid for in sends rather than hours - a two-arm test needs 200+ sends per variant, so roughly 400+ sends buys one answer - the test-design reference carries that floor and the multi-variant multiplier.

- expected gain: `angle test > casing > question form == emoji`
- send cost: `angle test == casing == question form == emoji` - each is one binary variable resolved by the same two-arm test at the same per-variant floor, so the send bill is identical and gain alone orders the queue
- efficiency: `angle test > casing > question form == emoji`

Run the angle test first, always. Angle carries the only large measured effects in this file - an internal-style noun roughly doubling opens, pitching in the subject costing -57% replies - while every parked question rests on a single unreconciled vendor claim.

| Parked question                   | What resolving it buys                                                                    | Why it ranks here                                                                                                                                 |
| --------------------------------- | ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| Casing (title case vs. lowercase) | A rule applying to every subject sent afterwards - a compounding asset, not a one-off win | The only parked question with no practitioner reconciliation at all, and the cleanest single variable: casing changes nothing else about the line |
| Question vs. statement            | A default for one angle, and a usable one already exists (specific works, generic fails)  | Hard to run cleanly - a question and a statement are usually different angles too, so the test confounds unless the angle is held identical       |
| B2C emoji                         | An audience-specific answer that does not transfer to another list                        | Same send bill, narrowest scope; the answer expires whenever the audience changes                                                                 |

`question form == emoji` on gain: both buy an answer scoped to one angle or one audience, where casing buys a rule over every future send.

Delete rather than queue: a list too small to power one two-arm test at the floor above has no test queue. Say that plainly and ship the defaults - hold casing constant, statements over questions, emoji off - instead of leaving a queue nobody can run.

## Rules identical in B2B and B2C

These apply unchanged to both regimes:

- One variable per test, with a written falsifiable hypothesis before any variant is drafted.
- Mobile truncation discipline - write for the smallest screen the audience uses.
- Never promise in the subject what the body does not deliver.
- Deceptive subject lines are illegal under CAN-SPAM in both regimes, not merely bad practice.
- Variants must differ by angle, never by synonym swap.
references/scoring-rubric.md
# Subject Line Scoring Rubric

Score every candidate variant in two passes: hard fails first, then the 10-point scorecard. Pass threshold: at least 8/10 with zero hard fails, for every shipped variant. Regenerate and re-score anything below that until the whole set clears.

## Hard fails - score 0, regenerate, no exceptions

- Fake "Re:" or "Fwd:" prefix, or any framing implying a prior conversation that never happened. This is the deceptive-subject-line case CAN-SPAM prohibits.
- Subject promises anything the email body does not actually deliver.
- Spam-trigger vocabulary in the subject: "free", "guarantee", "act now", "limited time", "click here". Flag the words here; route the full inbox-placement review to the deliverability skill.
- B2B cold only: the prospect's first name in the subject - a mail-merge signal correlated with fewer replies (attributed to a sales-engagement platform's data).
- Exceeds the regime's truncation budget: ~35 characters for B2B cold, ~50 for a B2C subject.
- Duplicate angle: the variant differs from another shipped variant only in wording, not in angle.

## 10-point scorecard, per variant

| Criterion                 | Points | Full marks means                                                                                                                  |
| ------------------------- | ------ | --------------------------------------------------------------------------------------------------------------------------------- |
| Length and truncation fit | 0-2    | B2B cold: 1-4 words, within ~35 chars. B2C: within ~50 chars, preheader 90-140 chars that extends (never repeats) the subject     |
| Angle execution           | 0-2    | Concretely expresses the chosen personalization angle - pain, trigger event, competitor, initiative - not a generic gesture at it |
| Reader's-world framing    | 0-2    | Anchored to the recipient's job or life, not the sender's product; no salesy verbs, no urgency words, no pitch                    |
| Body-promise match        | 0-2    | The body's opening fully honors what the subject implies                                                                          |
| Set distinctness          | 0-1    | Angle genuinely distinct from every other variant in the shipped set                                                              |
| Human sound               | 0-1    | Survives being read aloud as something a person would type; no templated AI phrasing                                              |

## Worked scored example - B2B cold

Context: recipient is a VP Operations; chosen angle is a trigger event (they posted five SDR job openings); body offers a benchmark on SDR ramp time.

| Check                     | "sdr ramp time"                                         | "Quick question?"                     |
| ------------------------- | ------------------------------------------------------- | ------------------------------------- |
| Hard fails                | none                                                    | none                                  |
| Length and truncation fit | 2 - 3 words, 13 chars                                   | 2 - 2 words, 15 chars                 |
| Angle execution           | 2 - names the exact consequence of their hiring trigger | 0 - no angle at all                   |
| Reader's-world framing    | 2 - their metric, no product                            | 1 - neutral but empty                 |
| Body-promise match        | 2 - body delivers the ramp benchmark                    | 1 - body content is a surprise        |
| Set distinctness          | 1                                                       | 0 - interchangeable with any campaign |
| Human sound               | 1                                                       | 1                                     |
| **Total**                 | **10/10 - ship**                                        | **5/10 - regenerate**                 |

## Worked scored example - B2C lifecycle

Context: cart-abandonment email; body contains the saved cart and a free-shipping offer.

| Check                     | "Your cart is saved - shipping's on us"                 | "You won't BELIEVE this deal!!!"                                                                                |
| ------------------------- | ------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- |
| Hard fails                | none                                                    | fails: excessive punctuation reads as spam bait and the body has one specific offer, not an unbelievable "deal" |
| Length and truncation fit | 2 - 38 chars; preheader extends with the offer deadline | -                                                                                                               |
| Angle execution           | 2 - direct pattern, names the concrete benefit          | -                                                                                                               |
| Reader's-world framing    | 2 - their cart, their saving                            | -                                                                                                               |
| Body-promise match        | 2 - body shows the cart and the offer                   | -                                                                                                               |
| Set distinctness          | 1                                                       | -                                                                                                               |
| Human sound               | 1                                                       | -                                                                                                               |
| **Total**                 | **10/10 - ship**                                        | **0 - hard fail, regenerate**                                                                                   |

## Delivery table template

Present the shipped set to the user in this shape (add a Preheader column for B2C):

| #   | Variant | Angle                  | Words / chars | Score | Notes                                   |
| --- | ------- | ---------------------- | ------------- | ----- | --------------------------------------- |
| A   | ...     | trigger event          | 3 / 14        | 9/10  | control                                 |
| B   | ...     | internal-style noun    | 2 / 11        | 8/10  | challenger                              |
| C   | ...     | specific pain question | 4 / 19        | 8/10  | park if volume only supports 2 variants |
references/test-design.md
# Subject Line Test Design

How to turn a scored variant set into a test that produces a trustworthy decision. Benchmarks cited here are vendor-published and not independently audited - use them to calibrate expectations, not as targets.

## Hypothesis

Write it before drafting any variant, in this form:

> Because [observation], we believe [change] will [effect] for [audience]. We'll know when [metric].

- Weak: "Let's test a new subject line."
- Strong: "Because replies stall when subjects sound vendor-written, we believe a plain-noun internal-style subject will raise reply rate for ops leaders. We'll know when reply rate over 200+ sends per variant beats the control."

A hypothesis that cannot lose is not a hypothesis. If no result could prove it wrong, rewrite it.

## One variable, no peeking

- Test exactly one variable: the subject-line angle (or, in B2C, the subject+preheader pair as one unit). Hold sender name, send time, body, and audience constant.
- Pre-commit to sample size and duration before launch. Checking results early and stopping on an apparent winner manufactures false positives - the most common way subject tests go wrong.
- Never change a variant mid-test. A changed variant is a new test.

## Sample size and duration

Practitioner rules of thumb - useful floors, but not real power calculations, and say so to the user:

- Outbound: 200+ sends per variant for a subject-line test; 500+ per variant for a send-time test.
- List-based ESP campaigns: ~20% of the list split across test variants, winner to the remainder. On a small list this can still be underpowered.

A real power calculation needs four inputs:

- Baseline rate for the decision metric.
- Minimum detectable effect.
- Significance level (95%).
- Statistical power (80%).

- efficiency: `rule-of-thumb floor > power calculation`. The floor costs near-zero and is right often enough to plan around; the calculation costs an hour and needs a baseline rate the user may not have yet.
- Promote the power calculation when the list is large enough that the extra sends are free, or when the decision is expensive enough that shipping an underpowered result would cost more than the hour.

- Duration: at least one full week to cover day-of-week effects; 2-4 weeks is typical.
- More variants multiply the sample: 3 variants need ~1.5x the total, 4 need ~2x. At low volume, cap the test at 2-3 variants and park the rest.
- Send-timing note - the one documented B2B/B2C split: B2B avoids weekends; B2C should test weekends.

## Metric choice - the decision that breaks most tests

- Open rate has been distorted since 2021: mail-privacy features pre-fetch tracking pixels through proxies, registering phantom opens the recipient never made. Treat open rate as directional only.
- B2B cold decision metric: reply rate, positive reply rate, or meetings booked. Never call a winner on opens alone.
- B2C lifecycle: open rate is still usable as the primary read, but pair it with a downstream metric (click or conversion) so a "winner" that attracts opens and repels action gets caught.

Which of the three to actually decide on, in B2B cold:

- trust: `meetings booked > positive replies > replies > opens`
- measurement effort: `meetings booked > positive replies > replies == opens`
  - Reply count and open count are both auto-tallied in the same sequencer report.
  - Classifying replies as positive needs a human pass over each one.
  - Meetings need the CRM wired to the sequencer, an hour of setup once.
- efficiency: `replies > positive replies > meetings booked > opens`

Default to raw reply rate: auto-counted, and trustworthy enough to call a subject-line winner. Promote positive replies when reply volume is high enough that "not interested" replies could carry the win, and meetings booked when the sequencer already writes to the CRM. Opens lose on efficiency despite costing nothing, because an untrustworthy metric buys no decision at any price.

## Guardrails - stop the test if these degrade

- Unsubscribe rate: keep under 0.5%.
- Spam-complaint rate: keep under 0.1%.
- Bounce rate: any significant rise means a list or sending problem, not a subject-line result.

A variant that wins the decision metric while tripping a guardrail loses.

## Calibration benchmarks (vendor-published; trust direction, not magnitude)

| Context       | Metric          | Average                 | Good                                                                       |
| ------------- | --------------- | ----------------------- | -------------------------------------------------------------------------- |
| B2B cold      | Open rate       | ~27.7%                  | 40-45% (excellent 50%+) - per outbound-agency and prospecting-tool figures |
| B2B cold      | Reply rate      | 4-5.8%                  | 5-10% - down from 7-8% in 2020-2022; expect continued decline              |
| B2C lifecycle | Open rate       | 20-40% by sequence type | -                                                                          |
| B2C lifecycle | Unsubscribe     | -                       | under 0.5%                                                                 |
| Both          | Spam complaints | -                       | under 0.1%                                                                 |

## Test plan template

Deliver and log every test in this shape:

- **Hypothesis:** the falsifiable statement above.
- **Variants:** IDs and angles from the scorecard (all at or above the quality gate).
- **Split:** e.g. 50/50, or 20% of list for the test with winner to remainder.
- **Sample per variant:** the pre-committed number, and whether it came from a rule of thumb or a power calculation.
- **Duration:** start date, end date, minimum one full week.
- **Decision metric:** the single metric that calls the winner.
- **Guardrails:** unsubscribe, complaints, bounce - with stop thresholds.
- **Result:** winner / loser / inconclusive, with the numbers. An honest inconclusive is a valid outcome; do not torture it into a winner.
- **Next hypothesis:** what this result suggests testing next - one variable, taken from the ranked test queue in the pattern library, or an explicit "volume does not support another test".
SKILL.md
---
name: cold-email-subject-line-tester
description: Generates, scores, and split-tests email subject lines, producing angle-distinct variants, a numeric scorecard, and a test plan with sample size and decision metric. Covers B2B cold outbound and B2C lifecycle/promotional email, and flags spam-trigger words inside the subject line only. Use whenever the user mentions a subject line, open rates, email A/B test, preview text, or "nobody opens my emails", even if they don't say "subject line" explicitly. Do NOT use for body copy or inbox-placement audits (mbfinotti/sales-skills@cold-email-deliverability), personalization angles (mbfinotti/sales-skills@sales-outreach-personalization), or cadence design (mbfinotti/sales-skills@sales-outbound-sequence).
license: MIT
metadata:
  author: Maya-Beth Finotti
  version: "1.4.3"
---

# Subject Line Tester

Generate subject line variants from an already-chosen personalization angle, score them against a numeric rubric, and design a statistically honest split test to pick the winner.

Two facts frame everything here:

- No named, cross-source framework exists for subject lines the way MEDDIC exists for qualification - only single-vendor branded methods and one unbranded convergence (keep B2B cold subjects short; avoid numbers and question marks in them).
- Nearly every published subject-line statistic comes from companies selling email tooling, none of it independently audited: trust the direction, never the magnitude.

Where sources flatly contradict each other (notably lowercase vs. title case), treat the question as a test candidate, never a rule.

Scope: the subject line (plus preheader for B2C) is the whole deliverable.

Never:

- Write body copy.
- Audit deliverability, headers, domains, or authentication.
- Research prospect signals.
- Plan touch-by-touch cadence.

Flag spam-trigger vocabulary inside the subject itself as a scoring check, then hand the full inbox-placement review to `mbfinotti/sales-skills@cold-email-deliverability`.

## Interview

Ask one question per message. Offer multiple-choice options. Skip anything already answered by context. If your harness has persistent memory and a prior run stored this user's answers, confirm them instead of re-asking.

1. Regime: B2B cold outbound, or B2C lifecycle/promotional? Ask this first - the two follow different rules.
2. Recipient persona, role, and seniority? Executive inboxes decide in seconds; the pattern set narrows sharply for them.
3. Which personalization angle is already chosen - pain point, competitor, company initiative, executive priority, shared context, trigger event? If none, stop and route to `mbfinotti/sales-skills@sales-outreach-personalization` first. Ranking these is step 4's job, not this question's.
4. What does the email body actually promise or deliver? The subject must never outrun it.
5. Can your sending platform (sequencer or ESP) split-test? At what volume (list size or sends per day)?
6. Which decision metric can you trust in your setup: replies or meetings, clicks or conversions, or only opens?
7. By what date must the result land? A deadline inside the week promotes the near-zero-research angles and cuts the parked-question tests entirely; a date far enough out to collect 200+ sends per variant keeps a test in scope.
8. One-off win on this campaign, or a compounding asset? Compounding promotes the reusable angles (internal-style noun, persona pain, executive anchor) and the casing test; one-off promotes whichever angle today's list already supports.
9. Effort ceiling: how many research minutes per prospect are affordable, and how many sends per day are available? Near-zero minutes deletes the trigger-event, competitor and initiative angles outright; a quota too small to reach the per-variant floor deletes the test queue.
10. B2C only: do you control the preheader text, and is emoji acceptable to the brand?

Re-rank the angle catalog against the answers to 7-9 before drafting, and tell the user which answer moved which angle.

## Workflow

1. Run the Interview. Confirm regime, persona, angle, body promise, test capacity, metric, deadline, payoff horizon, and effort ceiling before drafting.
2. Read the matching regime section of [references/pattern-library.md](references/pattern-library.md). Never apply B2B cold rules to B2C lifecycle or vice versa.
3. Write a falsifiable hypothesis before drafting: "Because [observation], we believe [change] will [effect] for [audience]. We'll know when [metric]." Reject vague forms like "let's test a new subject line."
4. Draft 3-5 candidate variants, working down the catalog's efficiency order - for B2B cold, `internal-style noun > specific pain question > executive-world anchor > competitor == company initiative > trigger event`, ranked by replies bought per research minute. Delete any angle the interview ruled out instead of ranking it last, and say which angle was deleted and why. Each variant must express a distinct angle, never a synonym swap of another. For B2C, draft a preheader with every variant - subject plus preheader is the tested unit.
5. Score every variant with [references/scoring-rubric.md](references/scoring-rubric.md): hard-fail checks first, then the 10-point scorecard.
6. Regenerate any variant below the pass threshold, then re-score. Repeat until the full set clears the quality gate.
7. Run the surviving set through your preferred humanizer skill, then re-score. Raw first-draft model output is never shipped - templated AI phrasing is exactly the smell that gets an email deleted unread.
8. If you can browse the web, verify any volatile fact a variant leans on (trigger event, competitor claim, initiative name). Otherwise, ask the user to confirm it before shipping.
9. Design the test with [references/test-design.md](references/test-design.md): one variable, pre-committed sample size, minimum duration, decision metric, guardrail metrics.
10. Deliver the output package (shape below). If your harness has persistent memory, record persona, chosen angle, hypothesis, and shipped variants so later runs skip the interview and build on results.
11. When results come back, log winner, loser, or inconclusive against the hypothesis, then take the next single-variable test from the ranked queue in [references/pattern-library.md](references/pattern-library.md) - or say plainly that the volume supports no further test.

## Invocation examples and expected output

Typical requests this skill handles:

- "Test my subject line for this cold email to a VP of Operations."
- "Write 5 subject line variants for our cart-abandonment campaign."
- "A/B test subject lines for next week's product launch email."
- "Why is my open rate low?" Diagnose the subject line only; route domain, authentication, and inbox-placement causes to `mbfinotti/sales-skills@cold-email-deliverability`.

Expected output package, in order:

1. Hypothesis statement in the falsifiable form above.
2. Variant scorecard table: variant, angle, word/character count, rubric score, notes (plus preheader column for B2C).
3. Test plan: split, sample per variant, duration, decision metric, guardrails.
4. Flagged risks: unverified facts, spam-trigger words found, angles deleted by the interview and why, and the parked questions in queue order with the send cost of resolving them.

Filled examples of the table and test plan live in [references/scoring-rubric.md](references/scoring-rubric.md) and [references/test-design.md](references/test-design.md).

## Quality gate

- Score every variant on the 10-point scorecard in [references/scoring-rubric.md](references/scoring-rubric.md).
- Pass threshold: every shipped variant scores at least 8/10 with zero hard fails.
- Hard fails (any one disqualifies the variant outright):
  - Fake "Re:"/"Fwd:" or other deceptive framing.
  - Subject promising what the body does not deliver.
  - Spam-trigger vocabulary.
  - Exceeding the regime's truncation budget.
  - Duplicate angle within the set.
  - Prospect first name in a B2B cold subject.
- Iterate: regenerate and re-score failing variants until the whole shipped set clears 8/10. Never ship a sub-threshold variant to "fill out" the test.

## KPIs and measurement

- B2B cold decision metric: reply rate, positive reply rate, or meetings booked - never opens alone. Mail-privacy proxies have pre-fetched tracking pixels since 2021, registering phantom opens; open rate is directional only.
- B2B calibration (vendor-published, not independently audited):
  - Average open ~27.7%, good 40-45%.
  - Average reply 4-5.8%, good 5-10%.
  - Falling year over year.
- B2C lifecycle: open rate is usable but must be paired with a downstream guardrail - click, conversion, unsubscribe under 0.5%, complaints under 0.1%.
- Success for this skill = the test reaches its pre-committed sample and produces a decision (winner, loser, or honest inconclusive) without any guardrail degrading.

## Failure modes

- Variant set is synonym swaps of one idea. Fix: force each variant onto a different angle from the pattern library.
- Capitalization enforced as a rule. Sources directly contradict each other (title case vs. lowercase, both vendor claims); hold it constant and park it in the ranked test queue.
- Efficiency order treated as law, or a ruled-out angle parked at the bottom of the set. The order is a default that shifts with list size, recipient seniority, and who executes it; an angle the interview ruled out gets deleted from the set and named, never demoted.
- Prospect first name in a B2B cold subject. It signals mail-merge and correlates with fewer replies; use a contextual token instead.
- Winner called on open rate in B2B cold. Phantom opens from privacy proxies manufacture false winners; decide on replies or meetings.
- Peeking and stopping early. Pre-commit to sample size and duration; an early "winner" is usually noise.
- Generic question subjects. Vendor data conflicts on questions overall; a specific pain question can work, a generic one fails. Default to statements.
- Subject outruns the body's promise. Opens rise, replies and trust collapse; the body-promise match is a scored criterion for this reason.
- B2B rules applied to B2C or vice versa. Numbers measure as a negative in B2B cold subjects yet perform fine in B2C lifecycle.
- Shipping raw model output. Always run the humanizer pass and re-score afterward.
cold-email-subject-line-tester · 熱門 Agent Skills | Mengbi