Skills に戻る
mbfinotti/revops-skillsチェック済み

SKILL DETAIL

lead-scoring

mbfinotti/revops-skills/lead-scoring

Design, validate, and recalibrate a lead scoring model - fit and engagement signals, weighting, negative scoring and exclusions, decay, MQL/PQL threshold and tier setting, backtesting against closed-won/closed-lost outcomes, and governance; also diagnoses an existing broken model. Use whenever the user mentions lead scoring, an MQL threshold, fit vs engagement scoring, product-qualified leads, score decay, "score my leads", "our lead scores are wrong", "sales rejects our MQLs", or "too many junk MQLs" - even if they never say "scoring model". Covers B2B sales-led and B2C/PLG, rules-based design and predictive readiness. Do NOT use for assigning scored leads to reps - use mbfinotti/revops-skills@lead-routing instead.

インストール · 162出典を見る

Installation

npx skills add https://github.com/mbfinotti/revops-skills --skill lead-scoring

スキルファイル

SKILL.md

最終同期 · 2026/09/15

evals/evals.json
{
  "skill_name": "lead-scoring",
  "evals": [
    {
      "id": 1,
      "prompt": "I'm the marketing ops lead at Harrowgate Systems - B2B workflow software, ACV around $22K, 9 AEs. We're building our first real lead scoring model. I pulled the last 90 days of inbound and the fit fields are patchy: employee count is populated on 43% of records, industry on 58%, job title on 91%. Our head of demand gen wants to buy an enrichment vendor to close the gap, but we have no procurement path this year - legal and finance are both frozen until January - and I'm the only person who touches the CRM, there's no analyst. The board wants this scoring live inside three weeks. What do we do about the fit fields?",
      "expected_output": "A coverage diagnosis that names the enrichment ceiling, an explicitly ranked three-option menu with enrichment deleted for lack of a procurement path, scoring unknowns 0 as the default, and the cohort bias that default introduces.",
      "files": [],
      "expectations": [
        "Identifies coverage below roughly 70% on a fit field as an enrichment-ceiling problem where most of the database can never reach the threshold",
        "Presents exactly three options for the coverage gap: score unknowns 0, enrich the field, drop the field",
        "States the ordering explicitly as a ranking by value per unit of effort rather than presenting an unordered list",
        "Ranks efficiency as: score unknowns 0 > enrich the field > drop the field",
        "Emits at least one separate ordering axis (value, effort, or compliance cost) whose order differs from the efficiency ordering",
        "Recommends scoring unknowns 0 as the default rung",
        "Deletes the enrichment option outright and says it is deleted, because the user reported no procurement path",
        "Names the bias scoring unknowns 0 buys: records that are thin for reasons other than fit, such as SMB, non-US, or privacy-conscious buyers, get demoted",
        "Recommends tracking those demoted records as a separate cohort and re-deriving the weight if they convert anyway",
        "States that dropping the field is the least reversible option in the menu because the signal stops being collected and its weight cannot later be re-derived",
        "Notes that neither scoring unknowns 0 nor dropping the field introduces a new data source, so neither triggers a consent or provenance review, while enrichment does",
        "Keeps job title at 91% coverage in the fit axis without proposing remediation for it",
        "Does not recommend delaying the model launch until enrichment is purchased"
      ]
    },
    {
      "id": 2,
      "prompt": "I inherited a lead scoring setup at Pellard Analytics when the last marketing ops person left. It's a single number out of 100: email open +2, email click +5, any page view +3, whitepaper download +10, webinar registration +10, blog visit +3, demo request +10, and job title containing 'director' or 'manager' +10. Threshold is 100. No decay, no caps, nothing is suppressed. Sales stopped opening MQL tasks around March and the VP Sales told me in a 1:1 that he thinks the score is astrology. I have a quarter and one analyst. My instinct is to rip it out and start over - worth it?",
      "expected_output": "A full symptom sweep across the described model naming several concurrent faults, with the fixes ordered by value per unit of effort rather than by the order the faults appear.",
      "files": [],
      "expectations": [
        "Checks the full symptom list against the described model rather than proposing a rebuild off the first fault found",
        "Names at least four distinct faults in the described model",
        "Flags that a demo request at +10 is outscored by accumulated low-intent activity because nothing is capped and nothing decays",
        "Flags email opens as structurally unreliable because mail-privacy proxies fire them automatically, and scores them 0 or near-0 with a hard cap on the email category",
        "Flags that fit and engagement are blended into a single number and must be split into two separate scores kept separate end to end",
        "States the routing consequence of blending: high-fit/low-engagement is a nurture case and low-fit/high-engagement is noise, and one number cannot distinguish them",
        "Flags the absence of exclusions and names at least three record types that must be excluded outright, drawn from competitor domains, students or .edu addresses, existing customers, unsubscribes, hard bounces, or internal test accounts",
        "States that absolute disqualifiers are excluded outright rather than given negative points, because a competitor browsing the pricing page daily out-engages any fixed penalty",
        "Flags the threshold of 100 as chosen by feel and requires re-setting it from sales capacity",
        "Recommends adding decay to engagement signals only, not to the fit signal",
        "Recommends re-deriving weights from 12-24 months of closed-won and closed-lost records rather than from intuition",
        "Names an accountable owner, a version or change log, and a v2 review date as part of the fix",
        "Presents the fixes in an order justified by value per unit of effort rather than in the order the faults were listed"
      ]
    },
    {
      "id": 3,
      "prompt": "Quick one. We run Tessellate, a $14/seat/month collaboration app, about 40,000 free signups a month, self-serve with a 3-person sales-assist team that chases accounts likely to buy 25+ seats. I want to flag the hottest free accounts for that team. Product already tracks: completed onboarding checklist, created 5+ boards, invited a teammate, hit the 3-project free cap, connected Slack or Drive, and days-active-in-a-row. Sales-assist can handle maybe 30 accounts a week. We rescore the whole base every Sunday night. What should the model look like?",
      "expected_output": "A PQL model that keeps a thin fit gate alongside usage, weights engagement-heavy, rescores and triggers at product speed, and sets the threshold from the 30-per-week sales-assist capacity.",
      "files": [],
      "expectations": [
        "Keeps a fit layer in the model rather than scoring on product usage alone",
        "Justifies the fit gate by the classic PQL failure: high-usage, low-fit users such as students or consumers on a business product being routed to sales",
        "Proposes a thin fit gate such as work-email domain or company size band rather than a full firmographic fit model",
        "Sets the fit/engagement weight split in the 30-40 fit / 60-70 engagement range for PLG",
        "Scores 'invited a teammate' among the strongest signals and exempts it from decay",
        "Replaces the weekly Sunday rescore with at least daily rescoring",
        "States that crossing the threshold should trigger instantly rather than waiting for the next batch run",
        "Sets decay horizons shorter than a B2B sales-led model because the decision window is days rather than quarters",
        "Sets the PQL threshold from the sales-assist capacity of about 30 accounts a week rather than from a percentile or a round number",
        "States that one user's behavior can qualify the lead because there is no buying committee to wait for",
        "Assigns exclusions covering existing paid workspaces plus .edu or competitor domains",
        "Recommends a recalibration cadence faster than quarterly, monthly or better, because the volume supports it"
      ]
    },
    {
      "id": 4,
      "prompt": "Norbeck Cloud here. The scoring matrix is built and backtested, lift looks fine. Now I just need to pick the MQL cutoff. My plan is to take the top 10% of the database by score and call that the bar - that's roughly 900 leads a month. We have 7 AEs. The top-decile approach seems to be the standard everyone uses. Separately, our VP Sales wants demo requests to clear the same bar as everything else so we stay consistent. Sound right?",
      "expected_output": "Threshold set from rep capacity first with the percentile demoted to a sanity check, the 900/month volume contrasted against 7 reps, and hand-raisers exempted from the bar.",
      "files": [],
      "expectations": [
        "Rejects the percentile as the primary threshold-setting method and demotes it to a sanity check",
        "States the three threshold methods ranked by value per unit of effort: sales capacity > conversion band > percentile sanity check",
        "Sets the threshold from sales capacity first, using reps multiplied by leads worked per rep per week",
        "Derives or asks for the capacity figure from the 7 AEs and a per-rep weekly working rate, and contrasts it with the 900-per-month volume the user proposed",
        "States that a volume far above rep capacity recreates the triage problem downstream instead of solving it",
        "Explains that the top 10% of a database is a statement about the database, not evidence those leads convert",
        "States that the percentile check should land near the top 15-20% of the active database and that a figure far outside that band usually signals a weighting problem rather than an unusual business",
        "Notes that on a 100-point scale most working models place the threshold between 60 and 80",
        "Rejects the VP Sales request and states that high-intent hand-raisers such as demo requests and contact-sales forms bypass the threshold entirely, qualifying on the action whatever the score",
        "Names the conversion band as the method to promote once roughly 60 days of scored history or a clean backtest exists",
        "Defines tiers where each tier maps to a distinct play rather than shipping a single MQL cutoff alone"
      ]
    },
    {
      "id": 5,
      "prompt": "Our new CRO came from a company that used an AI lead scoring vendor and wants us on predictive by end of quarter. Context: Ambrose Ledger, B2B fintech ops tooling, we closed 61 deals and lost 94 over the last 18 months, roughly 700 new leads a month. Two vendors quoted us; one says their documented minimum is 40 qualified plus 40 disqualified leads, so we clear it comfortably. Our rules model is 14 months old and reps mostly ignore it. Which vendor should we go with?",
      "expected_output": "A refusal to select a vendor: 155 closed outcomes sit below the predictive data floor, the pure-predictive rung is deleted rather than ranked last, and the rep-adoption problem is diagnosed as explainability and buy-in.",
      "files": [],
      "expectations": [
        "States that the user's roughly 155 closed outcomes sit below the data floor for predictive scoring of about 500-1,000 clean closed outcomes",
        "Treats the vendor's 40-plus-40 minimum as the permissive end of a published range rather than as evidence the user qualifies",
        "States that a model trained on this volume is a wrong option rather than a slower one",
        "Deletes the pure-predictive rung from the plan rather than ranking it last",
        "Presents the approaches ranked by value per unit of effort: rules-based > hybrid > pure predictive",
        "Emits at least one other ordering axis on which pure predictive leads, and qualifies that lead as holding only above the data floor",
        "Recommends repairing the existing rules model instead of selecting a vendor",
        "Names the black-box trust problem: when the model cannot explain a score, reps stop acting on it regardless of accuracy",
        "Attributes the reps-ignore-it symptom to explainability and missing sales buy-in rather than to model accuracy",
        "Describes the hybrid pattern as rules handling hard disqualification with a predictive layer ranking the remainder",
        "States that hard disqualifiers stay in rules even once a predictive layer is added",
        "Does not recommend either of the two vendors"
      ]
    },
    {
      "id": 6,
      "prompt": "Kestrel Grid, enterprise infrastructure security. Median sales cycle is 9 months and the buying committee runs 6 to 8 people. Our scoring model has no decay at all, so contacts from two years ago still look warm. A consultant told us to decay everything to zero after 30 days of inactivity and to reset any score above 50 back down to 50 after 45 idle days. Our marketing automation platform does store per-event timestamps. Our CTO also asked whether job titles should decay, since people move around a lot. Can you write us the decay rules?",
      "expected_output": "Decay applied to engagement only, horizon scaled to the 9-month cycle rather than 30 days, tiered brackets recommended over the consultant's threshold reset, and hand-raises left undecayed.",
      "files": [],
      "expectations": [
        "Applies decay to engagement signals only and states that fit signals do not decay",
        "Answers the job-title question with a clear no, on the grounds that a job title does not go stale on its own",
        "Rejects the consultant's 30-day decay-to-zero horizon as over-aggressive for a 9-month median cycle",
        "States that engagement should decay to zero at roughly 1-2x the median sales cycle, giving a horizon on the order of 9-18 months for this user",
        "Warns that over-aggressive decay silently kills slow-burn enterprise deals because the model forgets the buyer before the committee finishes deciding",
        "Keeps explicit hand-raises such as demo requests undecayed until a rep works them",
        "Presents the decay mechanisms ranked by value per unit of effort: tiered brackets > percentage per period > threshold reset",
        "Recommends tiered brackets as the default mechanism",
        "Rejects the consultant's threshold reset as the chosen mechanism and explains it cannot distinguish a 60-day-old pricing visit from a 6-day-old one",
        "Notes that because the platform stores per-event timestamps, both tiered brackets and percentage-per-period decay remain available",
        "Gives a concrete bracket schedule with at least three age bands and the share of points each retains",
        "States that the decay rate itself is set by sales-cycle length rather than by a value-per-effort ranking"
      ]
    },
    {
      "id": 7,
      "prompt": "Backtest results on our v1 model at Salvia Freight Systems, and I want a sanity check before we ship Monday. We retro-scored 9 months of leads. Top decile converts to opportunity at 7.1%, all-leads baseline is 5.4%. 22% of our closed-won deals scored under the threshold - digging in, most came through our partner referral program or the two trade shows we run. Also band 6 converts better than band 7, which is odd. Marketing wants to drop the threshold from 70 to 55 so more of those referral deals get picked up. Green light?",
      "expected_output": "A refusal to ship: lift computes to about 1.3x against a 2x floor, the 22% false-negative rate is fixed by adding the missing buying paths rather than lowering the bar, and the band inversion is found before launch.",
      "files": [],
      "expectations": [
        "Computes the lift as roughly 1.3x from 7.1% over 5.4% and states it fails the 2x pass floor",
        "Refuses to ship the model in its current state",
        "Rejects lowering the threshold from 70 to 55 as the fix for the false negatives",
        "States that a 22% false-negative rate is above the roughly 10-15% warning level and means the model ignores a real buying path",
        "Prescribes adding signals for the partner-referral and event buying paths instead of lowering the bar",
        "Identifies the band 6 over band 7 inversion as a symptom of a mis-weighted signal that must be found before shipping",
        "States that conversion by score band should be roughly monotonic",
        "Requires re-running the backtest after the signal changes rather than shipping and watching live",
        "Notes that the retro-score must exclude leads too young to have resolved, younger than one median sales cycle",
        "Recommends a spot-check holdout on named record types drawn from a recently accepted MQL, a recent closed-won, a recent closed-lost, a competitor contact, a high-engagement poor-fit contact, and an existing customer",
        "States that the competitor contact and the existing customer must come out excluded, not merely scored low",
        "Keeps false-negative rate as a standing KPI after launch rather than a one-time backtest figure"
      ]
    },
    {
      "id": 8,
      "prompt": "Need help framing something for our QBR. Our CMO's board metric is MQL volume and we're up 41% year on year - 2,100 MQLs last quarter against 1,490. Problem is sourced pipeline is flat and MQL-to-opportunity slid from 9% to 5.5%. Demand gen added three content syndication sources and lowered the score needed for a gated ebook download since last year. Sales acceptance is running around 38%. How do I present this?",
      "expected_output": "The pattern named as score inflation inside a Goodhart dynamic, with the reported metric changed to MQL-to-opportunity and pipeline contribution, and a weight/cap/threshold fix requiring sales re-sign-off.",
      "files": [],
      "expectations": [
        "Names the Goodhart dynamic: MQL volume became the target, so the qualification bar drifted down",
        "Identifies score inflation as the fault behind rising MQL volume alongside falling MQL-to-opportunity conversion",
        "Recommends reporting MQL-to-opportunity conversion and pipeline contribution upward instead of raw MQL volume",
        "States that the reporting change costs near-zero in hours but spends political capital with whoever set the MQL target",
        "Flags the 38% sales acceptance rate as far below the common 60% practitioner target",
        "Labels the 60% acceptance floor as a practitioner convention rather than a researched constant",
        "Notes that rejection above roughly 25-30% means the model lacks real sales buy-in",
        "Prescribes re-deriving weights from won/lost data, adding category caps, and raising the threshold as the inflation fix",
        "Requires fresh sales sign-off on any raised threshold",
        "Recommends monitoring the share of the active database sitting above threshold as a standing inflation alarm",
        "Treats content syndication as a signal category needing a cap so that channel alone cannot push a lead over the threshold"
      ]
    },
    {
      "id": 9,
      "prompt": "Working on weights for Brindle Robotics' scoring model. I exported all 240 closed-won and 380 closed-lost opportunities from the last 20 months. Interesting finding: 96% of won accounts have the 'annual revenue' field populated versus 44% of lost ones, so revenue band looks like our single strongest predictor - I was going to give it 25 of the 50 fit points. Webinar attendance shows up on 31% of wins and 29% of losses. Our CEO also read that responding within 5 minutes makes you 100x more likely to qualify a lead and wants that as the headline of the deck. And he wants to benchmark our MQL-to-SQL against the 45% industry number. Thoughts?",
      "expected_output": "The revenue-band finding rejected as the fields-populated-at-close artifact, webinar attendance demoted as measured noise, and both of the CEO's cited numbers flagged with their reliability caveats.",
      "files": [],
      "expectations": [
        "Flags the annual-revenue finding as the fields-populated-at-close trap: the field is complete on won records because sales filled it in shortly before closing",
        "Requires checking field history to confirm the value existed on the record at scoring time, rather than checking field completeness",
        "Rejects assigning 25 of the 50 fit points to revenue band on the current evidence",
        "Identifies webinar attendance at 31% of wins versus 29% of losses as roughly baseline, meaning noise that would be carrying points",
        "Demotes webinar attendance to a low or zero point value and says the measurement is what justifies the demotion, rather than deleting it silently",
        "Describes the weighting method as comparing each signal's frequency in the won group against the lost group and scaling points proportionally to the observed delta",
        "Rounds resulting point values to figures sales can reason about",
        "Flags the 5-minutes-100x claim as a single-vendor phone-outreach study from 2007 rather than a randomized trial",
        "Recommends treating fast response as an aggressive target for hand-raisers rather than a universal law",
        "Rejects benchmarking MQL-to-SQL against the industry number and states the published 13%-45% spread is definitional variance rather than performance variance",
        "Recommends comparing scored against unscored cohorts inside the company's own funnel instead of cross-company benchmarking",
        "States that every point value in the delivered spec carries its won-versus-lost evidence rather than standing as a bare number"
      ]
    }
  ],
  "trigger_queries": [
    { "query": "our lead scoring model is a mess, help me rebuild it", "should_trigger": true },
    { "query": "how do I set an MQL threshold", "should_trigger": true },
    { "query": "sales keeps rejecting the MQLs we hand them", "should_trigger": true },
    { "query": "too many junk MQLs in the queue every week", "should_trigger": true },
    { "query": "should fit and engagement be one score or two separate ones", "should_trigger": true },
    { "query": "how many points is a demo request worth", "should_trigger": true },
    { "query": "design a PQL model for our free tier", "should_trigger": true },
    { "query": "how fast should lead points decay", "should_trigger": true },
    { "query": "we give email opens +2, is that a problem", "should_trigger": true },
    { "query": "competitors keep showing up at the top of our hot list", "should_trigger": true },
    { "query": "our scores stopped predicting anything about six months ago", "should_trigger": true },
    { "query": "what number should we set the qualification bar at", "should_trigger": true },
    { "query": "how do I prove the scoring model actually works before we ship it", "should_trigger": true },
    { "query": "backtest our model against last year's closed won deals", "should_trigger": true },
    { "query": "students keep clearing our qualification bar", "should_trigger": true },
    { "query": "we have 300 leads above the bar and 6 reps to work them", "should_trigger": true },
    { "query": "do we have enough closed deals for predictive lead scoring", "should_trigger": true },
    { "query": "marketing says MQLs are up but nothing is closing", "should_trigger": true },
    { "query": "we closed deals that never hit MQL, what does that mean", "should_trigger": true },
    { "query": "how often should we recalibrate the qualification model", "should_trigger": true },
    { "query": "who needs to sign off on our qualification criteria", "should_trigger": true },
    { "query": "our newsletter subscribers outrank people asking for a demo", "should_trigger": true },
    { "query": "can you build me a scoring matrix for inbound", "should_trigger": true },
    { "query": "what should the penalty be for a personal gmail address on a B2B form", "should_trigger": true },
    { "query": "how do I stop one channel from pushing leads over the bar by itself", "should_trigger": true },
    { "query": "we bought intent data, how do I fold it into the model", "should_trigger": true },
    { "query": "rank our inbound so reps work the best ones first", "should_trigger": true },
    { "query": "which of these signups should a rep call today", "should_trigger": true },
    { "query": "help me decide who's actually worth calling out of 4000 signups", "should_trigger": true },
    { "query": "is a VP job title worth more than two pricing page visits", "should_trigger": true },
    { "query": "the free trial users who use the product the most are all students", "should_trigger": true },
    { "query": "this lead shows 88 but it's obviously garbage, what's broken", "should_trigger": true },
    { "query": "set up tiers so each band gets a different play", "should_trigger": true },
    { "query": "how do I know if our threshold is set too low", "should_trigger": true },
    { "query": "we're launching in the EU, anything to worry about with behavioural scoring", "should_trigger": true },
    { "query": "rebuild the inherited marketing engagement score nobody understands", "should_trigger": true },
    { "query": "should hand-raisers have to clear the same bar as everyone else", "should_trigger": true },
    { "query": "what fit to engagement weight split works for enterprise ACV", "should_trigger": true },
    { "query": "our enrichment only covers half the database, can we still score", "should_trigger": true },
    { "query": "reps don't trust the number the CRM shows on a lead", "should_trigger": true },
    { "query": "we want to move from rules to machine learning for qualification", "should_trigger": true },
    { "query": "should we use a letter grade plus a number like A1 and B3", "should_trigger": true },
    { "query": "MQL volume is our marketing target this year, is that a problem", "should_trigger": true },
    { "query": "our qualification bar hasn't been touched in three years", "should_trigger": true },
    { "query": "how do I tell a cold but perfect prospect apart from a noisy tire-kicker", "should_trigger": true },
    { "query": "put together a spec my marketing ops team can actually implement", "should_trigger": true },
    { "query": "what KPIs tell me our qualification model is healthy", "should_trigger": true },
    { "query": "we're a $12 a seat product, how do we decide who sales calls", "should_trigger": true },
    { "query": "job seekers browsing our careers page are coming through as qualified", "should_trigger": true },
    { "query": "round robin is handing one rep all the good leads", "should_trigger": false },
    { "query": "assign inbound demo requests to the right AE by territory", "should_trigger": false },
    { "query": "leads are sitting unworked in the CRM queue for days", "should_trigger": false },
    { "query": "design an SLA escalation path when a rep doesn't touch a lead", "should_trigger": false },
    { "query": "how do we match inbound leads to existing accounts", "should_trigger": false },
    { "query": "our green accounts keep churning, the health score is wrong", "should_trigger": false },
    { "query": "build a composite account health score with bands", "should_trigger": false },
    { "query": "which usage signals actually predict churn for us", "should_trigger": false },
    { "query": "champion departure keeps catching us by surprise", "should_trigger": false },
    { "query": "build a propensity model for which existing customers will upgrade", "should_trigger": false },
    { "query": "our pipeline stages don't mean anything anymore", "should_trigger": false },
    { "query": "audit our stage exit criteria against buyer-verifiable milestones", "should_trigger": false },
    { "query": "design a funnel stage set from scratch with conversion assumptions", "should_trigger": false },
    { "query": "define the marketing to sales handoff in our funnel model", "should_trigger": false },
    { "query": "why do our deals keep slipping out of the quarter", "should_trigger": false },
    { "query": "our commit number is never right, what's wrong with the forecast", "should_trigger": false },
    { "query": "clean up the stale deals in the pipeline before the QBR", "should_trigger": false },
    { "query": "flag deals whose close date has been pushed more than twice", "should_trigger": false },
    { "query": "who owns the industry field in our CRM and how often is it refreshed", "should_trigger": false },
    { "query": "we have custom field sprawl and nobody knows which system wins", "should_trigger": false },
    { "query": "finance and sales report two different ARR numbers", "should_trigger": false },
    { "query": "where is our metric definition supposed to live", "should_trigger": false },
    { "query": "build a discount approval matrix for non-standard deals", "should_trigger": false },
    { "query": "set a margin floor policy for our deal desk", "should_trigger": false },
    { "query": "which RevOps newsletters and podcasts should I subscribe to", "should_trigger": false },
    { "query": "write the scorecard for our first RevOps hire", "should_trigger": false },
    { "query": "design a take-home exercise and rubric for a sales ops candidate", "should_trigger": false },
    { "query": "am I ready to step up to a RevOps manager role", "should_trigger": false },
    { "query": "design a KPI tree for the board with guardrail metrics", "should_trigger": false },
    { "query": "our board deck reads like a data dump, restructure it", "should_trigger": false },
    { "query": "the CSM starts from zero every time a deal closes", "should_trigger": false },
    { "query": "where are we silently losing deals in this funnel", "should_trigger": false },
    { "query": "size the recoverable revenue from our unworked leads", "should_trigger": false },
    { "query": "we have way too many GTM tools, what do we cut", "should_trigger": false },
    { "query": "which revops skill do I need for this project", "should_trigger": false },
    { "query": "define our ICP tiers for the enterprise segment", "should_trigger": false },
    { "query": "which data enrichment vendor should we buy for the CRM", "should_trigger": false },
    { "query": "build a credit scoring model for loan applicants", "should_trigger": false },
    { "query": "score candidate resumes against this job description", "should_trigger": false },
    { "query": "design a scoring rubric for our hackathon judges", "should_trigger": false },
    { "query": "how should we score RFP vendor responses", "should_trigger": false },
    { "query": "set up NPS and a customer satisfaction score", "should_trigger": false },
    { "query": "score our SEO pages by traffic opportunity", "should_trigger": false },
    { "query": "rank our open support tickets by priority", "should_trigger": false },
    { "query": "write a cold email sequence for our inbound contacts", "should_trigger": false },
    { "query": "build a lead magnet to capture more emails", "should_trigger": false },
    { "query": "what's a good email open rate benchmark for B2B SaaS", "should_trigger": false },
    { "query": "write the compensation plan for our SDR team", "should_trigger": false },
    { "query": "score inbound partner applications for our reseller program", "should_trigger": false }
  ]
}
references/examples.md
# Worked Examples

Two passing specs and one broken model. The point values are illustrative shapes, not universal constants - every real spec re-derives them from the user's own won/lost data.

## Example 1 - Mid-market B2B, sales-led

```
SCORING MODEL SPEC - B2B SaaS, ACV ~$15K, 2026-03, v1
Motion       : sales-led mid-market - 50/50 fit/engagement, 100-pt scale (50+50)
Fit (50)     : primary industry +15 (2.1x won-vs-lost) | 51-1,000 employees +15 (1.8x)
               | manager-to-VP title in buying dept +12 | complementary tool +8 (coverage 71%)
Engagement   : demo/trial request +30, no decay until worked | pricing 2+ visits +20, -5/wk
(50, capped) : integration-docs view +15, -5/2wk (3.1x baseline - the surprise winner)
               | webinar attended +5, -5/mo (~1x baseline; kept low deliberately)
               | all email engagement capped at 10 total
Exclusions   : competitor domains, .edu, existing customers, unsubscribed, hard bounces
Negative pts : personal email -10 | careers-page-only -30 | consultant title -10
Threshold    : 65 -> projects ~35 MQL/wk against 8 reps x 5/wk = 40 capacity
Tiers        : T1 hand-raiser or 85+ -> fast human follow-up | T2 65-84 -> queue
               | T3 fit >=30, eng <15 -> nurture | T4 rest -> lifecycle only
Backtest     : 12-mo retro-score: top band 21% -> opportunity vs 6% baseline = 3.5x lift;
               9% of wins scored <65 (all event-sourced -> event-attendance signal added)
Guardrails   : acceptance floor 60% | alarm if >12% of active database sits above 65
Governance   : owner MOps lead | sign-off VP Sales, 2026-03-10 | v2 review 2026-05-10
```

Why it passes:

- Every weight cites a won/lost delta.
- A folklore signal (webinars) was measured, found to be noise, and demoted instead of deleted silently.
- The threshold comes from capacity math, then survives the conversion-band check.
- The false-negative finding produced a signal fix, not a lower bar.
- Lift 3.5x clears the 2x floor.

## Example 2 - PLG collaboration tool (mechanics identical, inputs and speed differ)

```
SCORING MODEL SPEC - PLG collaboration SaaS, $12/seat/mo, 2026-03, v1
Motion       : PLG - 30/70 fit/engagement; engagement fed by product events;
               rescored daily, threshold trigger fires instantly
Fit (30)     : work-email domain +10 | company 10-500 +10 (enrichment coverage 64%;
               unknown scores 0, never negative) | manager+ title +10
Engagement   : activation milestone +20 | core feature 3+ uses +20 | teammate invited
(70)         : +20, no decay | usage-limit hit +15 | pricing view +10, -5/wk
               | inactivity decay: engagement -25% per 14 idle days
Exclusions   : existing paid workspaces, .edu domains, competitor domains
Threshold    : PQL = 60 - one user's behavior qualifies; no committee to wait for
Tiers        : PQL -> sales-assist outreach | 40-59 -> in-product upgrade nudges
               | <40 -> lifecycle email only
Backtest     : 6-mo cohort: PQL band converted free->paid at 4.6x the all-signup
               baseline; false negatives 6%
Guardrails   : acceptance floor 60% for sales-assist queue | fit gate retained so
               high-usage consumer/edu users never reach sales
Governance   : owner growth ops | sign-off sales-assist lead | recalibration monthly
               (volume supports it) | v2 review +60 days
```

Why it passes:

- The fit layer is thin but present - the classic PQL failure (routing high-usage, low-fit users) is blocked by design.
- Decay and rescoring run at product speed.
- Monthly recalibration matches the data volume.

## Example 3 - Negative: a plausible-looking broken model

```
"MARKETING ENGAGEMENT SCORE" - v3 (inherited, undocumented)
Single blended score, no ceiling: email open +2 | email click +5 | any page view +3
| whitepaper +10 | webinar registration +10 | blog visit +3 | demo request +10
Fit          : title contains "manager" +10
Exclusions   : none    Decay: none    Caps: none
Threshold    : 100 ("felt right when we launched")
Validation   : never backtested    Owner: unclear    Change log: none
Status       : sales quietly stopped opening MQL tasks in Q2
```

Every flaw, decomposed:

- **Unbounded accumulation, no decay, no caps.** Six months of newsletter opens (+2 each) outscores a demo request (+10). The hottest signal in the model is worth three email clicks.
- **Vanity signals carry the model.** Opens and generic page views - the two least reliable signals, opens being auto-fired by mail-privacy proxies - contribute most crossings of the threshold.
- **Fit and engagement blended into one number**, and fit is one title keyword. The model cannot distinguish a cold-but-perfect prospect from an enthusiastic student - and has no exclusions, so competitors and students routinely cross 100.
- **Threshold chosen by feel**, never checked against sales capacity or a conversion band, never backtested. Volume swamped the reps; acceptance collapsed; the bar was then quietly lowered to keep the MQL chart up - the textbook Goodhart spiral.
- **No owner, no version log, no review date.** Three undocumented revisions in, nobody can say why any weight is what it is, so nobody can fix it - only distrust it.

The rebuild path is the main workflow, not patching:

- Split the axes.
- Re-derive weights from 12-24 months of won/lost outcomes.
- Move disqualifiers to exclusions.
- Add decay and category caps.
- Set the threshold from capacity.
- Backtest to a >= 2x lift.
- Put a name, a version, and a v2 date on the result.
references/signal-taxonomy.md
# Signal Taxonomy

Fit and engagement stay separate scores end to end. Blending them into one number loses the routing distinction that justifies scoring at all: high-fit/low-engagement is a cold prospect worth nurturing; low-fit/high-engagement is noise worth deprioritizing. Fit without intent is a cold prospect; intent without fit is noise.

## Fit signals - can they buy

Explicit data from forms and enrichment. Stable, so they never decay. Their usefulness is bounded by enrichment coverage, not by cleverness of the point values.

| Bucket             | Signals                                                                         | Notes                                                                                                              |
| ------------------ | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| Firmographic (B2B) | Industry, employee count, revenue band, geography, funding stage                | Strongest single gate at high ACV                                                                                  |
| Demographic        | Job title, seniority, department, buying-committee role                         | Normalize titles before scoring; "Head of" != VP everywhere                                                        |
| Technographic      | Uses complementary tool; uses a competitor; uses the tool this product replaces | Competitor usage is positive fit (they understand the category) but often an exclusion for outreach - decide which |
| B2C fit            | Age band, location, purchase history, account tenure                            | Same axis, consumer-grade inputs; identical mechanics                                                              |

## Engagement signals - are they about to buy

Behavioral data, first-party. Rank by proximity to a purchase decision, not by ease of tracking.

| Tier                       | Signals                                                                                                                                              | Typical points (100-pt scale) |
| -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------- |
| High intent                | Demo request, contact-sales form, trial start, pricing page 2+ visits, integration/API docs, ROI calculator                                          | +15 to +30                    |
| Medium                     | Webinar attendance, comparison page, case study 2+, multiple sessions in a week                                                                      | +8 to +15                     |
| Low                        | Single blog visit, newsletter click, single email click                                                                                              | +1 to +5                      |
| Vanity - score 0 or near-0 | Email opens (structurally unreliable since mail-privacy proxies fire them automatically), single content download, homepage-only visit, social likes | 0-2, hard-capped              |
| Negative                   | Careers-page-only visitor (job seeker)                                                                                                               | -30                           |

A diligent newsletter reader - or an automated inbox opening everything - must never be able to accumulate enough vanity points to look like a buyer. Hard-cap each low-tier category's total contribution so accumulation cannot substitute for intent.

## Third-party intent data

Real but noisy; treat as a layer, never a qualifier on its own.

- It is account-level: it identifies a company researching a topic, not a person ready to talk.
- Independent precision testing puts leading providers' topic-match accuracy in the low-to-high 80% range - a "surge" is a probability, not an intent; plenty of surges are an analyst building a market map.
- Use it to break ties and prioritize outreach among already-fit accounts, or to trigger monitoring. Never let a third-party signal alone push a lead across the MQL threshold.
- Privacy: third-party behavioral data carries the heaviest consent burden (GDPR profiling, browser third-party-cookie blocking already covers roughly a third of traffic).

## Exclusions vs negative points

Exclude absolute disqualifiers outright rather than assigning negative points - negative scoring rarely nets out cleanly against genuine engagement. A competitor employee browsing the pricing page daily will out-engage any fixed penalty.

**Exclusion list (suppressed from scoring and from the sales queue):** competitor email domains, students and .edu addresses (unless education is the market), existing customers, unsubscribes and spam complaints, hard bounces, internal test accounts.

**Negative points (soft demotions, not disqualifiers):** personal email address in a B2B motion (-10; relax for SMB), consultant/agency title (-10, may be evaluating for a client), IC-only contact at enterprise ACV (-15), invalid phone (-10).

## Decay

Engagement decays; fit does not - a job title does not go stale on its own. Without decay, every contact accumulates points indefinitely and stale contacts end up looking as hot as active ones. That decay must exist is not a choice; which mechanism delivers it is. Ranked by value per unit of effort, effort being admin config, engineering to build and run a recompute, and how much of the model reps can still explain afterwards:

- efficiency: tiered brackets > percentage per period > threshold reset
- value (stale points removed without forgetting a live buyer): percentage per period > tiered brackets > threshold reset
- effort: percentage per period > tiered brackets > threshold reset

| Mechanism             | Worked example                                                                                       | Effort                                                                                               |
| --------------------- | ---------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| Tiered brackets       | 0-30 days = 100% of points, 31-60 = 75%, 61-90 = 50%, older = context only (40 pts -> 30 -> 20 -> 0) | An hour - most scoring tools ship step rules natively, and a rep can read the bracket off the record |
| Percentage per period | An event worth 10 points at 50%/month decay contributes 5 after one month, ~0 after two              | A week to build a scheduled recompute over per-event history, then a standing job                    |
| Threshold reset       | Any score above 50 resets to 50 after 45+ days of inactivity                                         | Near-zero - one rule on the composite score, no per-event history needed                             |

Default to tiered brackets: near the smooth version's accuracy at admin-config effort, and auditable by the rep looking at the record. The threshold reset is last on efficiency despite being the cheapest, which is the whole point of ranking by ratio - it truncates a dormant record's score and never distinguishes a 60-day-old pricing visit from a 6-day-old one.

What this order starves: percentage-per-period decay, first on value and first on effort. Promote it where engagement carries the majority of the score and the decision window is days, not quarters (PLG and B2C), because bracket edges are too coarse there. Where the platform cannot store per-event timestamps, delete both brackets and percentage decay and say so - the threshold reset is then the only mechanism available, not the recommended one.

How fast to decay each signal is genuinely contested, and this one gets no ranking: it would be false precision, because the answer is set by the sales cycle rather than by any value/effort ratio. The rate that is correct on a 14-day PLG cycle is the rate that kills a nine-month enterprise deal. One camp decays high-intent actions fastest (pricing visits to zero in 15-30 days - heat matters only while fresh) and low-intent slowly (60-90 days); another never decays explicit hand-raises (a demo request stays actionable until a rep works it).

Choose by sales-cycle length: decay engagement to zero at roughly 1-2x the median cycle, and keep hand-raises undecayed until worked. Over-aggressive decay silently kills slow-burn enterprise deals - the model forgets the buyer before the committee finishes deciding.

## PLG and B2C product-usage signals

Same two-axis structure; the engagement axis is fed by product analytics instead of marketing touches. Everything above about caps, exclusions, and decay applies unchanged.

| Signal                                                                         | Typical points | Why it predicts                                                                             |
| ------------------------------------------------------------------------------ | -------------- | ------------------------------------------------------------------------------------------- |
| Activation milestone completed                                                 | +20            | Experienced the core value ("aha moment")                                                   |
| Core feature used 3+ times                                                     | +20            | Habit forming, not a tourist                                                                |
| Invited a teammate                                                             | +20, no decay  | Expansion inside the account; strongest PQL signal                                          |
| Hit a usage/plan limit                                                         | +15            | Concrete upgrade pressure                                                                   |
| Connected an integration                                                       | +15            | Workflow lock-in                                                                            |
| Daily-active streak                                                            | +10 to +20     | Sustained engagement                                                                        |
| B2C transactional: cart abandonment, repeat product-page visits, wishlist adds | trigger-grade  | Propensity signals that expire in hours - act on them near-real-time, not in a weekly batch |

Two PLG/B2C-specific rules:

- **Keep the fit layer even when usage dominates.** The classic PQL mistake routes high-usage, low-fit free users (students, consumers on a business product) to sales. A thin fit gate - work-email domain, company size band - filters them at near-zero cost.
- **Speed replaces committee dynamics.** With no buying group, one user's behavior qualifies the lead, and the decision window is days or minutes. Rescore at least daily, trigger instantly on threshold, and shorten every decay horizon accordingly.
references/validation-and-kpis.md
# Validation, KPIs, Recalibration, Governance

If the model does not beat no-scoring at predicting opportunity conversion, it is decoration. Everything in this file exists to test and keep testing that claim.

## Backtest method

1. Retro-score 6-12 months of historical leads whose outcome is known; exclude leads too young to have resolved (younger than one median sales cycle).
2. Bucket into score bands - deciles on the composite, or grid cells if using the two-axis grid.
3. Compute opportunity conversion per band. Expect a roughly monotonic curve; an inversion (a lower band converting better) points at a mis-weighted signal - find it before shipping.
4. **Lift** = top-band conversion / overall unscored baseline conversion. Pass floor: **>= 2x** (this skill's practical floor - the published practitioner rule only demands beating the baseline at all, but a thin edge on clean historical data will not survive live noise).
5. **False negatives**: share of closed-won deals that scored below threshold. Above ~10-15% means the model ignores a real buying path (partner referrals, events, product-led entry) - add the path, don't lower the bar.
6. **Spot-check holdout**: manually score a handful of known records - a recently accepted MQL, a recent closed-won, a recent closed-lost, a competitor contact, a high-engagement/poor-fit contact, an existing customer. Each must land where a human would put it; the competitor and the customer must be excluded, not merely low.

## KPI set

| KPI                                        | Definition                                   | Healthy signal               | Reliability note                                                                                             |
| ------------------------------------------ | -------------------------------------------- | ---------------------------- | ------------------------------------------------------------------------------------------------------------ |
| Top-band lift vs unscored baseline         | Top band conversion / all-leads conversion   | >= 2x, stable                | The decisive test                                                                                            |
| Sales acceptance rate                      | Accepted MQLs / delivered MQLs               | >= 60%                       | Practitioner convention, not a researched constant; rejection above 25-30% means the model lacks real buy-in |
| MQL-to-SQL conversion                      | SQLs / MQLs                                  | Trend inside the same funnel | Never benchmark cross-company - definitions differ (see below)                                               |
| False-negative rate                        | Closed-won deals below threshold             | Trending toward 0            | The cheapest early-warning signal a model has drifted                                                        |
| Score distribution                         | % of active database above threshold         | Stable over time             | Upward creep is the inflation alarm                                                                          |
| MQL-to-opportunity / pipeline contribution | Opportunities and pipeline from scored leads | The number to report upward  | Reporting raw MQL volume instead is what triggers the Goodhart spiral                                        |

## Benchmark claims and their reliability

Quote these with their flags; several widely-cited numbers are folklore.

| Claim                                                          | Number                     | Reliability                                                                                                                                                                  |
| -------------------------------------------------------------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| "Respond in 5 minutes = 21x/100x more likely to qualify"       | 2007                       | Single vendor phone-outreach study from 2007, not a randomized trial; treat as an aggressive target for hand-raisers, not a universal law                                    |
| "Fewer than 1% of MQLs convert to deals"                       | <1%                        | Genuine named-analyst Forrester blog, but rests on an uncited "our research"; rhetorical ammunition, not a planning number                                                   |
| PQL conversion advantage                                       | 5x                         | Real published figure (OpenView 2020 survey, 150+ SaaS companies); the 5-10x variants in circulation are inflated restatements                                               |
| Lead scoring case study: leads to sales -52%, conversions +79% | 2012                       | Named published source (MarketingSherpa/Bersin) but quarter-over-quarter figures, routinely misquoted with qualifiers stripped                                               |
| MQL-to-SQL "benchmarks"                                        | 13%-45%                    | The spread is definitional variance, not performance variance - two companies can both be right at 13% and 42%. Compare scored vs unscored cohorts inside one funnel instead |
| Data floor for predictive scoring                              | ~500-1,000 closed outcomes | Convergent practitioner guidance plus vendor-documented minimums; reliable as an order of magnitude                                                                          |

## Recalibration cadence

Models are "one-time build, ongoing drift" by default - the cadence is what prevents it.

- **Weekly**: score distribution vs threshold (the inflation alarm).
- **Monthly**: conversion by band; per-signal driver check.
- **Quarterly**: full reanalysis on multi-quarter cohorts; re-derive weights from the updated won/lost data. PLG and high-volume B2C teams run this monthly - more data, faster drift.
- **v2 review**: booked at launch for 60 days out, before v1 ships. Non-negotiable; "we'll revisit it later" is how models freeze.
- **Off-cycle triggers**: ICP change, product launch or pricing change, score inflation (MQL volume rising while conversion falls), a cluster of false negatives, sales acceptance dropping below the agreed floor.

## Governance and versioning

- **One accountable owner**, one source of truth - marketing ops at SMB/mid-market, RevOps at enterprise scale. Shared ownership is how models rot unnoticed.
- **Document the matrix**: every signal with its points, decay, cap, rationale, and evidence; plus the list of every system feeding signals into the score.
- **Version every change** with date, author, reason, and expected effect; announce changes to sales before they land. A rep who cannot interrogate why a lead scored what it did stops acting on scores - keep the per-lead score breakdown visible in the CRM, and keep a feedback path for reps to flag scores that look wrong.
- **Consent (EU-facing teams)**: behavioral lead scoring is profiling under GDPR. Document the lawful basis, time-bound consent where used, and make consent status visible to sales before outreach.
references/weighting-and-thresholds.md
# Weighting and Thresholds

## Derive weights from outcomes, not opinions

The intuited signal list and the list that actually correlates with revenue are usually different. Calibrate against the company's own history, never against industry averages.

1. Pull closed-won and closed-lost records from the last 12-24 months (12 months of history is the practical minimum for evidence-based weights).
2. For each candidate signal, compute its frequency in the won group vs the lost group; rank signals by the delta. Where volume allows, compute per-signal conversion vs baseline - one documented recalibration found an integrations-doc download converting at 3.2x baseline while webinar attendance sat at ~1x, i.e. pure noise that had been carrying points for years.
3. Scale points proportionally to the observed delta, then round to values sales can reason about.
4. Trap - fields populated at close: won records often carry complete firmographic data only because sales filled it in right before closing. A signal is only usable if it existed on the record at scoring time; check field history, not field completeness.
5. Below ~12 months of usable history, start from the template split below, label every weight "assumed", and pull the first recalibration forward to 60 days.

## Fit/engagement split by motion

| Motion                          | Split (fit/engagement) | Why                                                                                |
| ------------------------------- | ---------------------- | ---------------------------------------------------------------------------------- |
| Enterprise sales-led (high ACV) | 60/40                  | Fit is the gate; a wrong-market lead engaging heavily is still a wrong-market lead |
| Mid-market hybrid               | 50/50                  | Balanced motion, balanced evidence                                                 |
| PLG / self-serve                | 30-40 / 60-70          | Product usage is the evidence of value; fit is a thin filter                       |
| B2C transactional               | Behavior-dominated     | Fit reduces to a demographic gate; propensity signals expire in hours              |

The split mechanics are identical for B2B and B2C - only the ratio and the signal sources change.

## Keep the two axes visible - the grid

Report fit and engagement as separate scores and route on the combination (Marketo-style: letter grade A-D for fit x number 1-4 for engagement, quartile cutoffs; MadKudu-style: Fit band x Likelihood-to-Buy band combined via a matrix into a grade). A 75 from an A-fit account at demo stage is not the same lead as a 75 from a D-fit account browsing the careers page.

Check win rate per grid cell empirically once volume accrues - B cells sometimes outperform A cells, and that finding is the point of the grid. It exists for calibration, not decoration.

## Category caps

Cap each signal category's total contribution so no single channel can cross the threshold alone - e.g. all email engagement capped at 10-15 points of a 100-point scale. Two silent inflation bugs to check for:

- Repeatable low-value events (opens, blog visits) with no cap compounding past the threshold.
- Multi-value properties scored once per matching value instead of once per record (a contact with three matching job functions triple-counts the points).

## Setting the threshold

Three methods, ranked by value per unit of effort - effort being analyst hours plus the scored history each method needs before it can run at all. They cross-check each other, so run them in this order; the order is the ranking.

- efficiency: sales capacity > conversion band > percentile sanity check
- value (a threshold that still holds three months after launch): conversion band > sales capacity > percentile sanity check
- effort: conversion band > sales capacity > percentile sanity check

Effort in magnitudes: the percentile check is minutes against the current database; the capacity math is an hour of arithmetic available before a single lead is scored; the conversion band needs ~60 days of live scored history or a full backtest first, then a week of analysis.

1. **Sales capacity first.** Threshold where projected weekly MQL/PQL volume roughly matches what reps can actually work (reps x leads-per-rep-per-week). This is the honest constraint most guides underweight - a statistically perfect threshold that produces 3x workable volume just recreates the triage problem downstream. It leads on efficiency because it needs no scored history at all and still buys the outcome that matters most on day one: a queue reps finish.
2. **Conversion band.** After ~60 days of scored history (or on backtest data), chart sales acceptance and opportunity conversion per score band; set the threshold where the curve breaks upward. Highest value of the three - it is the only method that reads actual conversion instead of proxying it - and highest effort, since it cannot run until the data exists.
3. **Percentile sanity check.** The threshold should land near the top 15-20% of the active database; on a 100-point scale most working models sit at 60-80. Far outside that range usually signals a weighting problem, not an unusual business. Cheapest and last on purpose: it catches a gross weighting error but can never set a threshold, because a database's top 20% is a statement about the database, not evidence those leads convert.

What this order starves: the conversion band, first on value and first on effort, so a day-one build never reaches it. Promote it the moment 60 days of scored history or a clean backtest exists, and treat the v2 review as its standing invitation. Where the Interview gave no analyst and no backtest, delete it from the plan rather than carrying it - set from capacity, sanity-check the percentile, and book the band check for when someone can run it.

**Initial calibration before any live data:** retro-score the last 6-12 months of closed-won deals; find the natural breakpoint separating wins from losses; set the threshold just below where ~80% of wins would have scored; validate against closed-lost - if many losses land above it, tighten the criteria rather than lowering the bar.

High-intent hand-raisers (demo request, contact-sales) bypass the threshold entirely - they qualify on the action, whatever the score.

## Tiers

Each tier maps to a distinct play or it is not a tier:

| Tier | Definition                          | Play                                |
| ---- | ----------------------------------- | ----------------------------------- |
| T1   | Hand-raiser, or top band (e.g. 85+) | Immediate human follow-up under SLA |
| T2   | Above threshold                     | Standard sales queue                |
| T3   | Good fit, low engagement            | Orchestrated nurture + ads          |
| T4   | Below both bars                     | Hold; lifecycle marketing only      |

Tier-to-rep assignment, SLAs, and escalation are routing concerns - hand the tiers to `mbfinotti/revops-skills@lead-routing`.

## Rules-based vs predictive

Ranked by value per unit of effort, effort being the data volume required up front, analyst and vendor onboarding time, and what it costs to keep the thing running:

- efficiency: rules-based > hybrid (rules disqualify, a predictive layer ranks the remainder) > pure predictive
- value (ranking accuracy on the leads sales actually works): pure predictive > hybrid > rules-based - but only above the data floor below; under it, predictive's value is unmeasurable and frequently negative
- effort: pure predictive > hybrid > rules-based

Effort in magnitudes: a rules model is a week of design plus an hour per tuning pass; the hybrid adds the vendor layer on top of rules that already work; a pure predictive build is a quarter before it scores anything - data floor, onboarding, validation - and then a standing job of retraining, drift monitoring, and explaining scores to reps.

Default to rules-based, and promote the hybrid once the data floor is genuinely cleared. Delete the predictive rungs entirely below that floor rather than ranking them last. A model trained on 80 outcomes is not a slower option, it is a wrong one, and a rung left at the bottom of the list comes back as a vendor evaluation.

What this order starves: pure predictive, first on value and first on effort. Promote it where lead volume is high enough that a few points of ranking accuracy outweigh the transparency reps lose - and even there, keep the hard disqualifiers in rules.

- **Rules-based** works from day one, is transparent, and sales can audit it - which is most of why they trust it. It degrades without tuning.
- **Predictive/ML needs volume.** Vendor minimums range from 40 qualified + 40 disqualified leads at the permissive end to 1,000+ leads with 120+ conversions at the strict end; the vendor-neutral floor is roughly 500-1,000 clean closed outcomes. Below that, a governed two-axis rules model delivers most of the benefit at a fraction of the data requirement.
- **The black-box trust problem is real:** when the model cannot explain a score, reps stop acting on it regardless of accuracy. Prefer the hybrid pattern - rules handle hard disqualification, a predictive layer ranks the remainder - and resist stacking manual overrides on a predictive base until nobody can audit the result.
SKILL.md
---
name: lead-scoring
description: Design, validate, and recalibrate a lead scoring model - fit and engagement signals, weighting, negative scoring and exclusions, decay, MQL/PQL threshold and tier setting, backtesting against closed-won/closed-lost outcomes, and governance; also diagnoses an existing broken model. Use whenever the user mentions lead scoring, an MQL threshold, fit vs engagement scoring, product-qualified leads, score decay, "score my leads", "our lead scores are wrong", "sales rejects our MQLs", or "too many junk MQLs" - even if they never say "scoring model". Covers B2B sales-led and B2C/PLG, rules-based design and predictive readiness. Do NOT use for assigning scored leads to reps - use mbfinotti/revops-skills@lead-routing instead.
license: MIT
metadata:
  author: Maya-Beth Finotti
  version: "1.2.2"
---

# Lead Scoring

Design, validate, and maintain a model that ranks leads by likelihood to become revenue. The practitioner standard splits scoring into two separate axes:

- **Fit** (can they buy): firmographic, demographic, technographic.
- **Engagement** (are they about to buy): behavioral, in-product.

This two-axis model appears under several names: Marketo's A-D x 1-4 grade-and-score grid, MadKudu's Customer Fit x Likelihood to Buy, and OpenView's product-qualified lead (PQL) in PLG.

Scoring fails far more often from organizational neglect - no sales sign-off, no recalibration, MQL count treated as the goal - than from bad math, so validation and governance are part of the design here, not an afterthought. The finished score is an input to assignment: routing, territories, and rep matching belong to `mbfinotti/revops-skills@lead-routing`. Defining the account-fit criteria this skill's fit axis weights - firmographic and technographic ICP tiers - belongs to `mbfinotti/sales-skills@sales-icp-definition`; this skill owns turning that definition, plus behavioral engagement, into a working score.

The capture-score-route-nurture mechanics are structurally identical for B2B and B2C/PLG. What genuinely differs between B2C/PLG and B2B:

- Dominant signal source is in-product behavior, not forms and content.
- No buying committee - one user's behavior can qualify.
- Cycles run in days, not quarters.

Because of these differences, decay, thresholds, rescoring frequency, and recalibration all run faster in B2C/PLG.

## Interview

Ask before designing anything. One question per message; offer multiple-choice answers when possible. Skip a question only when the user already gave the answer.

- Which motion? (a) sales-led B2B, (b) PLG / self-serve with sales assist, (c) B2C transactional.
- Roughly how many new leads per month, and how many closed-won and closed-lost outcomes exist from the last 12-24 months? (Decides rules-based vs predictive, and how much backtesting is possible.)
- What is actually on a lead record at creation time - enrichment coverage per fit field, form fields, product events? (A score is downstream of data quality.)
- Does a model already exist? If yes: redesign, or diagnose first? (Diagnosis path: see Diagnosing a Broken Model.)
- How many leads can sales actually work per week? (The honest threshold constraint.)
- Who signs off - which sales leader, and what acceptance-rate floor will they hold the model to?
- Any EU or consent constraints? Behavioral scoring is profiling under GDPR; the lawful basis must be documented.
- By what date must the model be scoring live, and what is waiting on it? (A date inside two weeks rules out both enrichment procurement and the conversion-band threshold method; a quarter of runway makes both affordable.)
- One-off win or compounding asset? (A single campaign list to work stops at v1 weights and a capacity threshold; a compounding mandate is what justifies enrichment, the full backtest, and the recalibration cadence that keep the model honest after launch.)
- What is the effort ceiling - analyst hours, CRM/marketing-automation admin access, a procurement and legal path for a data contract, and how much of sales' attention you can spend? (No analyst rules out re-deriving weights from won/lost data and the conversion-band method; no procurement path deletes the enrichment rung outright rather than ranking it last.)

Carry these last three answers into every ranking in this skill and in its references; they are the only inputs that reorder the defaults for this user. Every ranking here is a default, not a law: it shifts with context and with who executes it. Re-rank before recommending a rung, for example:

- An enrichment contract already in place collapses that rung's effort to a field mapping and promotes it to the top.
- A CRM whose fit fields arrive self-reported on the form at high coverage removes the coverage menu entirely.
- No analyst promotes every rung a CRM admin can ship alone and demotes every rung that needs a cohort analysis.

## Workflow

1. Run the Interview; collect every answer before designing.
2. Audit the data first. Measure enrichment coverage per fit field over the last 90 days of leads. A fit field below roughly 70% coverage means most of the database can never reach the threshold - the enrichment-ceiling failure. Three ways out, ranked by value per unit of effort, where effort is analyst hours, admin work, procurement, and reversibility:

   - efficiency: score unknowns 0 > enrich the field > drop the field
   - value (fit signal the model can actually act on): enrich the field > score unknowns 0 > drop the field
   - effort: enrich the field > drop the field > score unknowns 0
   - compliance cost: enrich the field > score unknowns 0 == drop the field

   Effort in magnitudes:

   - Scoring unknowns 0: near-zero, a scoring-config change reversible in a click.
   - Dropping the field: an hour, and the least reversible option in this menu - the signal stops being collected and the next recalibration cannot re-derive its weight without first rebuilding the history.
   - Enrichment: a procurement and contracting step before a single lead scores, then a standing per-record dependency and a coverage number to monitor forever.

   Default to scoring unknowns 0: it keeps every thin record scoreable on engagement and ships this week. Justify the compliance tie: neither the 0 rule nor the deletion introduces a new data source, so neither triggers a consent or provenance review. Enrichment does - third-party fit data is profiling input under GDPR, and its lawful basis and provenance have to be documented before it scores anyone.

   Name the bias that default buys: scoring unknowns 0 demotes records that are thin for reasons other than fit - SMB, non-US, privacy-conscious buyers - so track their conversion as a separate cohort and re-derive the weight if they convert anyway. What this order starves is enrichment: first on value, first on effort, so a ratio never buys it, and it is the only rung that restores the signal instead of routing around it.

   Promote enrichment when:

   - The field is the ICP gate itself (employee count at enterprise ACV).
   - Coverage is low across the whole database rather than one segment.
   - A contract already exists.

   Delete the enrichment rung outright, and say it is deleted, where the Interview gave no procurement path - a rung parked at the bottom of the list returns as a quarter of vendor calls nobody approved.

3. Derive signals from outcomes, not intuition.

   - Pull 12-24 months of closed-won and closed-lost records.
   - Compare signal frequency between the two groups.
   - Keep signals with a real conversion delta over baseline.

   The gut list and the evidenced list are usually different lists. Trap: fields that look complete on won deals because sales filled them right before close - verify each signal existed at scoring time.

4. Build the signal catalogue from [references/signal-taxonomy.md](references/signal-taxonomy.md): fit and engagement as two separate scores, hard disqualifiers as exclusions (never negative points), soft demotions as negative points, decay on engagement only.
5. Weight per [references/weighting-and-thresholds.md](references/weighting-and-thresholds.md): start near 60/40 fit/engagement for sales-led, 50/50 hybrid, 30-40 fit / 60-70 engagement for PLG and B2C where behavior dominates. Cap each signal category so no single channel can cross the threshold alone.
6. Set the threshold by sales capacity first, conversion band second, percentile as sanity check - that is the efficiency ranking of the three methods, not just a sequence, and the same reference states each method's value and effort before choosing. Define tiers where each tier maps to a distinct play.
7. Backtest per [references/validation-and-kpis.md](references/validation-and-kpis.md): retro-score history, compute conversion by score band, compute the top band's lift over the unscored baseline, run the spot-check holdout set. Iterate signals, weights, and threshold until the pass threshold below clears.
8. Get written sales sign-off on the signal list, the threshold, and the acceptance floor. Models built without it are the most-cited failure in the field: marketing builds alone; sales rejects half the leads; both blame the other.
9. Emit the spec (next section), roll out staged - score the live flow before batch-scoring the backlog - and book the v2 review for 60 days post-launch before launching.
10. If your harness has persistent memory, memorize the approved spec - signals, weights, threshold, version, review date - so later recalibration and diagnosis runs start from it instead of re-interviewing.

## The Scoring Model Spec

Deliver every engagement as this artifact - a document the user can hand to marketing ops and sales. Fuller worked versions live in [references/examples.md](references/examples.md).

```
SCORING MODEL SPEC - <company/product>, <date>, v<n>
Motion       : sales-led | PLG | B2C - fit/engagement weight split, scale (e.g. 100-pt)
Fit signals  : attribute -> points | data source | coverage % | evidence (won-vs-lost delta)
Engagement   : signal -> points | decay rule | per-category cap | evidence
Exclusions   : hard disqualifiers, excluded outright (competitors, students, customers...)
Negative pts : soft demotions -> points
Threshold    : MQL/PQL score + the capacity math behind it
Tiers        : band -> play (T1 fast human follow-up, T2 queue, T3 nurture...)
Backtest     : period, conversion by band, top-band lift vs baseline, false-negative %
Guardrails   : sales acceptance floor, score-distribution inflation alarm
Governance   : owner, sales sign-off (name, date), v2 review date, change-log location
```

Three rules about the spec itself:

- **Every point value carries its evidence.** "Pricing page 2+ visits +20: converted 2.8x baseline in last-FY won analysis" is reviewable; a bare number is folklore waiting to be gamed.
- **The spec has a pass threshold.** In backtest, the top score band must convert to opportunity at **at least 2x the unscored all-leads baseline** (published guidance only demands "beats no-scoring"; treat 2x as this skill's practical floor, since a thin edge will not survive live noise). The backtest's false-positive rate must also make the agreed **sales acceptance floor - 60% is the common practitioner target, a convention rather than a researched constant** - plausible. Iterate until both clear; refuse to ship a model that does not beat the unscored baseline.
- **The model is versioned.** v1 is a starting position; the v2 review is booked before v1 ships; every change is logged and announced to sales. An unexplainable score kills rep adoption faster than a wrong one.

## Diagnosing a Broken Model

When the user arrives with a running model ("our scores are wrong", "sales ignores the score"), check every symptom below before proposing a rebuild - checking is cheap and most broken setups fail two or three, not one. Fixing is where the choice is, so the rows are ordered by value per unit of effort; work them top-down:

- efficiency: zero the vanity signals > report the right number > add decay > fix the coverage gap > add the missing buying paths > re-derive inflated weights > full recalibration
- value (opportunities the corrected model wins, or stops losing): full recalibration > add the missing buying paths > re-derive inflated weights > add decay > zero the vanity signals > fix the coverage gap > report the right number
- effort: full recalibration > re-derive inflated weights > add the missing buying paths > fix the coverage gap > add decay > report the right number > zero the vanity signals

| Symptom                                        | Likely fault                                                                                       | Fix                                                                                         | Effort                                                                                 |
| ---------------------------------------------- | -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- |
| Newsletter devotees outrank real buyers        | Vanity signals carry points (email opens are structurally unreliable since inbox privacy features) | Zero or floor low-intent signals; cap the email category                                    | Near-zero - point values already in the matrix, edited by the model owner alone        |
| MQL count celebrated while pipeline is flat    | Goodhart dynamic - MQL volume became the target, so the bar drifts down                            | Report MQL-to-opportunity conversion and pipeline contribution upward, never raw MQL volume | Near-zero in hours, but spends political capital with whoever set the MQL target       |
| Stale contacts look as hot as active ones      | No decay - points accumulate forever                                                               | Add decay to engagement only (mechanisms ranked in signal-taxonomy)                         | An hour of config, plus one rescore of the database                                    |
| Most of the database can never reach threshold | Enrichment ceiling - fit fields empty for most records                                             | Measure coverage, then apply the ranked menu in Workflow step 2 - score unknowns 0 first    | Near-zero for the default rung; a quarter and a standing job if enrichment is promoted |
| Closed-won deals that never hit MQL            | False negatives - model ignores real buying paths (referrals, events, product usage)               | Add the missing paths; track false-negative rate as a standing KPI                          | A week - find the path in won records, add signals, re-backtest                        |
| MQL volume up, MQL-to-opportunity down         | Score inflation - generous weights or vanity signals                                               | Re-derive weights from won/lost data; add category caps; raise threshold                    | A week of analyst work on won/lost cohorts, plus sales re-sign-off on the new bar      |
| High scores stopped predicting wins            | Stale model - ICP, product, or content shifted since calibration                                   | Recalibrate on latest multi-quarter cohort; install drift triggers                          | A quarter to rebuild, then a standing job to keep                                      |

What this order starves: the full recalibration - first on value, first on effort, so a ratio never schedules it. Promote it above everything when ICP, pricing, or product changed since the last calibration, or when two rows above have already been fixed and conversion still does not track the score.

Where the Interview gave no analyst, delete the two cohort-analysis rungs - re-deriving weights and the full recalibration - and say they are deleted rather than leaving them at the bottom of a plan nobody can execute. Against inflation, that leaves category caps and a raised threshold, which a CRM admin can ship alone and which buy back most of the precision.
lead-scoring · 人気上昇中の Agent Skills | Mengbi