返回 Skills 目錄
mbfinotti/revops-skills已通過檢查

SKILL DETAIL

customer-health-score

mbfinotti/revops-skills/customer-health-score

Design, validate, and govern a composite customer health score - product usage and engagement signals combined into one weighted, normalized, decayed, banded account score that flags both churn risk and expansion readiness, backtested against real churn outcomes and recalibrated. Use whenever the user mentions a customer health score, health scoring model, account health, signal weighting, score bands, at-risk account flagging, churn risk score, expansion risk, or says "green accounts keep churning" - even if they never say "health score". Covers B2B account-level rollup and B2C/PLG continuous scoring. Do NOT use for discovering which indicators actually predict churn - use mbfinotti/revops-skills@customer-churn-signals instead.

安裝量 · 163查看來源

Installation

npx skills add https://github.com/mbfinotti/revops-skills --skill customer-health-score

技能檔案

SKILL.md

最近同步 · 2026年9月15日

evals/evals.json
{
  "skill_name": "customer-health-score",
  "evals": [
    {
      "id": 1,
      "prompt": "I run CS ops at Kelfast, B2B workflow automation, 620 accounts, about $14M ARR. My VP wants a customer health score live before the QBR in six weeks. I already sketched the model in a workshop with the CSM leads: logins 40%, NPS 20%, ticket count 20% (more tickets = worse), CSM gut rating 20%, and five bands A through E. We have three years of CRM and product analytics data, and 47 accounts churned in the last 24 months. Can you turn this into something we can actually ship? Mostly I want to know whether the weights look right.",
      "expected_output": "A rebuilt model: weights re-derived from churned-vs-retained separation rather than the workshop, three bands, per-seat rate normalization, ticket trend instead of ticket count, decay toward neutral, a capped category structure with a stakeholder-continuity input that can pull the band down alone, a structured replacement for the gut rating, and validation framed on precision/recall against a defined churn event.",
      "files": [],
      "expectations": [
        "Rejects weights set in a workshop or by intuition and instructs deriving them from the separation between the churned and retained cohorts over the trailing 12 to 24 months.",
        "Requires defining the churn event, the entity level, and the prediction window before any further signal or weight discussion.",
        "Reduces the band count from five to three.",
        "Replaces raw login counts with a per-seat or per-active-user rate, stating that raw counts make large accounts look healthy and small accounts look sick by construction.",
        "Flags that scoring tickets as 'more tickets = worse' punishes engaged accounts, and replaces it with a rate-normalized ticket trend or severity and resolution-time trend.",
        "States that zero tickets can indicate disengagement rather than health.",
        "Adds a decay rule so engagement-type signals drift toward neutral as they age, and specifies drift toward neutral rather than toward red.",
        "Requires at least one stakeholder-continuity input with enough downside weight to pull an account's band down on its own.",
        "Caps every category's share so no single input can move an account across a band by itself.",
        "Replaces the unstructured CSM gut rating with structured rating definitions that state what each rating requires, or with gated overrides carrying reason codes.",
        "Specifies validation on precision and recall rather than raw accuracy, naming class imbalance as the reason."
      ]
    },
    {
      "id": 2,
      "prompt": "Hoyt Systems, 1,100 accounts, we have had a health score running for about two years. Problems, all at once: the CSMs basically ignore it and call it 'the weather', our big enterprise logos are permanently red because they open tons of tickets, accounts bounce green to yellow and back week to week, and 88% of the book sits green so the number never tells us anything. We have roughly 60 churns over the past two years to learn from. Where do we start? I have one analyst for about a day a week.",
      "expected_output": "A diagnosis of all four symptoms with an explicit fix ordering by churn prevented per hour of analyst time, leading with wiring bands to plays and normalization, and deferring cohort re-derivation on effort grounds while naming it high value.",
      "files": [],
      "expectations": [
        "Diagnoses every reported symptom rather than only the first or the most severe.",
        "States the fix ordering explicitly and ranks it by churn actually prevented per hour of analyst effort.",
        "Places wiring each band to a play with an owner and an SLA at or near the top of the fix order.",
        "Attributes the CSMs' distrust to unwired plays, absent gated overrides, and unpublished weight evidence.",
        "Attributes the permanently red enterprise accounts to raw counts used instead of rates, and prescribes rate normalization per seat or against similar-size accounts.",
        "Attributes weekly band flapping to single-point noise treated as signal, and requires two consecutive readings or a trailing average before a band change fires.",
        "Attributes the 88% green book to missing decay, and names more than roughly 80% green as a broken-model red flag independent of backtest results.",
        "Prescribes rate-normalizing tickets and reading them as a trend rather than as counts.",
        "Places re-deriving weights from the churned-versus-retained cohort late in the fix order on effort grounds while acknowledging it is among the highest-value fixes.",
        "Does not order the fixes by severity of fault or in the order the symptoms were reported.",
        "Gives each fix an order-of-magnitude effort (near-zero, an hour, a week, a standing job) rather than a currency amount or a precise hour count."
      ]
    },
    {
      "id": 3,
      "prompt": "Brightsill, Series A, 140 customers all on annual contracts, 18 months since our first paying customer. We have lost 5 accounts total. The board wants a churn-risk score by end of quarter. Our data scientist wants to train a logistic regression on what we have and output a 0-100 probability per account. Does that work, and what would you build instead if it doesn't?",
      "expected_output": "A rules-based proxy score labeled as hypothesis-weighted and explicitly non-predictive, with the outcome definition, score logging, and monthly CSM interview validation set up now, and statistically fitted weights deferred until tens of churn events per segment exist.",
      "files": [],
      "expectations": [
        "Rejects fitting a statistical model on 5 churn events and names the insufficient number of churn events as the reason.",
        "Recommends a rules-based proxy score instead, built from inputs such as onboarding-milestone completion, activation or adoption thresholds, recency tripwires, and champion-contact recency.",
        "Labels every weight in the rules-based score as a hypothesis rather than an evidenced weight.",
        "Refuses to present the initial score as predictive, or states plainly that it carries no predictive claim.",
        "Still requires defining the churn event, the entity level, and the prediction window now, so validation can begin once events accumulate.",
        "Instructs logging the score at fixed intervals during the first 2 to 4 quarters so later backtests can retro-read it honestly.",
        "Prescribes monthly interview validation with experienced CSMs on which accounts the score got wrong.",
        "Names tens of churn events per segment as the threshold that promotes statistically fitted weights.",
        "States that the fitted score is the higher-value option and is demoted on effort ratio rather than on quality.",
        "Warns against tuning the model to individual churn anecdotes while event counts are small.",
        "Does not propose treating a backtest against the 5 existing churn events as validation of the model."
      ]
    },
    {
      "id": 4,
      "prompt": "Vantail, B2B, around 900 accounts, $60M ARR. Our CRO wants an 'expansion readiness score' separate from the churn health score we already run, so CS can hand AEs a ranked list of accounts ready to upsell. He also asked me for a number he can put in the board deck on how much expansion lift a green account gives us. We do not track whitespace anywhere today: product usage sits in our analytics tool, everything else in the CRM. What should the expansion model look like?",
      "expected_output": "A recommendation to keep one composite score with the top band gating expansion plays, a refusal to quote an expansion lift figure, and a qualification layer on whitespace, fit, plan-limit pressure, and budget or contract timing.",
      "files": [],
      "expectations": [
        "Recommends one composite score whose top band gates expansion plays, instead of building a second dedicated expansion model.",
        "Names the systematically tracked expansion inputs that would justify a second model, such as whitespace mapping, multi-team adoption spread, and plan-limit trajectories.",
        "States that a second model built before those inputs are tracked is a second unvalidated score.",
        "Refuses to provide a numeric expansion lift figure for a top-band or green account.",
        "States that no rigorous published study shows high health causally predicting expansion, and that the claim rests on vendor case studies.",
        "Frames health as necessary but not sufficient for expansion.",
        "Requires qualifying top-band accounts on non-health conditions including whitespace, fit, plan-limit pressure, and budget or contract timing.",
        "Describes the deliverable as a better-qualified expansion queue rather than a promised expansion lift.",
        "Names the cost of a second model as roughly a quarter to build plus a standing second validation loop to maintain.",
        "Notes that legitimate score splits are more commonly by lifecycle stage or segment than by risk versus expansion outcome.",
        "Wires the top band to a named expansion or advocacy play with an owner."
      ]
    },
    {
      "id": 5,
      "prompt": "Two of our biggest accounts churned last quarter at Nerrin Cloud, and both were sitting at 82 and 79 in our health score the week they gave notice. In both cases the exec who originally bought us left months earlier, but usage never dropped because their ops team kept running the same weekly reports. Our current model is usage 60 / support 20 / NPS 20, aggregated as total events per account. We have about 400 accounts averaging 45 seats. How do we stop this from happening again?",
      "expected_output": "A model change adding a weighted stakeholder-continuity component with band-down power, a usage cap, per-user normalized rollup with role weighting, and an engagement-breadth check.",
      "files": [],
      "expectations": [
        "Names the pattern as silent churn: flat aggregate usage hiding a departed or disengaged champion.",
        "Adds a relationship or stakeholder-continuity category as its own weighted component rather than folding it into aggregate usage.",
        "Requires that at least one stakeholder-continuity input be able to pull the band down on its own.",
        "Caps the usage category so that no combination of usage greens can hold an account green after a champion departs.",
        "States the deliberate asymmetry: a continuity input may force the band down alone, while no single positive signal may force green.",
        "Replaces total per-account event aggregation with per-user normalized rates rolled up to the account level.",
        "Applies stakeholder-role weighting in the rollup so champion and decision-maker signals count more than end-user signals.",
        "Adds an engagement-breadth check, flagging engagement concentrated in one person as fragile even when the aggregate is high.",
        "States that usage describes behavior and never intent.",
        "Names concrete continuity inputs such as executive-sponsor contact recency, responsiveness trend, or stakeholder breadth across roles.",
        "Notes that continuity inputs depend on humans logging meetings and sponsor changes, making the category a standing job rather than a one-time wire-up."
      ]
    },
    {
      "id": 6,
      "prompt": "I run growth at Tumbledry, a $12/month self-serve design tool, about 40,000 paying subscribers, month to month, and we have no CSMs at all. Monthly churn is around 4.5% and I want a health score so we can catch people before they cancel. I was going to reuse the model from my last company, enterprise SaaS: sponsor engagement 30%, QBR attendance 20%, adoption 30%, support 20%, reviewed by the CSM team every quarter. What changes?",
      "expected_output": "A continuous, system-driven subscriber score: sponsor and QBR inputs removed, an involuntary-churn billing layer added, growth/product ownership, monthly recalibration, no rollup for single-user subscribers, automated band plays, and explicit confirmation of which B2B mechanics carry over unchanged.",
      "files": [],
      "expectations": [
        "Removes the sponsor-engagement and QBR-attendance inputs as inapplicable without a CSM layer or a renewal cliff.",
        "Replaces renewal-anchored periodic evaluation with continuous, system-driven scoring.",
        "Adds an involuntary-churn layer covering failed payments and dunning outcomes into the score itself rather than handling it outside the score.",
        "Justifies the billing layer by payment failure correlating with engagement decline and being an intervention trigger in its own right.",
        "Assigns ownership of the scoring logic to growth or product rather than to CS ops.",
        "Sets recalibration to monthly rather than quarterly, on the grounds of higher volume and faster drift.",
        "Skips user-to-account rollup for single-user subscribers, treating the subscriber as the scored entity.",
        "States explicitly that trend-over-level scoring, decay, band-to-play wiring, and the backtest loop carry over from the B2B model unchanged.",
        "Defines the churn event for this business as cancellation or no reactivation within a stated number of days.",
        "Uses a recency and frequency read of usage rather than a cumulative usage total.",
        "Wires each band to an automated play such as a win-back sequence, dunning retry, lifecycle nudge, or upgrade offer, rather than to a human motion."
      ]
    },
    {
      "id": 7,
      "prompt": "We just finished backtesting the new health score at Orcastead, two years of data. Results: overall model accuracy 91%, the red band churns at 2.1x the portfolio base rate, and 44% of the accounts that actually churned were still green 60 days before they left. 78% of the book currently sits green. We run one model across everyone, enterprise accounts averaging 300 seats with a dedicated CSM, and a self-serve SMB tier averaging 6 seats with no CSM where about half are still in onboarding. My VP says 91% accuracy means we are ready to launch. Thoughts?",
      "expected_output": "A refusal to launch: accuracy discarded on class-imbalance grounds, both the 3x red-band floor and the two-thirds pre-churn recall floor shown as failed, the 78% green distribution flagged as near the red-flag line, and segment-specific weights and thresholds required before relaunch.",
      "files": [],
      "expectations": [
        "Rejects the 91% accuracy figure as evidence, naming class imbalance and noting that a model calling everything green scores high accuracy while predicting nothing.",
        "States that the red band fails the pass floor of at least 3x the portfolio base rate, citing the 2.1x result.",
        "Converts the 44% still-green figure into 56% of churned accounts having sat below green, and states that this fails the floor of at least two-thirds.",
        "Refuses to approve the launch, or refuses to present the score as predictive in its current state.",
        "Notes that 78% green sits just under the roughly 80% distribution red flag and must be monitored rather than treated as a clean result.",
        "Requires segment-specific weights and thresholds for the enterprise and self-serve SMB tiers instead of one global model.",
        "States that a score healthy for a mature enterprise account can be a red flag for an SMB still in onboarding.",
        "Requires segmenting at minimum by size or touch model and by lifecycle stage.",
        "Warns that each segment needs its own validation cohort, so the segment count must stay as small as the differences justify.",
        "Prescribes identifying which signal category would have caught the accounts that read green until the end, and reweighting accordingly.",
        "Notes that low-touch segments run naturally lower engagement, shifting weight there toward sentiment or support responsiveness."
      ]
    },
    {
      "id": 8,
      "prompt": "The model is finally agreed at Pellhurst Software: usage 35, relationship 30, support 15, commercial 10, sentiment 10, backtest done and it held up. Now I have to hand it over to three groups. Our CS ops lead builds it, the CSM team uses it every day, and analytics engineering owns the pipeline. Last time we rolled out a scoring thing it drifted within six months because every team quietly added their own field to it. What do I actually hand them? Also, the CSMs keep asking to be able to mark an account healthy when they think the score is wrong, and I would rather not let them do that.",
      "expected_output": "A complete written health score spec with evidence attached to each weight, bands wired to plays with owners and SLAs, named single owners for logic and data, versioned change logging, plus a reversal of the user's position on overrides in favor of reason-coded overrides feeding recalibration.",
      "files": [],
      "expectations": [
        "Delivers a single written health score spec artifact rather than loose prose guidance.",
        "The spec states the outcome definition: churn event, entity level, and prediction window.",
        "The spec lists segments together with their per-segment weight and threshold overrides.",
        "The spec attaches outcome-correlation evidence to each weight rather than stating the percentage alone.",
        "Labels any weight lacking outcome evidence as a provisional hypothesis carrying a validation date.",
        "The spec maps each band to a play, an owner, and an SLA.",
        "The spec includes decay rules, the user-to-account rollup method, validation results, band-change triggers, and a governance block.",
        "Reverses the user's position on overrides and recommends allowing CSM overrides rather than blocking them.",
        "Requires a mandatory reason code on every override.",
        "Treats override clusters as labeled disagreement data reviewed at recalibration, and a repeated reason code as a model bug filed by the field.",
        "Names one owner for the scoring logic, one data owner, and the CSMs as acting owners, against the diffuse ownership that caused the previous drift.",
        "Requires versioning every change with date, author, reason, and expected effect, announced to the CS team before it lands."
      ]
    },
    {
      "id": 9,
      "prompt": "Kessler Rowe, B2B compliance software, 260 accounts, renewals clustered heavily in Q1. Our product already streams every click into the warehouse, billing is in Stripe, support in Zendesk. Nobody logs meetings: the entire CS team is two people and they live in their inbox. Board review is in seven weeks and they want an at-risk list out of this. Realistically those two CSMs can work maybe 12 accounts a month on a save motion. Which signals do we wire first, and where do we set the red cutoff?",
      "expected_output": "A build sequence starting from usage and billing signals, with continuity promoted despite poor ratio because renewals are date-anchored, sentiment surveys deferred, renewal proximity weighted inside the score, and the red band sized to the stated 12-accounts-per-month capacity.",
      "files": [],
      "expectations": [
        "Sizes the red band to the stated intervention capacity of roughly 12 accounts per month before tuning it on statistics.",
        "States that the score is then improved until the same capacity catches more of the actual churn, rather than widening the red band.",
        "Starts from product usage and adoption plus commercial or billing signals, citing near-zero effort because the product and billing systems already emit those events.",
        "Defers the sentiment survey category, citing roughly a week to stand up plus a standing job to run, and unstable results from low response rates.",
        "Explicitly re-ranks the default build order against the user's stated constraints and names which constraint moved which signal category.",
        "Promotes stakeholder continuity despite its poor effort ratio because this is a B2B book with contract renewal dates.",
        "Resolves the conflict between nobody logging meetings and the continuity requirement by falling back to whatever the CRM captures on its own, such as email or calendar-derived contact recency.",
        "Weights renewal proximity inside the score rather than applying it as a filter after scoring.",
        "Rejects integrating every available signal category at once.",
        "Drops any candidate signal that could not trigger a play, naming it a vanity signal."
      ]
    },
    {
      "id": 10,
      "prompt": "We are about to validate the score at Aldermist. The plan: pull the current warehouse tables, compute the score for every account that churned in the last 18 months using their full history, compare against the accounts that renewed, then report to the exec team what percentage of accounts have a score at all plus the overall hit rate. After that we would review the model once a year at annual planning. Anything missing?",
      "expected_output": "A corrected backtest that retro-scores at a fixed past point using only data available then, adds the 60-90 day pre-churn retro check and hand spot-checks, replaces coverage reporting with save and expansion outcomes, and replaces the annual review with monthly-then-quarterly recalibration plus off-cycle triggers.",
      "files": [],
      "expectations": [
        "Rejects scoring churned accounts on their full history and requires retro-scoring at a fixed point in the past, such as 90 days before each churn or renewal event.",
        "Requires using only data that existed at that past point, and names signals backfilled after the outcome as the classic leak.",
        "Rejects score coverage, the percentage of accounts carrying a score, as a success metric reported upward.",
        "Replaces coverage with save and expansion outcomes: red-band plays that retained the account, and top-band plays that expanded.",
        "Adds the retro check on what churned accounts scored at 60 and 90 days before churning, with a post-mortem for each account that read green until the end.",
        "Adds hand spot-checks of known accounts covering a recent churn, a recent expansion, a known-shaky renewal, and a champion-departure case that must not read green.",
        "Replaces the annual review with monthly review for the first 90 days after launch, then a standing quarterly recalibration.",
        "Names off-cycle recalibration triggers such as a product or pricing change, an ICP or segment shift, green creep in the distribution, a cluster of green-account churns, or a rising override rate.",
        "Reports a KPI set covering the red-band churn multiple, missed-churn rate, score distribution by segment, play completion within SLA, and override rate clustered by reason code."
      ]
    }
  ],
  "trigger_queries": [
    { "query": "design a customer health score for our B2B SaaS", "should_trigger": true },
    { "query": "our green accounts keep churning and nobody can explain it", "should_trigger": true },
    { "query": "how should I weight the inputs in our account health model?", "should_trigger": true },
    { "query": "we need a way to flag at-risk accounts before the renewal conversation", "should_trigger": true },
    { "query": "what should the bands be on a health score?", "should_trigger": true },
    { "query": "the CSMs completely ignore the health score we built last year", "should_trigger": true },
    { "query": "help me build one composite number that tells us which customers are drifting away", "should_trigger": true },
    { "query": "our score says 90% of the book is healthy, that cannot be right", "should_trigger": true },
    { "query": "how do I validate a customer health score?", "should_trigger": true },
    { "query": "should we have a separate expansion score or just one health score?", "should_trigger": true },
    { "query": "we score customers 0 to 100 but nobody trusts the number", "should_trigger": true },
    { "query": "customer health scoring model design", "should_trigger": true },
    { "query": "how many health tiers should we use, three or five?", "should_trigger": true },
    { "query": "our biggest accounts always look red, something is off with the scoring", "should_trigger": true },
    { "query": "recalibrate our churn risk score, it has drifted", "should_trigger": true },
    { "query": "we want one number per account that tells the CSM what to do next", "should_trigger": true },
    { "query": "backtest our account health model against last year's churn", "should_trigger": true },
    { "query": "how do I stop the score flapping between green and yellow every week?", "should_trigger": true },
    { "query": "what is a sensible split between product usage and relationship signals in a health score?", "should_trigger": true },
    { "query": "we need to know which customers are ready for an upsell and which are about to leave", "should_trigger": true },
    { "query": "the champion left and our score stayed green for two months", "should_trigger": true },
    { "query": "health scoring for a PLG product with no CSMs", "should_trigger": true },
    { "query": "should health thresholds be different for enterprise and SMB accounts?", "should_trigger": true },
    { "query": "our score has no decay so everything just accumulates green forever", "should_trigger": true },
    { "query": "who should own the health scoring logic in a RevOps team?", "should_trigger": true },
    { "query": "should CSMs be allowed to override the health score?", "should_trigger": true },
    { "query": "write the spec for our customer health scoring model", "should_trigger": true },
    { "query": "how do I roll user-level product usage up into one account health number?", "should_trigger": true },
    { "query": "we only have five churns ever, can we still build a risk score?", "should_trigger": true },
    { "query": "how do I prove to the exec team that our health score actually predicts churn?", "should_trigger": true },
    { "query": "the account health model only ever tells us things we already knew", "should_trigger": true },
    { "query": "I want a red yellow green system across our customer base", "should_trigger": true },
    { "query": "what should happen automatically when an account drops a band?", "should_trigger": true },
    { "query": "customer health index design for a subscription business", "should_trigger": true },
    { "query": "the CS team wants a health score and I have no idea where to start", "should_trigger": true },
    { "query": "is NPS a good input into a health score or not?", "should_trigger": true },
    { "query": "we are building post-sale account scoring, not lead scoring", "should_trigger": true },
    { "query": "how often should we re-derive the weights in our health model?", "should_trigger": true },
    { "query": "I need a systematic way to grade every account in the book from healthiest to most at risk", "should_trigger": true },
    { "query": "the number next to each account means something different to every team here", "should_trigger": true },
    { "query": "we want to catch subscribers before they cancel based on how they use the app", "should_trigger": true },
    { "query": "how do we decide how many accounts should land in the red band?", "should_trigger": true },
    { "query": "score said green, the renewal came back no, what is broken?", "should_trigger": true },
    { "query": "design the scoring model CS uses to prioritise its week", "should_trigger": true },
    { "query": "tune our at-risk threshold, there are way too many false alarms", "should_trigger": true },
    { "query": "our health score predicts nothing, I want to rebuild it from scratch", "should_trigger": true },
    { "query": "add payment failures and dunning outcomes into how we score subscriber health", "should_trigger": true },
    { "query": "account health scoring for enterprise accounts anchored on renewal dates", "should_trigger": true },
    { "query": "which product events actually predict churn for us? rank them by lift over base rate", "should_trigger": false },
    { "query": "build an early warning system of churn indicators with thresholds, windows and lead times", "should_trigger": false },
    { "query": "score inbound leads on fit and engagement so marketing knows what counts as an MQL", "should_trigger": false },
    { "query": "our MQL threshold is letting junk through, recalibrate the lead model", "should_trigger": false },
    { "query": "route new demo requests to the right rep by territory and capacity", "should_trigger": false },
    { "query": "check the health of our Kubernetes pods and tell me what is degraded", "should_trigger": false },
    { "query": "add a /healthz endpoint to this service", "should_trigger": false },
    { "query": "diagnose why this cloud resource is reporting degraded health", "should_trigger": false },
    { "query": "score our codebase health and flag the worst modules for refactoring", "should_trigger": false },
    { "query": "build a metric tree from board-level ARR down to what an individual SDR owns", "should_trigger": false },
    { "query": "what goes in the sales to CS handoff packet at closed won?", "should_trigger": false },
    { "query": "design our funnel stages from first touch to closed won", "should_trigger": false },
    { "query": "find where we are losing revenue between MQL and closed won and size it in dollars", "should_trigger": false },
    { "query": "audit our pipeline stage exit criteria against buyer-verifiable milestones", "should_trigger": false },
    { "query": "our sales forecast is always wrong, figure out why", "should_trigger": false },
    { "query": "flag stale deals in the pipeline and give me a disposition for each one", "should_trigger": false },
    { "query": "who owns the industry field in our CRM and how often must it be refreshed?", "should_trigger": false },
    { "query": "which system should be source of truth for subscription objects across our stack?", "should_trigger": false },
    { "query": "we run 40 GTM tools, help me decide which ones to cut at renewal", "should_trigger": false },
    { "query": "set up the discount approval matrix for non-standard deals", "should_trigger": false },
    { "query": "what metrics belong in the board revenue deck and in what order?", "should_trigger": false },
    { "query": "am I ready for a senior RevOps role? review my resume", "should_trigger": false },
    { "query": "write the job description and interview loop for a RevOps manager", "should_trigger": false },
    { "query": "which RevOps newsletters and podcasts should I be following?", "should_trigger": false },
    { "query": "write a win-back email sequence for customers who already cancelled", "should_trigger": false },
    { "query": "design a cancellation flow with a save offer in it", "should_trigger": false },
    { "query": "improve activation in our product onboarding so more trials convert", "should_trigger": false },
    { "query": "how do I run my first customer QBR next week?", "should_trigger": false },
    { "query": "draft the renewal negotiation email for this account", "should_trigger": false },
    { "query": "build an NPS survey and help me word the follow-up question", "should_trigger": false },
    { "query": "what is a good monthly logo churn rate for a seed-stage B2B SaaS?", "should_trigger": false },
    { "query": "calculate our net revenue retention for last quarter", "should_trigger": false },
    { "query": "set up a BI dashboard showing product usage by account", "should_trigger": false },
    { "query": "write a customer success career ladder from CSM to director", "should_trigger": false },
    { "query": "our support queue is backed up, how do we triage tickets faster?", "should_trigger": false },
    { "query": "segment our customer base by industry for a marketing campaign", "should_trigger": false },
    { "query": "predict which free trial users will convert to paid", "should_trigger": false },
    { "query": "price our new enterprise tier", "should_trigger": false },
    { "query": "how do we work out whether our CS team is understaffed?", "should_trigger": false },
    { "query": "instrument the product with analytics events for the new feature", "should_trigger": false },
    { "query": "our warehouse sync is failing and the account table is stale", "should_trigger": false },
    { "query": "build a scorecard for evaluating the vendors we are buying from", "should_trigger": false },
    { "query": "score our channel partners by tier and performance", "should_trigger": false },
    { "query": "run a post-mortem on the outage that hit 30 customers yesterday", "should_trigger": false },
    { "query": "write the CS playbook for onboarding a new enterprise customer in the first 30 days", "should_trigger": false },
    { "query": "forecast next quarter's renewal revenue by account", "should_trigger": false },
    { "query": "how do I get our CSMs to actually log their meeting notes in the CRM?", "should_trigger": false },
    { "query": "compare the main customer success platforms and tell me which to buy", "should_trigger": false }
  ]
}
references/signal-categories.md
# Signal Categories Inside a Composite

Boundary note: this file covers how each category behaves once it is _in_ the score - normalization, trend windows, rollup, traps. Which specific indicators to pick and how predictive each is belongs to `mbfinotti/revops-skills@customer-churn-signals`; take its output as this file's input.

Sections run in the build ranking from SKILL.md § Build order, best value-per-hour first. Re-rank against the user's own data before following it - the order assumes the product emits usage events and that nobody is logging sponsor meetings yet.

## Product usage and adoption

Typically the heaviest-weighted category (40-50% in circulating SaaS examples - illustrative, not a standard; derive the real weight from the book's own outcomes).

- Measure:
  - Breadth: share of core capabilities adopted.
  - Depth: intensity per active user.
  - Recency/frequency: an RFM-style read.
  - Seat utilization: active vs purchased.
- Always rate-normalize: divide activity by user count. Raw counts make big accounts look healthy and small accounts look sick by construction.
- Score the trend, not the level - a 30% decline from a high baseline outranks a stable low baseline as a risk read.
- Trap: usage describes behavior, never intent. An account can show strong usage while evaluating a competitor or absorbing a sponsor change the dashboard can't see. This is why usage can never be the whole score - cap its share.

## Commercial / financial

- Signals: payment timeliness drift (net-30 sliding toward 75 days), contract-value trajectory, discount/renegotiation frequency, invoice disputes.
- This category surfaces risk when usage looks fine: daily logins coexist with disputed invoices and a resigned champion.
- B2C/PLG: include the involuntary-churn layer - failed payments and dunning outcomes belong in the score, since payment failure correlates with engagement decline and is itself an intervention trigger.

## Support

- Ticket volume alone is not a risk signal. Rate-normalize by account size, read it as a trend, and read severity/escalations and resolution-time trend rather than counts.
- An account filing one critical bug a quarter can be healthier than one filing zero tickets - zero can mean disengagement.
- Context-blindness trap: an SLA miss during an outage that hit dozens of accounts is not the same signal as a miss specific to one account.

## Relationship and engagement (champion/sponsor)

The highest-value and hardest-to-automate category.

- Signals: meetings held, executive-sponsor contact recency, responsiveness trend, stakeholder breadth across roles and functions.
- That combination is what puts it fourth on ratio and first on value; build it early anyway wherever renewals hinge on a named buying organization.
- In B2B, weight champion/decision-maker engagement as its own component rather than folding it into aggregate usage - flat aggregate activity hides a champion who stopped logging in while junior users keep running routine tasks ("silent churn").
- Design rule: at least one stakeholder-continuity input must be able to pull the band down on its own; no combination of usage greens may hold an account green after a champion departs.
- Structure CSM sentiment where it feeds the score: define what each rating requires ("cannot be green without an executive-buyer meeting in the last 2 months and regular active users" - the ChurnZero-style definitional approach). Unstructured gut ratings are not comparable across a portfolio.

## Sentiment (survey and qualitative)

- Signals: relationship surveys (NPS-style), post-interaction satisfaction, effort scores, ticket-text tone.
- Sentiment is context for behavior, not a standalone predictor: strong usage with falling sentiment often means the customer uses the product out of necessity while hunting alternatives.
- Read as trend across periods; a single response is noise. Low response rates make this category unstable - weight accordingly.

## Rollup: user-level to account-level (B2B)

- Aggregate as normalized weighted averages (activity per user, not totals), with stakeholder-role weighting: champion and decision-maker signals count more than end-user signals.
- Breadth matters: engagement concentrated in one person is fragile even when the aggregate is high; multiple inactive admins are a signal one inactive user is not.
- B2C and single-user PLG skip rollup entirely - subscriber level is the account level. Everything else in this file applies unchanged.
references/validation-and-kpis.md
# Validation, Cold Start, KPIs, Governance

A composed-but-unvalidated score is the industry's default failure - 73% of CS leaders say theirs doesn't reliably predict churn (ChurnZero 2025 study). Everything here exists to put the user's score in the other 27%.

## Backtest procedure

1. Retro-score accounts as they stood at a fixed point in the past (e.g. 90 days before each renewal or churn event), using only data that existed at that time. Signals backfilled after the outcome are the classic leak.
2. Compare predicted band against actual outcome over the prediction window, per segment.
3. Use precision/recall, never raw accuracy: churn is class-imbalanced, so a model that calls everything green scores high accuracy while predicting nothing.
   - **Precision of red** = share of red accounts that actually churned (wasted-intervention cost when low).
   - **Recall of red** = share of churned accounts the red band caught (missed-churn cost when low).

   Tune the red threshold to the costlier error: recall-leaning for high-value segments, precision-leaning when interventions are expensive. No ranking between the two leanings; which error costs more is a property of the user's book, not of the method, so a default order here would be false precision.

4. Run the retro check: for every account that churned, what did it score 60 and 90 days before? Churned accounts that read green until the end mark the model's blind spot - identify which category would have caught them and reweight.
5. Sanity-check the distribution: more than ~80% of the book sitting green is a documented practitioner red flag for a broken model, whatever the backtest says.
6. Spot-check a handful of known accounts by hand. Each must land where an experienced CSM would put it, and the champion-departure case must not be green.
   - A recent churn.
   - A recent expansion.
   - A known-shaky renewal.
   - A champion-departure case.

## Pass floor (this skill's convention)

No standards body publishes a numeric ship-it bar for health scores; the floors below are this skill's convention, set where a model becomes worth acting on and trustable by CSMs:

- Red band churns at **>= 3x the portfolio base rate** (precision lens).
- **>= 2/3 of churned accounts** sat below green 60-90 days pre-churn (recall lens).
- **< 80% of the book green** (distribution guard).

Iterate signals, weights, decay, and thresholds until all three hold. If they can't be met with available data, ship a rules-based score labeled as such (below) - never present an unvalidated composite as predictive.

## Cold start: too few churn events

Backtesting needs churned accounts to compare against; young or low-churn books don't have enough.

efficiency: rules-based proxy > statistically fitted score. Value runs the other way - the fitted score is the better model - which is exactly why the ratio, not the value, decides here.

- **Rules-based proxy.** An hour to a week: onboarding-milestone completion, activation/adoption thresholds, recency tripwires (no qualifying activity in N days), champion-contact recency - all readable off data the product already emits, and legible enough that a CSM can argue with a named rule instead of distrusting a black box. Label every weight a hypothesis. Buys a working at-risk list now, and no predictive claim.
- **Statistically fitted score.** A quarter of analyst time plus tens of churn events per segment, then a standing job to re-fit. Buys the better model: weights derived from what churned accounts actually did rather than from what someone expected.
- Promote fitting the moment tens of churn events exist per segment; until then resist tuning to individual anecdotes.
- When the book has no churned-account history at all, drop the fitted option from the menu rather than parking it at the bottom, and say so plainly - a ruled-out option left sitting in the list reappears as scope at the next planning round.

1. Still define the churn event, entity level, and prediction window now - so validation can start the moment events accumulate.
2. Treat the first 2-4 quarters as data collection. Log the score at fixed intervals so future backtests can retro-read it honestly.
3. Interview-validate meanwhile: monthly, ask experienced CSMs which accounts the score got wrong. Several named misses a week means the model is noise - fix the named causes.

## KPIs - is the score working?

| KPI                     | Definition                                                 | Healthy signal                                                                |
| ----------------------- | ---------------------------------------------------------- | ----------------------------------------------------------------------------- |
| Red-band churn multiple | Red-band churn rate / portfolio base rate                  | >= 3x, stable across quarters                                                 |
| Missed-churn rate       | Churned accounts that were green 60-90 days out            | Trending toward zero; each miss gets a post-mortem                            |
| Score distribution      | Band shares over time, per segment                         | Stable; green creep = decay or calibration failure                            |
| Play completion         | Band-change alerts that produced the wired play within SLA | High - otherwise it's a dashboard, not a score                                |
| Override rate + reasons | Overrides / accounts, clustered by reason code             | Low and explainable; a repeated reason code is a model bug filed by the field |
| Save/expansion outcomes | Red plays that retained; top-band plays that expanded      | The numbers to report upward - never report score coverage as success         |

## Recalibration cadence

- **Monthly for the first 90 days** after launch or major change: distribution by segment, CSM-named misses, play completion.
- **Quarterly standing.** Faster (monthly) for PLG/B2C - more volume, faster drift.
  - Rerun the backtest on the newest cohort.
  - Re-derive weights where separations shifted.
  - Retire signals that lost predictive power.
- **Off-cycle triggers:**
  - Product or pricing change.
  - ICP/segment shift.
  - Green-creep in the distribution.
  - A cluster of green-account churns.
  - Override rate climbing.

## Governance and overrides

- Diffuse ownership - every team contributes a metric, nobody owns the formula - is the classic decay path.
  - One owner for the logic (CS ops or RevOps).
  - One data owner for the pipeline.
  - CSMs as the acting owners.
- Overrides are a credibility mechanism, not a bug: a CSM who thinks "this is wrong, but whatever" has already downgraded the score to a report. Allow one-click overrides gated by a mandatory reason code; review override clusters at each recalibration as labeled disagreement data.
- Version every change with date, author, reason, and expected effect; announce to the CS team before it lands. Keep the per-account score breakdown visible where CSMs work - an unexplainable score loses the field faster than a wrong one.
references/weighting-and-bands.md
# Weighting, Normalization, Decay, Bands

## Deriving weights from outcomes

1. Pull the churned and retained/expanded cohorts over the trailing 12-24 months (or as far as data exists).
2. For each candidate signal, compare its behavior in the two cohorts over the prediction window. Keep signals with a real separation; weight proportionally to separation strength.
3. The working rule: if one signal separates churned from retained several times more strongly than another, equal weights make the score deliberately wrong. Weights are an empirical claim, not a design preference.
4. Normalize each kept signal to a common 0-100 scale with defined thresholds, then compute the composite as the weighted sum. Weights total 100.
5. Anti-pattern - the performative score: overweighting whatever is easy to measure (logins) and underweighting what is predictive but hard (champion presence, adoption depth). The 40/25/20/15-style splits circulating in vendor examples are illustrative starting hypotheses at best; every published source that offers one also says weighting is unique to each company.

## Caps and the champion rule

- Cap every category's share so no single input can move an account across a band alone. Convergence of two moderate signals should outweigh one strong one.
- Require at least one stakeholder-continuity input with enough downside weight to pull the band down by itself. This is the one deliberate asymmetry in the model: nothing may hold green after a champion departs, but no single positive signal may force green either.

## Decay

Every engagement-type signal drifts toward neutral as it ages. A score that never drifts accumulates false-positive greens: healthy six months ago, silent for weeks, still reading green.

Two implementations. efficiency: rolling windows > per-signal half-lives. Value runs the same way for a first score, so one line covers both axes.

- **Rolling windows.** Near-zero effort: one date filter per signal, computed wherever the score already runs, and a CSM can restate the rule in a sentence. Buys most of the decay benefit outright - green creep stops. A trailing 30-day window is a common practitioner default - a convention, not a researched constant; match the window to the book's sales/usage rhythm.
- **Per-signal half-lives.** An hour to fit, then a standing job: each half-life is a claim about that signal's shelf life that has to be re-checked at every recalibration. Buys smoother movement and a truer read where shelf lives genuinely differ - a sponsor meeting ages nothing like a login. Costs legibility: a CSM cannot reconstruct a decayed score by hand, and the trust that costs is not always bought back by the accuracy.
- Promote half-lives when window boundaries are themselves causing band flapping, or when one category's shelf life is obviously out of scale with the rest.
- Decay toward neutral, not toward red: silence is uncertainty, not confirmed risk. Let the trend component carry the alarm.
- Fit/firmographic-type inputs don't decay; only behavioral ones do.

## Trend and level, both

- Score movement is where the insight lives: a falling 75 is a worse account than a stable 60. Carry a trend component (delta over 1-2 periods) alongside the level, and display the score with a trend arrow and top risk driver.
- Guard against single-point noise: one bad month on one metric is noise more often than trend. Require two consecutive readings (or a trailing average) before a band change fires.

## Bands

- Three bands. Four or five sound precise and add confusion; if the team can't instantly decide what to do with a score, the bands aren't working.
- Size the red band to intervention capacity first, statistics second: if the team can work 30 accounts a month, a red band flagging 200 is a list nobody works. Then improve the model until the same capacity catches more of the actual churn.
- Every band has a wired play, an owner, and an SLA. A band with no play should not exist.
  - Top band: expansion/advocacy motion.
  - Middle: structured value intervention within a stated SLA.
  - Red: escalation with leadership review.
- B2B renewal proximity is a dimension of the score, not a filter after it: the same signals weigh heavier inside the renewal window, and a red account near renewal outranks a red account mid-contract for attention.

## Segment calibration (mandatory)

- One global model over unlike segments produces false negatives by construction: a 70 healthy for a mature enterprise account is a red flag for an SMB still onboarding. Low-touch segments run naturally lower engagement - shift weight from engagement to sentiment/support-responsiveness there.
- Override per segment: weights, thresholds, and expected signal ranges. Segment at minimum by size/touch model and by lifecycle stage (onboarding vs mature).
- Keep the segment count as small as the differences justify - each segment needs its own validation cohort.

## One score or two (expansion)

efficiency: one composite with a gating top band > a second dedicated expansion model.

- **One composite, three bands, top band gates expansion plays.** Near-zero marginal effort - it is the score already being built and validated, so it costs one model, one validation loop, and one number the field learns to read. The dominant practitioner pattern; a healthy top band is an expansion signal, not a rest state.
- **A second, dedicated expansion model.** A quarter to build and a standing job to keep: expansion-specific inputs (whitespace mapping, multi-team adoption spread, plan-limit trajectories) must be systematically tracked, and someone must maintain two validation loops instead of one. Buys better signal - the composite predicts churn and is a blunt proxy for growth.
- Promote it once those inputs are already tracked and the account base is large enough to validate a second model against its own outcomes. Below that it is a second unvalidated score, which is worse than none.
- No ranking on expected expansion lift, deliberately: no rigorous published study shows high health causally predicting expansion - only vendor case studies. Any ratio quoted for expansion payoff would be invented precision. Promise a better-qualified expansion queue, never expansion lift.
- The more common legitimate split is by lifecycle stage or segment - not by risk-vs-expansion outcome.

Health is necessary but not sufficient: the top band earns the expansion conversation. Qualify it on non-health conditions:

- Whitespace (unowned seats/products/teams).
- Fit.
- Plan-limit pressure: accounts pushing plan limits with strong engagement are the classic usage-derived expansion trigger.
- Budget and contract timing.
references/worked-examples.md
# Worked Examples

Illustrative specs with invented-but-realistic numbers. Every weight and threshold below must be re-derived from the user's own outcome data - copying these values verbatim recreates the performative-score failure this skill exists to prevent.

## Example 1: CSM-covered B2B (renewal-anchored)

```
HEALTH SCORE SPEC - Meridian Analytics (B2B data platform), 2026-03-10, v2
Outcome      : non-renewal at contract anniversary | account level | churn within 120 days
Segments     : Enterprise (CSM 1:15) | Mid-market (CSM 1:60) - separate weights and thresholds
Signals      : usage/adoption 35% - per-seat weekly active rate + core-capability breadth, trend
                 over 2 periods | evidence: declined in 11 of 14 churned accounts by day -90
               relationship 30% - exec-sponsor meeting recency, champion activity, stakeholder
                 breadth | evidence: 9 of 14 churned accounts lost sponsor contact by day -120
               support 15% - severity-weighted ticket trend per 100 seats, resolution-time trend
                 | evidence: escalation clusters preceded 6 of 14 churns
               commercial 10% - payment-timeliness drift, contraction requests | evidence: weak
                 separation; kept for the invoice-dispute tripwire
               sentiment 10% - survey trend + structured CSM rating (green requires exec-buyer
                 meeting < 60 days + >= 3 weekly-active users) | provisional - low response rate
Decay        : engagement/relationship inputs on trailing 30-day windows, drift to neutral;
               firmographics static
Rollup       : per-user rates weighted by role - champion 3x, admin 2x, end-user 1x; breadth
               floor: engagement concentrated in 1 user caps the relationship component at 60
Bands        : Green >= 70 | Yellow 45-69 | Red < 45 (mid-market: 65/40 - lower-touch norm)
               Red -> save play, CSM + manager, 5 business days; Yellow -> value review, 14 days;
               Green -> expansion review at next sync
Expansion    : Green for 2+ consecutive months AND (seat utilization > 80% OR new-team usage
               spread) -> expansion-qualified queue; AE qualifies on whitespace + budget
Overrides    : CSM one-click with reason code (sponsor-change | data-gap | context | other);
               clusters reviewed quarterly
Validation   : FY25 cohort, scored at day -90: red band churned 4.1x base rate; 71% of churned
               accounts below green at day -90; 62% of book green. Miss post-mortems: 4 green
               churns, all champion-departure cases -> v2 added the continuity input
Triggers     : band drop -> CRM task + owner alert; red inside 120-day renewal window -> weekly
               leadership review
Governance   : logic: RevOps manager | data: analytics eng | acting: CSM team | quarterly
               recalibration | change log in the RevOps runbook
```

## Example 2: PLG / B2C subscription (continuous)

```
HEALTH SCORE SPEC - Loopnote (prosumer note app, $9-19/mo self-serve), 2026-02-01, v1
Outcome      : cancellation or no reactivation within 30 days of expiry | subscriber level |
               churn within 60 days
Segments     : Individual | Team (2-10 seats, team plan gets a breadth component)
Signals      : usage recency/frequency 45% - RFM-style: days since last session, sessions/week
                 trend | evidence: recency decay preceded 78% of cancellations in H2 cohort
               depth 25% - core-action count per session, feature breadth | evidence: shallow
                 usage churned 2.9x deeper usage
               billing 20% - payment-failure events, dunning outcome, downgrade clicks |
                 evidence: soft-decline subscribers churned at 3.5x base
               engagement 10% - support/community/education touchpoints | provisional
Decay        : all behavioral inputs on trailing 21-day windows (fast product rhythm)
Rollup       : none for individuals; team plan adds seat-breadth (active seats / paid seats)
Bands        : Green >= 65 | Yellow 40-64 | Red < 40 - continuous evaluation, no renewal anchor;
               Red -> automated win-back sequence + dunning retry logic; Yellow -> lifecycle
               nudge (re-activation email, feature education); Green -> upgrade/annual-plan offer
Expansion    : Green + hitting plan limits (storage/seats) -> upgrade prompt; no human motion
Overrides    : none (no CSM layer); support agents can flag mis-scored subscribers, flags
               reviewed monthly as label data
Validation   : H2 cohort: red churned 3.8x base; 69% of churned subscribers below green at
               day -45 (60-90-day check compressed to the shorter lifecycle); 58% green
Triggers     : band drop -> lifecycle-automation event; payment failure -> immediate red review
Governance   : logic: growth PM | data: product analytics | monthly recalibration (volume
               supports it) | change log in the growth repo
```

## Example 3: a plausible-looking broken score, decomposed

The model: `health = logins 40% + NPS 20% + tickets 20% (more tickets = worse) + CSM gut rating 20%`, five bands, no decay, one global model. It looks reasonable and fails on every axis this skill checks:

- **Logins 40%, raw count**: big accounts read permanently green, small ones permanently sick; no trend, so a 50% decline from a high base stays green. Fix: per-seat rate, trend-scored, capped.
- **No relationship/continuity input**: a champion departure changes nothing until logins finally sag months later - the exact silent-churn miss. Fix: stakeholder-continuity input with band-down power.
- **Tickets scored as "more = worse"**: the most engaged accounts get punished; the disengaged account filing nothing reads healthy. Fix: rate-normalized trend, zero-ticket disengagement check.
- **CSM gut rating, unstructured**: relationship warmth overrides risk ("great call three weeks ago" holds green through an open escalation); ratings aren't comparable across CSMs. Fix: structured rating definitions plus gated overrides instead of a free-form input.
- **No decay**: every green accumulates; the book creeps toward >80% green and the score stops moving. Fix: drift-to-neutral windows.
- **Five bands, no plays**: nobody can say what band 2 vs band 3 means for action, so nobody acts. Fix: three bands, each wired to a play with owner and SLA.
- **Never backtested**: weights came from a workshop, not from churned-vs-retained separation - a performative score. Fix: run the backtest, re-derive, and hold it to the pass floor before trusting it.
SKILL.md
---
name: customer-health-score
description: Design, validate, and govern a composite customer health score - product usage and engagement signals combined into one weighted, normalized, decayed, banded account score that flags both churn risk and expansion readiness, backtested against real churn outcomes and recalibrated. Use whenever the user mentions a customer health score, health scoring model, account health, signal weighting, score bands, at-risk account flagging, churn risk score, expansion risk, or says "green accounts keep churning" - even if they never say "health score". Covers B2B account-level rollup and B2C/PLG continuous scoring. Do NOT use for discovering which indicators actually predict churn - use mbfinotti/revops-skills@customer-churn-signals instead.
license: MIT
metadata:
  author: Maya-Beth Finotti
  version: "1.0.1"
---

# Customer Health Score

Design, validate, and maintain one composite score that tells a post-sale team which accounts are drifting toward churn and which are ready for an expansion conversation. In ChurnZero's 2025 Customer Revenue Leadership Study (~800 CS and post-sales leaders surveyed), 73% said their current health score does not reliably predict churn.

- **Performative score.** Someone picks ten metrics, weights them by intuition; CS ignores the result because it predicts nothing.
- **Predictive score.** Derives its weights from what churned accounts actually did before churning, then proves itself in backtest.

This skill owns the composition and validation that turns one into the other:

- Normalization
- Weighting
- Decay
- Banding
- Segment calibration
- Override governance
- Wiring every band to a play

Which indicators to feed it, and how predictive each one is, is discovery work owned by `mbfinotti/revops-skills@customer-churn-signals`. Treat signals here at the category level (product usage, relationship/engagement, support, commercial, sentiment) and take the discovered indicator list as input.

Draw on the field's named canon where it genuinely helps:

- **Gainsight's DEAR framework** (Deployment, Engagement, Adoption, ROI): a balanced-scorecard model of 4-6 weighted metric groups summed into bands. Gainsight cautions that weighting "is highly unique to each company".
- **Lincoln Murphy's Success Potential**: ties expansion to achieved health ("No Customer Success... No Expansion Revenue").
- **RFM-style reads**: recency, frequency, and depth of usage.
- **Leading-vs-lagging split**: adoption trend, sponsor engagement, and response-time trend lead; churn, NRR, and renewal rate only confirm afterward. Score the leading ones.

One composite score, or a separate expansion model:

- **One composite, top band gates expansion (default).** Already the score being built and validated; a second model costs a quarter to build and two validation loops to keep. Gating expansion plays on the top band is the dominant practitioner pattern: low bands trigger retention, the top band triggers expansion/advocacy motions.
- **A separate expansion model.** Justified only when expansion-specific inputs are systematically tracked; otherwise it is a second unvalidated score.

Only vendor case studies support the claim that high health causally predicts expansion; no rigorous published study does. Treat health as necessary but not sufficient for expansion: the top band earns an account the conversation, and fit, whitespace, and budget decide it.

B2B and B2C/PLG share most of the mechanics: per-seat normalization, trend-over-level, an action wired to every band, and the backtest loop are identical in both. Say so when asked. What genuinely differs:

- **B2B.** Rolls user-level signals up to account level with extra weight on champion/sponsor continuity: renewals hinge on the buying organization, not one user, and flat aggregate usage can hide a disengaged champion (the "silent churn" pattern). Risk concentrates around a renewal date, so weight renewal proximity inside the score.
- **B2C/PLG.** No CSM and no renewal cliff: scoring is continuous and system-driven, ownership sits with growth/product rather than CS ops, and an involuntary-churn layer (failed payments, dunning) belongs in the score because payment failures correlate with engagement decline.

## Interview

Ask before designing. One question per message; offer multiple-choice answers when possible. Skip anything already answered.

- Which motion? (a) CSM-covered B2B with contract renewal dates, (b) PLG/self-serve continuous subscription, (c) B2C consumer subscription.
- How many accounts, and roughly how many churn events in the last 12-24 months? (Decides statistical weighting vs the rules-based cold-start path.)
- Which data sources are actually connected today - product usage data, engagement/meeting records, support system, billing? Rough coverage of each?
- Has signal discovery been done - is there an evidenced list of indicators that preceded past churn (see `mbfinotti/revops-skills@customer-churn-signals`), or only a gut list?
- Does a health score already exist? If yes: diagnose first or rebuild? (Diagnosis path: see Diagnosing a Broken Score.)
- What segments exist - enterprise vs SMB, onboarding vs mature? (A 70 that is healthy for enterprise is a red flag for an onboarding SMB; segment calibration is mandatory, not optional.)
- How many at-risk accounts can the team actually work per month? (The red band is sized to this, not to statistics.)
- Who owns the scoring logic, who acts on the score day-to-day, and who may override it?
- By what date must the score be live and acted on? (A hard date - a renewal cycle, a board review - promotes the rules-based proxy and the top two signal rungs; a distant date makes statistically fitted weights and a survey category affordable.)
- One-off win or compounding asset - this quarter's at-risk list, or a model that keeps getting better? (A one-off list drops decay, segment calibration, and the standing recalibration cadence; a compounding mandate promotes all three plus fitted weights.)
- What is the effort ceiling - analyst hours, whose data-engineering time, and how much manual logging CS will genuinely do? (Manual logging is what makes stakeholder-continuity signals real; if nobody will log, that category drops to whatever the CRM captures on its own.)

## Build order

Build the score rung by rung, ranked by churn actually prevented per hour of analyst time - not by what is cheapest to wire, and not by what scores best in backtest. Effort here is analyst hours, data the product must already emit, ongoing re-fitting, and how legible the score stays to the people who must act on it: a score CSMs cannot reconstruct is expensive no matter how accurate it is.

- efficiency: usage/adoption > commercial/billing > support > stakeholder continuity > sentiment survey
- value (churn caught): stakeholder continuity > usage/adoption > support == commercial/billing > sentiment survey
- effort:
  - usage/adoption and commercial/billing: near-zero - the product and billing system already emit the events.
  - support: an hour to map tickets to rates and trends.
  - stakeholder continuity: a standing job - humans must log meetings and sponsor changes for the input to stay true.
  - sentiment survey: a week to stand up, plus a standing job to run - low response rates keep it unstable.

Support and commercial tie on value because each catches a blind spot the other cannot - escalation clusters, invoice disputes - and neither substitutes for the other.

Start from the top two rungs plus whatever discovery already evidenced, validate that, then add rungs.

**What this order starves:**

- **Stakeholder continuity.** The highest-value input, the one that catches silent churn, and the worst ratio because it depends on humans logging contact. Promote it to rung one in any B2B book with renewal dates, or the moment one green account churns after a champion left - the continuity rule in step 5 makes it mandatory regardless of its ratio.
- **Statistically fitted weights and a second, dedicated expansion model.** Both lose every round on ratio and are still right eventually. Both need an account base large enough to fit and validate on its own outcomes (tens of churn events per segment), and both buy better signal once it exists.

This ordering is a default, not a law. It shifts with context and with who executes it - re-rank it against everything already known about the user before proposing it:

- A product already emitting rich usage telemetry locks usage at rung one for near-zero effort.
- A customer-success team of one demotes everything that needs manual logging and promotes automated usage and billing signals.
- Fewer than a hundred accounts takes fitted weights off the table entirely (see the cold-start path) and makes monthly CSM interview-validation the strongest check available.

## Workflow

1. Run the Interview; collect every answer before designing.
2. Define the outcome before any signal talk. A score that predicts an undefined event cannot be validated.
   - Churn event: non-renewal, cancellation, or no reactivation within N days for B2C.
   - Entity level: account for B2B, subscriber for B2C.
   - Prediction window: what the score claims to predict, e.g. churn within 90 days.
3. Take the indicator list as input - from the customer-churn-signals discovery work or, failing that, the user's best-evidenced candidates - and group it into the categories in [references/signal-categories.md](references/signal-categories.md). Audit data coverage per category, then build in the Build order ranking above, re-ranked against the user's answers. Never integrate everything at once, and drop any signal that could never trigger a play - if it can't trigger a play, it's a vanity signal.
4. Normalize per [references/signal-categories.md](references/signal-categories.md):
   - Rate-normalize by seats/users (500 logins from 50 users is not 500 logins from 1,000 users).
   - Read support and sentiment as trends, not counts.
   - In B2B, roll user-level signals up to account level with champion/sponsor weighting.
5. Weight per [references/weighting-and-bands.md](references/weighting-and-bands.md): derive weights from correlation with the defined outcome in the churned-vs-retained cohort, never from intuition. Cap every category so no single input can move an account across a band alone, and require at least one stakeholder-continuity input so no combination of usage signals can hold an account green after a champion departs.
6. Add decay: every engagement-type signal drifts toward neutral as it ages, so silence reads as uncertainty instead of accumulated green. Score the trend and the level - a falling 75 is a worse account than a stable 60.
7. Band: three bands, no more - more bands sound precise and add confusion. Size the red band to intervention capacity first, then improve the score until the same capacity catches more of the churn. Set segment-specific thresholds and weight overrides.
8. Validate per [references/validation-and-kpis.md](references/validation-and-kpis.md): retro-score the historical cohort, use precision/recall (never raw accuracy - churn is class-imbalanced), and run the retro check on what churned accounts scored 60-90 days pre-churn. Iterate weights, decay, and thresholds until the pass floor in the spec section clears. Too few churn events → take the cold-start path in the same reference.
9. Wire every band to a play with an owner and an SLA - a band change that triggers nothing is a report, not a health score. Define the override policy: any CSM may override with a mandatory reason code; overrides feed recalibration as labeled disagreements.
10. Emit the spec (next section), pilot with a small group before full rollout, review score distribution and misses monthly for the first 90 days, then move to the standing recalibration cadence.
11. If your harness has persistent memory, memorize the approved spec - categories, weights, bands, segment overrides, validation results, review date - so later recalibration and diagnosis runs start from it instead of re-interviewing.

## The Health Score Spec

Deliver every engagement as this artifact - a document the user can hand to CS ops, RevOps, and the CSM team. Fully worked versions live in [references/worked-examples.md](references/worked-examples.md).

```
HEALTH SCORE SPEC - <company/product>, <date>, v<n>
Outcome      : churn event definition | entity level | prediction window
Segments     : list + per-segment weight/threshold overrides
Signals      : category -> weight | inputs (from discovery) | normalization | evidence (outcome correlation)
Decay        : per-category drift-to-neutral rule and window
Rollup (B2B) : user->account aggregation method, champion/sponsor weighting
Bands        : 3 bands -> score ranges -> wired play, owner, SLA per band
Expansion    : top-band gate + the non-health conditions (fit, whitespace, contract timing)
Overrides    : who may override, reason codes, how overrides feed recalibration
Validation   : backtest window | red-band churn multiple vs base rate | % churned accounts
               below green at 60-90 days pre-churn | precision/recall at the red threshold
Triggers     : band-change alert -> task/workflow; renewal-proximity weighting (B2B)
Governance   : logic owner | acting owner | data owner | review cadence | change log location
```

Three rules about the spec itself:

- **Every weight carries its evidence.** "Adoption-depth trend 35%: churned accounts declined here in 7 of 9 pre-churn quarters analyzed" is reviewable. A bare percentage is folklore. If a weight has no outcome evidence yet, label it a provisional hypothesis with a validation date.
- **The spec has a pass threshold.** No standards body publishes a numeric ship-it bar for health scores; these floors are this skill's convention, set because a model that clears less will not survive live noise or earn CSM trust.
  - In backtest, the red band must churn at **at least 3x the portfolio base rate**.
  - **At least two-thirds** of accounts that actually churned must have sat below green 60-90 days before churning.
  - No more than **~80%** of the book may sit green (a documented practitioner red flag if crossed), independent of backtest results.

  Iterate until all three hold. Refuse to ship a score that reads as decoration.

- **The score is versioned and owned.** Diffuse ownership produces a score that changes meaning per reader: CS wants adoption, RevOps wants CRM activity, product wants feature usage, and everyone gets a field.
  - One owner for the logic (CS ops/RevOps).
  - CSMs act and override.
  - One data owner.
  - Every change logged and announced.

## Diagnosing a Broken Score

When the user arrives with a running score ("green accounts keep churning", "the team ignores it"), diagnose every symptom, then fix in the order below - not in the order the symptoms were reported, and not worst-fault-first. The ratio is churn actually prevented per hour, and the deadliest fault is rarely the one that pays back first.

- efficiency: wire band->play > normalize per seat > add decay > cap categories + continuity input > noise guard == ticket-read fix > segment calibration > re-derive weights from the cohort > qualify the expansion gate
- value: re-derive weights == wire band->play > cap categories + continuity input > normalize per seat == add decay > segment calibration > noise guard == ticket-read fix > qualify the expansion gate
- effort:
  - decay and the noise guard: near-zero.
  - normalization, caps, the ticket read, and the expansion gate: an hour each.
  - segment calibration and the cross-team coordination that wires plays to owners and SLAs: a week.
  - cohort re-derivation: a standing job, which also needs churned-account history the book may not have.

Ties justified:

- An unacted score and a wrongly-weighted one both prevent zero churn - they tie on value and split on effort.
- Normalization and decay tie because each removes one systematic misread of the same book.
- The noise guard and the ticket read are the same trailing-trend machinery applied twice, each buying back a different flavor of CSM distrust.

The table is ordered to match.

| Symptom                                             | Likely fault                                                                                     | Fix                                                                                              |
| --------------------------------------------------- | ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------ |
| CSMs ignore or silently distrust the score          | No wired plays; no gated override; unexplainable weights                                         | Wire band->play; allow overrides with reason codes; publish the weight evidence                  |
| Large accounts always look sick (or always healthy) | Raw counts instead of rates                                                                      | Normalize per seat/user; benchmark against similar-size accounts                                 |
| Score rarely moves; most of book green              | No decay - old positive signals accumulate forever                                               | Decay engagement signals toward neutral; recheck distribution monthly (>80% green = broken)      |
| Champion left; account stayed green for weeks       | Single-category dominance - usage alone can hold green                                           | Cap category shares; require a stakeholder-continuity input that can pull the band down alone    |
| Score flaps band-to-band weekly                     | Single-point noise treated as signal                                                             | Score trailing-window trends; require two consecutive readings before a band change fires        |
| High-touch accounts flagged red on ticket volume    | Ticket count read as risk                                                                        | Rate-normalize tickets, read the trend; treat zero tickets as possible disengagement, not health |
| Same score means different things across accounts   | One global model over unlike segments                                                            | Segment-specific weights and thresholds (enterprise vs SMB, onboarding vs mature)                |
| Green accounts keep churning                        | Performative weights - intuition, not outcome correlation; missing relationship/sentiment inputs | Re-derive weights from the churned-vs-retained cohort; add a stakeholder-continuity input        |
| Top-band accounts never expand                      | Health treated as sufficient for expansion                                                       | Gate expansion plays on top band, but qualify on fit, whitespace, and contract timing separately |

Re-rank this too: a book with no churned-account history has no cohort to re-derive from, so that rung leaves the list rather than sitting last (see the cold-start path).

## Reference

- See [references/signal-categories.md](references/signal-categories.md) for how each signal category behaves inside a composite: normalization, trend windows, rollup, and per-category traps. Category mechanics only; indicator discovery lives in `mbfinotti/revops-skills@customer-churn-signals`.
- See [references/weighting-and-bands.md](references/weighting-and-bands.md) for deriving weights from outcomes, caps, decay mechanics, banding, segment overrides, and the one-score-vs-two decision in full.
- See [references/validation-and-kpis.md](references/validation-and-kpis.md) for the backtest procedure, precision/recall framing, pass floors, the cold-start path, KPIs, recalibration cadence, and override governance.
- See [references/worked-examples.md](references/worked-examples.md) for a CSM-covered B2B spec and a PLG/B2C spec worked in full, plus a plausible-looking broken score decomposed.
- See `mbfinotti/revops-skills@lead-scoring` for pre-sale scoring - lead scoring predicts who will buy; this skill predicts who will stay and grow. Different outcome, different model.
- See `mbfinotti/revops-skills@sales-to-cs-handoff` for the sales-to-CS transition where the first health baseline gets set.
- See `mbfinotti/revops-skills@revenue-kpi-framework` for the org-wide metric tree this score reports into - the composite can serve as a leading input under the retention branch, and that skill sets which level owns it and what guardrail rides alongside it.