SKILL DETAIL
customer-churn-signals
mbfinotti/revops-skills/customer-churn-signals
Assemble, define, validate, and rank leading indicators of customer churn from account activity, product usage, and support data into a ranked churn signal register - each signal carrying event, threshold, window, baseline, lift over base rate, lead time, and coverage. Use whenever the user mentions churn signals, churn indicators, leading indicators of churn, an early warning system for at-risk accounts, champion departure, seat utilisation drops, payment failure, or says "our health score misses churn" - even if they never say "signal". Covers B2B and B2C subscription. Do NOT use for combining signals into one composite score, weights, bands, or tiers - use mbfinotti/revops-skills@customer-health-score instead.
Installation
npx skills add https://github.com/mbfinotti/revops-skills --skill customer-churn-signals
Fichiers du skill
SKILL.md
Dernière synchronisation · 15 sept. 2026
evals/evals.json›
{
"skill_name": "customer-churn-signals",
"evals": [
{
"id": 1,
"prompt": "I run RevOps at Tellwright — B2B compliance platform, 640 active accounts, annual contracts, 52 cancellations in the last 18 months. Our marketing analyst went through every account we lost and found that 68% of them had a declining marketing-email open rate across their final 90 days. Nothing else came back anywhere near that high, so we want \"open rate down three months running\" to be the headline alert on the new at-risk dashboard. Two things worth mentioning: whenever a CSM thinks an account is wobbling they fire a re-engagement email sequence at it, and we cut total send volume roughly in half in March when we switched email platforms. Can you help me write this up properly before I take it to our VP?",
"expected_output": "A rejection of the open-rate signal that names the coverage/lift confusion, computes the base rate, demands the retained-account denominator, and identifies the vendor-artefact, baseline, and actionability failures.",
"files": [],
"expectations": [
"Identifies the 68% figure as coverage, not lift, because it was computed only on accounts that churned",
"States that lift requires the non-churned group as a denominator: churn rate among signal-showing accounts divided by the base rate",
"States or computes the book's base rate from 52 churns out of 640 accounts (approximately 8%) and uses it as the comparison point",
"Asks the user to pull, or instructs them to pull, the share of retained accounts that also showed a declining open rate",
"Refuses to ship the open-rate signal as the headline alert on the strength of the 68% figure",
"Flags the CSM re-engagement sequence as a vendor-outreach artefact that inflates measured engagement, and requires counting customer-initiated activity only",
"Names the mid-period send-volume halving as a denominator or baseline change that makes the decline non-attributable to the customer",
"Notes that email open rate has no honest baseline because inbox privacy features inflate or mask opens independently of the customer's behaviour",
"Notes the signal fails the actionability test because the intervention available is email, the very channel being ignored",
"Applies a lift floor of at least 2x the base rate before any signal earns a ship verdict",
"Proposes replacement candidates drawn from customer-initiated activity in sources the vendor cannot pollute"
]
},
{
"id": 2,
"prompt": "Hi — new director of CS ops at Varenholm Systems. B2B enterprise, average contract $180k, annual and two-year terms, 31 accounts lost in the past 24 months out of 410. Leadership wants an early-warning system live before the next board meeting. I've drafted the trigger list already: invoice unpaid past 15 days, customer asks for a seat reduction, procurement pushes the renewal date, someone runs a full data export, auto-renew switched off in the admin panel, and any visit to the billing or cancellation page. I ranked them by strength — data export came out on top, failed invoice second. What am I missing?",
"expected_output": "Every drafted trigger reclassified as confirmatory, the register diagnosed as containing zero leading signals, enterprise lead-time and reconstruction windows applied, and genuine leading candidates proposed.",
"files": [],
"expectations": [
"Classifies all six drafted triggers as confirmatory tripwires rather than leading indicators",
"States plainly that the drafted register contains zero leading signals",
"Explains that confirmatory tripwires formalize a decision the customer already made weeks earlier",
"Keeps the tripwires in the register, flagged as confirmatory, as a triage-urgency input rather than deleting them",
"States that confirmatory tripwires are exempt from the lead-time bar but must carry the confirmatory flag",
"Declines to rank the tripwires against one another, on the grounds that each costs near-zero to instrument and buys zero lead time, so any ordering is false precision",
"Applies a median lead-time bar of 90 days for this enterprise annual-contract motion rather than the 30-day default",
"Reconstructs candidate signal state at 120-180 days pre-churn for this enterprise motion, not only at 30/60/90 days",
"Sets a register-level floor requiring the shipped set to have fired on at least 70% of past churned accounts at minimum lead time",
"Applies a per-signal coverage floor of at least 25% of past churned accounts",
"Proposes actual leading candidates such as login recency, seat utilisation, support-pattern, or champion departure",
"Notes that the annual and multi-year contract terms delay the visible churn event, so behavioural signals run far ahead of the commercial one"
]
},
{
"id": 3,
"prompt": "Kestrelbrook, seed-stage B2B SaaS, 280 paying accounts. We've had 22 accounts cancel in the last 18 months. My CEO saw a conference talk and now wants \"a churn prediction model\" — logistic regression at minimum, ideally gradient boosting, and she wants to see the accuracy number in the deck. I'm the only analyst here, I have maybe 15 hours of spreadsheet time, and honestly once this ships nobody is retraining anything. Renewals open in 7 weeks and she wants it in front of the team before then. Where do I start?",
"expected_output": "Regression, Cox, and WoE/IV deleted from the menu with the specific rule cited for each, backtest-and-lift recommended as sufficient, and accuracy rejected as the validation metric.",
"files": [],
"expectations": [
"Deletes logistic regression and Cox from the method menu outright rather than demoting, deferring, or parking them as a stretch goal",
"Cites the absence of anyone to refit the model as an independent reason to delete logistic regression and Cox at any sample size",
"Deletes Weight of Evidence / Information Value as well, because the churned sample of 22 falls below the roughly 50-churn threshold",
"States the approximate 50-churned-account threshold below which WoE/IV and regression become inadmissible",
"Deletes every method below retention curves because the 7-week deadline falls inside the next renewal cycle",
"Recommends backtest plus churn-rate lift as the first method and, here, the sufficient one",
"States the method order as backtest and lift, then retention curves, then WoE/IV, then logistic regression and Cox, stopping when the register clears",
"Labels the resulting backtest numbers as directional given the sample size",
"Rejects accuracy as the validation metric, explaining that churners are a rare class so predicting nobody churns scores high",
"Names precision, recall, and PR-AUC as the metrics to use if any scoring model ever exists",
"Recommends k-fold cross-validation over a single train/test split on a sample this small",
"Does not deliver or recommend gradient boosting as the output of this engagement"
]
},
{
"id": 4,
"prompt": "Norsholt Analytics — B2B, 1,100 accounts, mix of monthly and annual. In our last CS meeting the team shouted out four things they say they see before an account leaves: logins go down, they stop using the reporting module as much, the seat count looks off, and NPS drops. I want to turn those four into something I can actually query next week. Worth knowing: about a fifth of our book grew headcount two to three times this year, and a handful downsized hard after layoffs. Our usage also dips every northern-hemisphere summer.",
"expected_output": "Each of the four rewritten into a five-part operational definition with the correct baseline flavour per signal, seat and seasonality normalization applied, and NPS demoted.",
"files": [],
"expectations": [
"Rewrites each of the four into an operational definition rather than accepting any as stated",
"Names all five required parts of an operational definition: event, threshold, observation window, baseline, and normalization",
"States that a candidate missing any of the five parts is not testable and must be fixed or dropped",
"Uses the account's own trailing history as the baseline for the decline signals (login frequency, reporting-module usage)",
"Uses segment peers as the baseline for the adoption-depth reading, to catch accounts that never activated rather than only accounts that declined",
"Normalizes the relevant candidates per licensed or provisioned seat",
"Explains that without seat normalization growing accounts look healthier and downsizing accounts look sicker than they are, tying this to the headcount growth and the post-layoff downsizing described",
"Applies a seasonality adjustment for the recurring summer usage dip",
"Defines the seat candidate as active seats divided by provisioned seats, trended, and calls out reconciling what counts as active against the provisioning record",
"Ranks the NPS candidate below product engagement, citing survey response bias against the quietly disengaging",
"Defines the NPS candidate as a score trend across waves rather than a single reading, or treats non-response after prior participation as its own candidate"
]
},
{
"id": 5,
"prompt": "I'm at Dunmarrow, PLG B2B — 3,400 workspaces, mostly monthly self-serve with an annual cohort. Our product has written a per-account event for every meaningful action into our warehouse for two years already. My VP wants the at-risk list to be a compounding thing we keep sharpening, not a one-quarter fire drill, and our annual cohort doesn't renew for another 11 months. His instruction was: start with the cheap stuff — CRM activity, billing records, the login table — and bolt product data on later once we've proved the concept. We've got a data engineer with spare cycles. What order should we build in?",
"expected_output": "The VP's cheap-first order overturned, product usage and feature-adoption breadth promoted to the front with the three promotion conditions named, and the privacy review flagged.",
"files": [],
"expectations": [
"Promotes product usage and feature-adoption breadth to the front of the build order instead of leaving them at their default efficiency positions",
"Gives the already-existing per-account event stream as the reason the instrumentation effort collapses to near-zero and the efficiency ratio inverts",
"Gives the compounding-asset mandate as an independent reason to promote the telemetry categories",
"Notes that the 11-month runway to the annual renewal leaves room for instrumentation to land before the next renewal window opens",
"Contradicts the VP's cheap-first instruction explicitly rather than complying with it",
"Explains that a register built strictly by value-per-effort ratio ends up entirely CRM-derived and lagging relative to what the customer actually does in the product",
"States that product usage and feature-adoption breadth carry the highest lift and the longest lead times of the categories",
"States that product usage and feature-adoption breadth never separate on effort because both are gated on the same per-account event stream",
"Names the in-house data engineer as a factor that collapses the telemetry effort estimate",
"Flags that product usage, feature-adoption breadth, champion tracking, and customer business context each require a privacy review before the first event is stored, and that the other categories carry no such cost",
"Notes that the compliance exposure is one-way: events collected without a lawful basis must be deleted, not relabelled"
]
},
{
"id": 6,
"prompt": "Meadowlark is a consumer meditation app — $8.99 a month, 47,000 active subscribers, we lost about 2,300 last quarter. My previous job was enterprise software and I still have the at-risk playbook from there, so I was going to port it across: champion departure, exec sponsor engagement, QBR attendance, buying-committee coverage, plus login recency and NPS. Our CS team of three can work maybe 40 flagged subscribers a week. How should I adapt it?",
"expected_output": "The B2B relationship categories deleted outright, billing promoted for the consumer book, compressed lead times handled as a documented deviation, and the shipped set sized to the stated capacity.",
"files": [],
"expectations": [
"Deletes champion departure, exec sponsor engagement, QBR attendance, and buying-committee coverage outright rather than ranking them low or parking them for later",
"Gives the absence of an internal buying committee in a consumer subscription as the reason those categories cannot apply",
"Warns that a ruled-out category left in the list reappears as scope in a later session",
"Promotes billing and payment behaviour on this consumer book, stating that payment failure is a proportionally larger churn driver in B2C than in B2B",
"Recommends catching card or payment failure pre-emptively rather than only once the payment has already failed",
"States that B2C lead times compress because cancellation is a one-person, one-tap decision with no procurement cycle",
"Treats any relaxation of the 30-day lead-time floor for this book as a deliberate, documented deviation rather than a silent change",
"Names churn by indifference — a forgotten low-cost subscription — as a B2C-specific pattern with no B2B analogue",
"Keeps usage and login decline and feature-adoption breadth with an identical method, changing only the unit of analysis to the individual subscriber",
"Sizes the shipped signal set against the stated capacity of 40 flagged subscribers a week, recording the false-alarm side (signal-showing subscribers who retained) per signal",
"Notes that the churned sample of roughly 2,300 is large enough that no ranking method is deleted on sample-size grounds"
]
},
{
"id": 7,
"prompt": "Cortholm, B2B SaaS, 900 accounts, $2.4M ARR. Can you build me a customer health score? I want every account scored 0-100 with red/amber/green bands, and I'm thinking weights along the lines of usage 40%, support 20%, NPS 20%, billing 20%. Then a playbook for the CSMs on what to do when an account goes red. We have 12 months of churn history, 63 accounts lost. Nothing has been validated — we picked those four inputs because they felt right.",
"expected_output": "Composite scoring, weights, bands, and CSM playbooks routed elsewhere; a validated ranked signal register delivered instead, with the four proposed inputs backtested against the 63 known churns.",
"files": [],
"expectations": [
"Declines to set the weights, the 0-100 scale, or the red/amber/green bands as part of this engagement",
"Routes composite scoring, weighting, and banding to the customer-health-score skill by name",
"Routes the CSM playbooks and save motions out of scope as well",
"Delivers a ranked, validated churn signal register as the endpoint of this engagement instead",
"Refuses to carry the four proposed inputs forward unvalidated and backtests them against the 63 known churned accounts first",
"Applies the per-signal floors of at least 2x lift, at least 30 days median lead time, and at least 25% coverage",
"Applies the register-level floor of at least 70% of past churned accounts caught at minimum lead time",
"Names the handoff of the shipped signals to composite scoring as the final step of this engagement",
"Recommends a deliberately small shipped set of roughly four to six signals rather than tracking everything measurable, crediting Gainsight",
"Recommends segmenting by account tier, because health looks different for SMB and enterprise accounts",
"Notes that the 63-churn sample clears the roughly 50-churn threshold, so WoE/IV is admissible on this book"
]
},
{
"id": 8,
"prompt": "Braithsend, B2B SaaS. We've had a health score for two years and it's useless — accounts sit green and then cancel on us. 58 churns in the last 20 months out of 720 accounts. I ran a WoE/IV analysis on \"days since last admin login over 45\" and got an IV of 1.8, way past the 0.3-0.5 strong band — easily the best number in the whole analysis, so I want to make it the backbone of the rebuilt score. Separately, our support lead swears that ticket volume going up is the clearest tell we have.",
"expected_output": "The IV of 1.8 read as a leakage warning rather than strength, the survivorship check run on the existing score's inputs, and raw ticket volume rejected as one-way.",
"files": [],
"expectations": [
"Reads the IV of 1.8 as a leakage warning rather than as unusually strong predictive evidence",
"Explains that an IV far above the interpretation bands usually means the variable is a consequence of a churn decision already made — a tripwire rather than a leading indicator",
"Recommends re-binning the days-since-login variable more tightly, with finer bins below 30 days, and checking the measured lead time before shipping it",
"Runs the survivorship check by asking, for each of the 58 churned accounts, whether the candidate was actually firing 60-90 days before the churn",
"States that a signal still green at 60-90 days pre-churn on most churned accounts is a blind spot, not a predictor",
"Backtests the existing health score's own input signals against the real churn outcomes rather than rebuilding the score blind",
"Reports the IV interpretation bands (under 0.02 not useful, 0.02-0.1 weak, 0.1-0.3 medium, 0.3-0.5 strong) as credit-scoring convention rather than law",
"Notes that the 58-churn sample sits right at the roughly 50-churn boundary, so the WoE/IV output is directional",
"Rejects raw ticket volume as a one-way signal, noting that engaged customers file more tickets while silently disengaging ones file none",
"Defines silence after a ticket spike as its own separate candidate signal",
"Requires pairing ticket volume with severity mix, unresolved-ticket age, and sentiment trend rather than shipping volume alone"
]
},
{
"id": 9,
"prompt": "Hallowbridge Software — B2B, 380 accounts, annual contracts averaging $60k, 44 churns over the last two years. I've backtested five candidates and here is what came back. (a) failed payment: lift 6.1x, median 7 days lead, 18% coverage. (b) admin login gap of 30+ days: lift 2.4x, 61 days lead, 44% coverage. (c) named champion leaves the company: lift 4.3x, 51 days lead, 29% coverage. (d) support tickets triple then stop: lift 2.8x, 55 days lead, 24% coverage. (e) distinct core features used drops from 4 to 1: lift 3.9x, 88 days lead, 38% coverage — but we'd need engineering to build the event pipeline, about a quarter of work. My plan is to ship the top two by lift and move on. Our two CSMs can chase maybe 10 accounts a week.",
"expected_output": "The lift-only ranking rejected, each candidate verdicted against lift, lead time, coverage and effort, the tripwire flagged, and register-level coverage checked against the 70% floor.",
"files": [],
"expectations": [
"Rejects the plan to ship the top two by lift",
"Ranks on the triple of lift, lead time, and coverage together rather than on lift alone",
"Flags candidate (a) failed payment as confirmatory, exempt from the lead-time bar but carrying the confirmatory flag, and not counted among the leading signals shipped",
"States that a high-lift signal with roughly 7 days of lead time is un-actionable as an early warning",
"Rejects candidate (d) on the coverage floor, because 24% falls below the 25% minimum",
"Applies the 90-day lead-time bar for this annual-contract motion, or explicitly justifies relaxing it, rather than silently applying the 30-day default",
"Divides the value triple by what each candidate costs to stand up, and records candidate (e)'s quarter of engineering work in the register's effort line",
"Does not reject candidate (e) purely on cost, weighing its lift, 88-day lead time, and 38% coverage against the register-level floor",
"Computes or states the combined register-level coverage of the proposed shipped set and tests it against the 70% floor",
"Sizes the resulting flag volume against the stated capacity of 10 accounts a week, reporting the false-alarm side of each signal",
"Notes that the recorded-champion field needs a standing maintenance job, because a stale field silently produces a signal that never fires",
"Notes that champion departure is typically discovered at the renewal call rather than when it happens, so the detection itself must be instrumented"
]
}
],
"trigger_queries": [
{ "query": "help me build an early warning system for at-risk B2B accounts", "should_trigger": true },
{ "query": "what are the leading indicators of churn?", "should_trigger": true },
{ "query": "our health score misses churn entirely", "should_trigger": true },
{ "query": "which signals actually predict that a customer is about to leave", "should_trigger": true },
{ "query": "how do I know an account is at risk before it's too late", "should_trigger": true },
{ "query": "we want to spot wobbling accounts earlier than we do now", "should_trigger": true },
{ "query": "rank our churn indicators by how predictive they actually are", "should_trigger": true },
{ "query": "I need to validate whether our at-risk triggers work at all", "should_trigger": true },
{ "query": "backtest our risk flags against the accounts we already lost", "should_trigger": true },
{ "query": "does champion departure really predict non-renewal?", "should_trigger": true },
{ "query": "seat utilisation is sliding across the book — is that a churn signal?", "should_trigger": true },
{ "query": "we keep getting blindsided by cancellations, what should we be watching", "should_trigger": true },
{ "query": "build me a churn signal register", "should_trigger": true },
{ "query": "what should we measure to catch churn 90 days out", "should_trigger": true },
{ "query": "how much lead time does declining product usage actually buy us", "should_trigger": true },
{ "query": "is a failed payment a leading or a lagging churn indicator", "should_trigger": true },
{ "query": "our CSMs say they can feel when an account is going to leave — I want that measurable", "should_trigger": true },
{ "query": "what's the lift over base rate on days-since-last-login as a predictor", "should_trigger": true },
{ "query": "define the threshold and observation window for a usage decline alert", "should_trigger": true },
{ "query": "we have 60 accounts that cancelled last year — what can I learn from them", "should_trigger": true },
{ "query": "which of these at-risk triggers are actually worth shipping", "should_trigger": true },
{ "query": "help me work out what precedes cancellation in our own data", "should_trigger": true },
{ "query": "early warning indicators for subscription cancellation", "should_trigger": true },
{ "query": "our renewals keep falling through and we never see it coming", "should_trigger": true },
{ "query": "how do I tell a real predictor from a coincidence", "should_trigger": true },
{ "query": "should NPS be part of our at-risk detection?", "should_trigger": true },
{ "query": "we use ticket volume as a risk flag, is that right", "should_trigger": true },
{ "query": "how far ahead of cancellation does narrowing feature adoption show up", "should_trigger": true },
{ "query": "what data do I need to predict churn without a data science team", "should_trigger": true },
{ "query": "set up leading indicators for our consumer subscription app", "should_trigger": true },
{ "query": "b2c subscription — what tells me someone is about to cancel", "should_trigger": true },
{ "query": "which usage metrics matter for retention risk", "should_trigger": true },
{ "query": "I want to stop being surprised by the renewals we lose", "should_trigger": true },
{ "query": "how do you measure whether a risk flag fires early enough to act on", "should_trigger": true },
{ "query": "work out coverage and lift for each of our warning flags", "should_trigger": true },
{ "query": "we think login drop-off predicts cancellation, how do I prove it", "should_trigger": true },
{ "query": "what's a good early indicator for an annual contract business", "should_trigger": true },
{ "query": "identify the behaviours that come before a customer leaves us", "should_trigger": true },
{ "query": "put together the list of things we should monitor so we catch losses early", "should_trigger": true },
{ "query": "can you help me figure out why we can't see cancellations coming", "should_trigger": true },
{ "query": "how many days before cancellation does a support escalation typically show up", "should_trigger": true },
{ "query": "validate our warning triggers against real cancellations from last year", "should_trigger": true },
{ "query": "we have product event data — what risk indicators can we build from it", "should_trigger": true },
{ "query": "what counts as a testable definition for a usage-decline alert", "should_trigger": true },
{ "query": "which of our warning flags are just confirming a decision already made", "should_trigger": true },
{ "query": "help me pick the handful of things to watch that catch most of our losses", "should_trigger": true },
{ "query": "combine our churn signals into a single 0-100 health score", "should_trigger": false },
{ "query": "what weights should each health score input carry", "should_trigger": false },
{ "query": "set the red, amber and green bands on our account health score", "should_trigger": false },
{ "query": "how often should we recalibrate the customer health score", "should_trigger": false },
{ "query": "apply score decay so old activity stops propping up the number", "should_trigger": false },
{ "query": "write the CSM playbook for when an account turns red", "should_trigger": false },
{ "query": "design a save offer for subscribers who hit the cancel button", "should_trigger": false },
{ "query": "build a win-back campaign for customers who already left", "should_trigger": false },
{ "query": "write the cancellation flow copy for our app", "should_trigger": false },
{ "query": "set up dunning emails for failed card payments", "should_trigger": false },
{ "query": "design our lead scoring model", "should_trigger": false },
{ "query": "sales keeps rejecting our MQLs, fix the scoring", "should_trigger": false },
{ "query": "what should the MQL threshold be set to", "should_trigger": false },
{ "query": "route inbound leads to the right rep", "should_trigger": false },
{ "query": "our round robin is lopsided, some reps get nothing", "should_trigger": false },
{ "query": "audit our pipeline stage exit criteria", "should_trigger": false },
{ "query": "our stages don't mean anything, rewrite them", "should_trigger": false },
{ "query": "design our revenue funnel stages from scratch", "should_trigger": false },
{ "query": "build a bowtie funnel model for the plan", "should_trigger": false },
{ "query": "why did we miss the forecast this quarter", "should_trigger": false },
{ "query": "our reps are sandbagging the commit number", "should_trigger": false },
{ "query": "clean up the stale deals in the pipeline before the QBR", "should_trigger": false },
{ "query": "find where revenue is leaking out of our funnel", "should_trigger": false },
{ "query": "which system is the source of truth for ARR", "should_trigger": false },
{ "query": "finance and sales report two different ARR numbers", "should_trigger": false },
{ "query": "who owns each field in our CRM", "should_trigger": false },
{ "query": "our CRM fields contradict each other", "should_trigger": false },
{ "query": "build the KPI tree for our revenue org", "should_trigger": false },
{ "query": "what should our North Star metric be", "should_trigger": false },
{ "query": "explain how NRR is calculated and who owns the branch", "should_trigger": false },
{ "query": "write the revenue section of the board deck", "should_trigger": false },
{ "query": "our investor update reads like a data dump", "should_trigger": false },
{ "query": "design the discount approval matrix for non-standard deals", "should_trigger": false },
{ "query": "who has to sign off on a 30% discount", "should_trigger": false },
{ "query": "we have too many GTM tools, what do we cut", "should_trigger": false },
{ "query": "design the handoff packet from sales to CS after closed-won", "should_trigger": false },
{ "query": "our CSMs start from zero after the deal closes", "should_trigger": false },
{ "query": "which revops newsletters and podcasts should I follow", "should_trigger": false },
{ "query": "how do I break into RevOps from a sales role", "should_trigger": false },
{ "query": "write the scorecard for our first RevOps hire", "should_trigger": false },
{ "query": "which revops skill do I need for this project", "should_trigger": false },
{ "query": "build a monthly cohort retention chart for the investor deck", "should_trigger": false },
{ "query": "forecast next quarter's churn rate for the financial model", "should_trigger": false },
{ "query": "write an NPS survey and pick the question wording", "should_trigger": false },
{ "query": "segment our customers for a lifecycle email program", "should_trigger": false },
{ "query": "help me negotiate with an account that says it wants to leave", "should_trigger": false }
]
}
references/examples.md›
# Worked Examples
Every figure below is invented to illustrate the computations. Never quote them as benchmarks or defaults - the user's own backtest replaces all of them.
## Worked lift computation
Book: 520 active B2B accounts; 46 churned in the trailing 12 months; base rate 8.8% per 12-month window.
Candidate: "weekly active users down >= 40% vs the account's own trailing-90-day median, sustained 3 consecutive weeks."
- Fired on 61 accounts during the period. Of those, 19 churned within the following 90 days: 19/61 = 31.1%.
- Lift = 31.1 / 8.8 = **3.5x** base rate. Passes a 2x floor.
- Coverage = 19 of 46 churns fired it = **41%**.
- Median days from first fire to churn across those 19 = **74 days** lead time.
- False-alarm note: 42 of 61 firing accounts retained - at ~5 flags/month this fits a CSM team that can absorb 8/week.
- Verdict: **ship**.
## Worked WoE/IV computation
Candidate variable: days since last admin login, binned. Same book: 46 churned, 474 retained - right at the ~50-churn boundary where WoE/IV becomes admissible at all, so every number below is directional.
| Bin | Churned (share) | Retained (share) | WoE = ln(c/r) | (c - r) x WoE |
| ---------- | --------------- | ---------------- | ----------------------- | ---------------------- |
| 0-13 days | 8 (17.4%) | 322 (67.9%) | ln(0.174/0.679) = -1.36 | (-0.505)(-1.36) = 0.69 |
| 14-29 days | 14 (30.4%) | 108 (22.8%) | ln(0.304/0.228) = 0.29 | (0.076)(0.29) = 0.02 |
| 30+ days | 24 (52.2%) | 44 (9.3%) | ln(0.522/0.093) = 1.73 | (0.429)(1.73) = 0.74 |
IV = 0.69 + 0.02 + 0.74 = **1.45** - far above the 0.3-0.5 "strong" band, which is itself the finding: at 30+ days of admin silence this is drifting from leading indicator toward tripwire. Re-bin tighter (0-6 / 7-13 / 14-29 / 30+) and check the lead time before shipping; the 14-29 bin alone is the honest early-warning zone.
## B2B register (abbreviated)
```
CHURN SIGNAL REGISTER - illustrative B2B book, v1
Base rate : 8.8% per 12 months; churned sample = 46 (24 months)
Method : backtest + lift + WoE/IV (46 churns - boundary call, read as directional)
regression/Cox deleted: nobody to refit the model after ship
Rows : ordered by value per unit of instrumentation effort
signal : active/provisioned seats < 40% for 60d
category : seats | lift 2.6x | lead 95d | coverage 28% | effort near-zero | verdict SHIP
signal : ticket spike (>=3x own median) then zero tickets for 30d
category : support | lift 2.9x | lead 52d | coverage 26% | effort an hour | verdict SHIP
signal : recorded champion departs (job-change event)
category : relationship | lift 4.1x | lead 48d | coverage 33% | effort a standing job | verdict SHIP
confidence: small n (15 departures observed) - directional
signal : WAU decline >=40% vs own 90d median, 3 consecutive weeks
category : usage | lift 3.5x | lead 74d | coverage 41% | effort near-zero | verdict SHIP
confidence: event stream already existed, so the highest-value signal was also cheap -
on a book without one this drops below the cheap categories
signal : failed payment
category : billing | lift 6.0x | lead 9d | coverage 15% | effort near-zero | verdict CONFIRMATORY
confidence: exempt from lead-time bar; escalation trigger only
signal : NPS drop >=3 pts between waves
category : survey | lift 1.6x | lead 80d | coverage 17% | effort near-zero | verdict WATCH
confidence: response bias - 61% of churned accounts never answered the wave
Register-level : shipped set fired on 76% of past churns at >=30d - clears 70% floor
Handoff : -> mbfinotti/revops-skills@customer-health-score
```
## B2C register (compact)
```
CHURN SIGNAL REGISTER - illustrative B2C subscription app, v1
Base rate : 5.1% per month; churned sample = 1,240 (12 months)
session recency > 14d (vs subscriber's own weekly habit) : usage
lift 3.2x | lead 21d | coverage 58% | effort near-zero | SHIP (B2C lead times compress
- 30d floor relaxed to 14d for this book, documented as a deliberate deviation)
card approaching expiry with no update : billing
lift 2.8x | lead 30d | coverage 22% | effort near-zero | WATCH (billing climbs on a
consumer book - pre-emptive here - but 22% misses the coverage floor)
core-feature breadth drops to 1 of 4 within a month : adoption
lift 2.4x | lead 26d | coverage 31% | effort a week | SHIP
cancellation-page visit : behavior
lift 9x | lead 2d | coverage 44% | effort near-zero | CONFIRMATORY
Champion and committee categories deleted, not ranked: no B2B equivalent on this book.
```
## A plausible-looking bad signal, decomposed
Proposed: "marketing-email open rate declining -> churn risk." Rejected on four grounds:
1. **No honest baseline** - inbox privacy features inflate or mask opens per mail client, so the metric moves for reasons unrelated to the customer; the "decline" is not attributable to disengagement.
2. **Coverage masquerading as lift** - the analyst computed "70% of churned accounts had declining opens" without checking retained accounts; retained accounts showed 62%, so lift was ~1.1x.
3. **Vendor artefact exposure** - send volume changed mid-period, moving the denominator.
4. **Fails actionability** - even were it real, the team's intervention (an email) is the very channel being ignored.
The fix is not a better threshold - it is replacing the signal with customer-initiated activity from a source the vendor does not pollute.
references/ranking-methods.md›
# Ranking Methods
How to rank candidate signals by predictive strength without a data science team - and when to escalate to one. Every method that survives the deletion rules below runs in a spreadsheet or a few queries.
## Which method first
Run the methods in this order and stop when the register clears its Pass Threshold - each later one refines the ordering of signals the earlier ones already found, and none of them finds a signal the backtest missed:
`backtest + lift > retention curves > WoE/IV > logistic regression == Cox`
Value and effort rank these identically, so one line carries both axes. Effort in orders of magnitude, all of it analyst hours on data that already exists - no new instrumentation:
| Method | Effort | What the effort buys |
| ------------------------------ | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| Backtest + lift (§1-2) | a week to assemble the churned-account list once, then an hour per candidate | The shippable register: which signals precede churn and by how much |
| Retention curves (§3) | an hour, on the same pull | Where the curves diverge, which estimates lead time - the attribute lift alone cannot give |
| WoE/IV (§4) | an hour per candidate, most of it binning judgment | A strength band that separates two candidates the lift numbers tie |
| Logistic regression / Cox (§8) | a week to fit, then a standing job to refit as the book turns over | Exposes a candidate that only looked strong because it co-moves with a stronger one |
Logistic regression `==` Cox because they consume the same labeled dataset, carry the same refit burden, and differ only in what they model - churn probability versus time-to-churn - not in how much churn either lets you catch earlier.
Delete, do not demote, a method the user's constraints rule out - a method parked at the bottom of a list gets attempted anyway, and its output is confidently wrong rather than obviously missing:
- **Fewer than ~50 churned accounts in 12-24 months**: drop WoE/IV, logistic regression, and Cox from the menu entirely. Report the backtest numbers as directional.
- **Nobody will refit the model after ship**: drop logistic regression and Cox at any sample size. An unrefitted model decays into a signal ranking that describes a book that no longer exists.
- **A hard date inside the next renewal cycle**: drop everything below retention curves. The first two methods answer "which signals ship" on their own.
Small samples are the norm - a 3% monthly churn book of 300 accounts produces roughly a hundred churns a year. A thin sample blocks only the math below WoE/IV; the spreadsheet stack still ships a defensible register, and produces the labeled dataset a data team will want anyway.
This ordering is a default, not a law.
Re-rank it against what you already know about the user:
- An in-house analyst who fits models weekly moves regression up two rungs.
- A book with 2,000 churned accounts and no lead-time question makes retention curves the first stop rather than the second.
## 1. Backtest against known churned accounts
Pull every churn from the period. For each, reconstruct the candidate signal's state at 30, 60, and 90 days pre-churn (120-180 for enterprise motions).
A predictive signal was already firing at those checkpoints on most churned accounts; a signal still green at 60-90 days pre-churn on most of them is a blind spot regardless of how sensible it sounds. This same pull doubles as the survivorship-bias check on any existing health score.
## 2. Churn-rate lift over base rate
The workhorse. Requires both groups - accounts that showed the signal and churned, and accounts that showed it and did not. Computing only on churned accounts yields coverage, not lift; the distinction is the most common mistake in this exercise.
```
base rate B = churned accounts / all active accounts, per outcome window
signal rate S = churned among signal-showing accounts / all signal-showing accounts
lift = S / B
```
Report in the canonical phrasing: "of accounts showing X, N% churned within the window, vs base rate B%." Also record the false-alarm side (signal-showing accounts that retained) - a 5x-lift signal firing on half the book floods whatever intervention capacity exists.
## 3. Cohort / retention-curve comparison
Split accounts into cohorts by signal presence at a point in time (or by onboarding month, plan tier, segment), then plot % still active at 30/60/90/180/365 days per cohort and compare the curves. This is the practitioner analogue of Kaplan-Meier survival analysis: it handles still-active accounts naturally instead of discarding them, and shows _when_ churn accelerates, not just whether.
A widening gap between signal-present and signal-absent curves is the visual form of lift; the point where the curves diverge estimates lead time. If statistical tooling exists, the log-rank test formalizes whether two curves genuinely differ.
## 4. Weight of Evidence / Information Value
Borrowed from credit-risk scoring; spreadsheet-computable univariate ranking, no model training required.
Per candidate variable: bin it (e.g. days-since-last-login into 0-13 / 14-29 / 30+), then per bin compute the share of all churned accounts and the share of all retained accounts falling in that bin:
```
WoE(bin) = ln( %of_churned_in_bin / %of_retained_in_bin )
IV = sum over bins of (%of_churned_in_bin - %of_retained_in_bin) x WoE(bin)
```
Interpretation bands, by convention (credit-scoring practice, not a law):
| IV | Predictive strength |
| ---------- | ------------------- |
| < 0.02 | Not useful |
| 0.02 - 0.1 | Weak |
| 0.1 - 0.3 | Medium |
| 0.3 - 0.5 | Strong |
An IV far above this scale is a leakage warning: the "signal" is probably a consequence of the churn decision already made (a tripwire), not a leading indicator.
WoE/IV also:
- Handles missing values without imputation (missingness itself can be scored as a bin, often informative).
- Works on categorical and continuous variables alike.
## 5. Lead time and coverage - ranked attributes, not footnotes
For every churned account that fired a signal, record days between first fire and churn; the median is the signal's lead time. Coverage is the share of all past churned accounts that fired it.
Rank on the triple, lift x lead time x coverage, because each fails alone:
- A high-lift tripwire with 3 days of lead time is un-actionable.
- A 90-day-lead signal catching 8% of churns is trivia.
- Broad coverage at 1.2x lift is noise.
## 6. Divide that value score by what the signal costs to stand up
The triple is the value axis only. Complete the ratio with the effort of emitting the signal at all:
- Engineering work to instrument it.
- Data latency before it fires.
- The standing maintenance of whatever keeps it alive.
Two signals with identical lift are not the same decision when one is a query against yesterday's login table and the other needs a product event stream that does not exist. Score both, then order candidates by value per unit of effort.
Take the per-category effort estimates and the resulting default build order from the signal taxonomy reference rather than re-deriving them here; that file is where the candidate list is chosen, so the ordering lives beside the choice. Record the call in the register's `effort` line, so a later reader can see that a signal ranked below a weaker one because it cost a quarter of engineering time, not because it predicted worse.
## 7. Once any scoring model exists: precision/recall, not accuracy
Churners are a rare class, so accuracy is misleading - predicting "nobody churns" scores in the 90s on most books.
Evaluate with:
- Precision: of flagged accounts, how many churned.
- Recall: of churned accounts, how many were flagged.
- PR-AUC: when comparing model versions.
On small samples use k-fold cross-validation rather than a single train/test split - one split can give a misleading read either direction.
## 8. Logistic regression and Cox - only when the deletion rules above left them standing
- **Logistic regression** ranks signals by coefficient contribution to churn probability, controlling for the others - it exposes candidates that only looked strong because they co-move with a stronger one.
- **Cox proportional hazards** is the multivariate extension of the retention-curve comparison: it ranks covariates by their effect on time-to-churn while handling still-active accounts.
references/signal-taxonomy.md›
# Signal Taxonomy
Category-by-category candidate signals, the operational-definition template every candidate must satisfy, lead-time expectations, and the frameworks worth crediting. All tool-agnostic: "product event stream" means whatever telemetry the user has, "support ticket system" whatever they file tickets in.
## Table of Contents
- [The five-part operational definition](#the-five-part-operational-definition)
- [Categories, ranked by value per unit of instrumentation effort](#categories-ranked-by-value-per-unit-of-instrumentation-effort)
- [Lead-time expectations](#lead-time-expectations)
- [Framework credits](#framework-credits)
- [Confirmation tripwires](#confirmation-tripwires)
## The five-part operational definition
A signal is testable only when all five parts are written down:
| Part | Question it answers | Example sketch |
| ------------------ | ----------------------- | ------------------------------------------------------------------------------------ |
| Event | What is observed? | Weekly active users in the account |
| Threshold | How much change counts? | Down 40% or more |
| Observation window | Over what period? | Sustained 3 consecutive weeks |
| Baseline | Compared to what? | The account's own trailing-90-day median |
| Normalization | Adjusted for what? | Per licensed seat; seasonality-adjusted for known cycles (holidays, fiscal-year-end) |
Baselines come in two flavors:
- The account's own history: catches decline in a healthy account. Prefer this for decline signals.
- Segment peers: catches an account that never activated. Prefer this for adoption-depth signals.
Without seat normalization, every growing account looks healthier and every downsizing account looks sicker than it is.
## Categories, ranked by value per unit of instrumentation effort
Work the categories in the efficiency order below and stop when the register clears its register-level floor.
Effort here is never money. It is:
- Engineering work to emit the signal at all.
- The analyst hours to define and backtest it.
- The latency before it can fire.
- Whatever standing job keeps it alive afterwards.
The axes disagree sharply, so each gets its own line:
- **value** (churn caught early enough that someone can still act): `product usage > feature-adoption breadth > champion departure (B2B) > login recency > seat utilisation > support-ticket pattern == customer business context > CRM engagement > billing behaviour > survey movement`
- **effort**, cheapest first: `login recency == billing behaviour == survey movement > seat utilisation == CRM engagement > support-ticket pattern > champion tracking > customer business context > product usage == feature-adoption breadth`
- **efficiency**, the default build order: `login recency > seat utilisation > support-ticket pattern > CRM engagement > champion departure (B2B) > product usage > feature-adoption breadth > billing behaviour > customer business context > survey movement`
- **compliance cost**: product usage, feature-adoption breadth, champion tracking, and customer business context only - each collects behavioural or personal data on named individuals, so each triggers a privacy review before the first event is stored, and the exposure is one-way: events collected without a lawful basis have to be deleted, not relabelled. Every other category reads a record the business already keeps for billing or support, and carries none.
Ties, justified:
- **login == billing == survey** on effort: each is a query against a system of record that already stores the field, with nothing new to emit.
- **seats == CRM**: both cost an hour of definition work rather than plumbing, reconciling provisioned against active seats and filtering CRM activity down to customer-initiated.
- **product usage == feature-adoption breadth** on both effort and compliance: they are blocked on the same missing artifact, a per-account product event stream. Once it exists both collapse to near-zero together, so they never separate.
- **support-ticket pattern == customer business context** on value: each catches a churn population the usage signals structurally miss, the frustrated heavy user and the healthy-usage account killed by a budget freeze, and neither covers much of the book alone.
### 1. Login recency and frequency - effort: near-zero
The RFM lens (Recency/Frequency/Monetary), remapped from retail to subscription:
- Recency: days since last login.
- Frequency: usage or touchpoint count over the window.
- Monetary: tier and expansion history.
Across sources applying RFM to churn, recency carries the strongest predictive influence of the three. Login logs exist wherever authentication does, which is what makes this the first candidate on every book.
### 2. Seat / licence utilisation - effort: near-zero to an hour
Gap between provisioned and active seats. A wide gap means low switching cost at renewal even if the active users are happy.
Define as active seats / provisioned seats, trended. The hour goes to reconciling what "active" means against the provisioning record, not to new data collection.
### 3. Support tickets and escalations - effort: an hour
Volume, severity mix, unresolved-ticket age, and sentiment trend - plus formal escalations as a discrete event. Genuinely ambiguous: an engaged customer files more tickets because they care; a silently disengaging one files none.
Never ship raw volume alone - pair it with sentiment/severity, and define "silence after a spike" (tickets surge then stop) as its own candidate. Sentiment scoring is what lifts this above near-zero.
### 4. CRM meeting/email engagement - effort: an hour
Customer-initiated meetings booked, email response rate and latency, QBR attendance. Artefact warning: a CSM emailing an at-risk account makes measured "engagement" rise - count customer-initiated activity only, or the vendor's own save motion pollutes the signal. That filter is the whole hour, and skipping it is what turns this category into noise.
### 5. Champion and executive-sponsor engagement/departure - effort: a week to stand up, then a standing job
B2B-only. Repeatedly cited as among the strongest single B2B signals; Lemkin puts non-renewal odds above 50% when the sponsoring executive departs.
Two distinct signals:
- Departure: a job-change event on the recorded champion.
- Disengagement: champion response rate/meeting attendance declining.
Champion departure is commonly cited as preceding churn by 30-60 days - and typically discovered at the renewal call, not when it happens, so instrument the detection, not just the field. The standing job is keeping the recorded-champion field true; a stale field silently produces a signal that never fires.
### 6. Product usage / telemetry - effort: near-zero if the product already emits per-account events, otherwise a quarter
The most commonly cited earliest-detectable family, and the highest-value category on the list. Decline is usually gradual - daily use becomes weekly becomes biweekly - not a cliff, so define trends against the account's own baseline rather than absolute floors.
Usage/engagement decline is commonly cited by practitioners as appearing 60-90 days before churn. It ranks sixth on efficiency only because of that instrumentation gate; see the starvation note below before accepting the position.
### 7. Feature-adoption breadth - effort: same gate as §6, plus an hour
Narrowing to fewer features, or plateauing instead of expanding into new teams and use cases, is flagged as a risk even when raw engagement looks fine (Jason Lemkin, SaaStr: usage that is not growing or spreading signals the customer is not seeing incremental value). Define as count of distinct core features used per window vs the account's prior window and vs segment peers. Needs the event stream of §6 plus a maintained definition of which features count as core.
### 8. Billing and payment behaviour - effort: near-zero
Failed payments, invoice disputes, late renewals, downgrade requests.
- **B2B**: late-stage/confirmatory - by the time payment friction surfaces, the account is usually already evaluating alternatives, which is what drops a near-zero-effort category this far down.
- **B2C**: payment failure is a proportionally larger churn driver and worth catching pre-emptively, so it climbs several rungs on a consumer book; it is still confirmatory of intent only when the customer lets it fail.
### 9. Customer business context - effort: a week to stand up, then a standing job
Layoffs, budget freeze, leadership change, M&A at the customer. Externally sourced (news, job-change alerts, the champion saying so); low frequency but high severity, and often the only warning for accounts whose product usage looks fine. Promote it on an enterprise book where losing one account is material - low coverage stops mattering when each covered account is worth a quarter of the number.
### 10. Survey movement (NPS/CSAT) - effort: near-zero
Weaker than assumed: point-in-time, and response-biased - the happiest and angriest over-respond while the quietly disengaging do not respond at all, making it a weaker churn predictor than product engagement in most subscription businesses. Counterpoint worth knowing: Dave Kellogg argues a direct "intent to renew" survey question captures churn-predictive signal that NPS misses.
Candidate definitions: score trend across waves (not one reading), and non-response itself after prior participation. It lands last despite costing nothing because the accounts it misses are exactly the quietly-disengaging ones the register exists to catch.
### What this order starves
Product usage and feature-adoption breadth carry the highest lift and the longest lead times on the list, and cost the most to instrument. So a register built strictly by ratio is all CRM-derived, all lagging relative to what the customer is actually doing in the product.
Promote both to the front of the build order when any of these holds:
- The product already emits per-account events. Effort collapses to near-zero and the ratio inverts outright - this is the common case in product-led books, so check it before accepting the default order.
- The Interview answer was "compounding asset" and the renewal cycle is long enough that the instrumentation lands before the next renewal window opens.
- A register assembled from the cheap categories fails the 70% register-level floor, which is what happens on a book whose churn is usage-driven: no volume of CRM-derived signals substitutes for never having watched the product.
Delete a category outright rather than parking it at the bottom when the Interview ruled it out:
- An unreachable data source.
- The B2B-only relationship categories on a B2C book, where there is no buying committee to lose a member from.
A ruled-out category left in the list reappears as scope two sessions later.
Treat this order as a default, not a law - it shifts with the book, with the product, and with who executes it.
Re-rank it against what you already know about the user:
- A data engineer on the team collapses the telemetry effort.
- An unmaintained CRM inflates the cheap categories' real cost.
- A book where every account is already wired for events makes the entire ordering above moot.
## Lead-time expectations
Commonly cited practitioner figures, not laws - measure the user's own median lead times in the backtest and prefer those:
- Usage/engagement decline: 60-90 days before churn.
- Enterprise intervention windows: extended to 120-180 days, to leave room for procurement cycles and multi-stakeholder re-alignment.
- Champion departure: 30-60 days before churn.
## Framework credits
Credit these where they genuinely fit; never force one onto a book of business it does not match:
- **Leading vs lagging indicators** - the foundational distinction this whole skill runs on.
- **Lincoln Murphy's Success Milestones** - usage is not success; track progress toward the customer's own defined outcome, and run structured churn-reason analysis on every loss.
- **Gainsight's guidance** - a deliberately small weighted signal set (4-6 signals) beats tracking everything, segmented by account tier because health looks different for SMB vs enterprise; Nick Mehta's DEAR frame (Deployment, Engagement, Adoption, ROI) names the leading-indicator stack he recommends over NPS alone.
- **RFM (Recency/Frequency/Monetary)** - remapped for subscription as above; recency strongest.
- **Survival/cohort analysis and WoE/IV** - ranking machinery, covered in the ranking reference.
## Confirmation tripwires
Class these as confirmatory - they formalize a decision usually made weeks earlier. They still belong in the register (flagged) because they escalate triage urgency, but they never satisfy the lead-time bar.
Deliberately unranked, unlike the categories above: every tripwire costs near-zero to instrument and buys zero lead time, so any ordering between them would be false precision. Instrument whichever ones the billing and product systems already expose, and treat the set as one triage input rather than a menu to choose from.
- Failed payment / lapsed card (B2B)
- Downgrade or seat-reduction request
- Procurement delay or renewal-date slippage
- Data export initiated
- Cancellation- or billing-page visits (B2C: often only days of warning)
- Auto-renewal switched off
SKILL.md›
---
name: customer-churn-signals
description: Assemble, define, validate, and rank leading indicators of customer churn from account activity, product usage, and support data into a ranked churn signal register - each signal carrying event, threshold, window, baseline, lift over base rate, lead time, and coverage. Use whenever the user mentions churn signals, churn indicators, leading indicators of churn, an early warning system for at-risk accounts, champion departure, seat utilisation drops, payment failure, or says "our health score misses churn" - even if they never say "signal". Covers B2B and B2C subscription. Do NOT use for combining signals into one composite score, weights, bands, or tiers - use mbfinotti/revops-skills@customer-health-score instead.
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.2.6"
---
# Churn Signals
Identify, operationally define, validate against actually-observed past churn, and rank the signals that precede churn in the user's own accounts. The organizing distinction is leading versus lagging:
- **Lagging** (reports what already happened): churn rate, lost MRR, a falling NPS.
- **Leading** (appears weeks before cancellation and leaves time to act): login decline, narrowing feature adoption, a champion going quiet.
Every candidate must pass a two-part actionability test before it is worth shipping:
1. It fires early enough that someone can still intervene.
2. Someone actually can intervene on it.
Late-firing events - failed payments, downgrade requests, procurement delays, renewal-date slippage - are confirmation tripwires that formalize a decision made weeks earlier. Class them as confirmatory, never as leading.
Lincoln Murphy's Success Milestones frame is the standing caution over the whole exercise: usage is not success. Track progress toward the customer's own outcome, because accounts can look busy right up to an "unexpected" churn.
This skill ends at a ranked, validated signal register; whoever assembles those signals into a composite score - weighting, score scales, red/yellow/green tiers - or the CSM playbooks, save offers, and renewal plays that act on them should use `mbfinotti/revops-skills@customer-health-score`. Pipeline stage design, pre-sale lead scoring, and the sales-to-CS handoff are also out of scope.
## Interview
Ask before proposing anything. One question per message; offer multiple-choice answers where possible; skip anything already answered.
- B2B, B2C/consumer subscription, or both books of business?
- How many churned accounts exist from the last 12-24 months? This decides which ranking method is even possible: backtest-and-lift works on small samples; WoE/IV and any regression need volume (see ranking reference).
- Which data sources are actually reachable, and at what granularity: product event stream, login logs, support ticket system, CRM activity (meetings, emails), billing records, survey scores? Per-account per-day, or only aggregates?
- Contract and renewal shape: monthly self-serve, annual, multi-year? Typical renewal cycle length?
- What intervention capacity exists - who acts on a flagged account, and how many flags per week can they absorb? A signal nobody can act on is not worth ranking.
- Does a health score already exist? If yes: validate its input signals against real churn first, or build the register fresh?
- By what date must the register be live? A date inside the next renewal cycle deletes every signal needing new instrumentation and every method below retention curves.
- A one-off win - save this quarter's renewals - or a compounding asset the team keeps recalibrating? "Compounding" promotes product telemetry and feature-adoption breadth to the front of the build order; "one-off" leaves the register CRM-derived.
- What is the effort ceiling: analyst hours available, engineering time you can actually requisition, and whether anyone will refit a model after ship? No engineering time deletes the product-telemetry categories; nobody to refit deletes logistic regression and Cox at any sample size.
Re-rank the candidate order against these three answers before proposing anything, and say out loud which answer moved which option - a ranking the user cannot trace back to their own constraints reads as arbitrary.
## Workflow
1. Run the Interview; collect every answer before proposing signals.
2. Build the candidate list from [references/signal-taxonomy.md](references/signal-taxonomy.md), working its categories in the efficiency order stated there - login recency first, survey movement last, product telemetry sixth unless the product already emits per-account events. Delete the categories the Interview ruled out instead of listing them for later. Mark confirmatory tripwires as such from the start.
3. Write the operational definition for every candidate: the event, the threshold, the observation window, the comparison baseline, and normalization for account size and seasonality. A signal missing any of the five is not testable; fix the definition or drop the candidate.
4. Pull the churned-account list and reconstruct each candidate's state at 30/60/90 days pre-churn (extend to 120-180 for enterprise motions). If you can query the data directly, compute it; otherwise emit the queries or spreadsheet steps for the user and work from their results.
5. Rank per [references/ranking-methods.md](references/ranking-methods.md): churn-rate lift over base rate, median lead time, and coverage for every candidate, divided by what that candidate costs to stand up. Run the methods in that file's stated order - backtest and lift first - and stop when the register clears; delete the methods the Interview's sample-size and maintenance answers ruled out rather than keeping them as a stretch goal. Lead time is a first-class ranked attribute alongside strength - a strong signal that fires days before cancellation loses to a weaker one that fires 90 days out.
6. Run the survivorship check: for the accounts that churned, was each surviving candidate actually firing 60-90 days before? A signal still green pre-churn on most churned accounts is a blind spot, not a predictor.
7. Fill the register (shape below), verdict every signal ship/watch/reject against the Pass Threshold, and iterate definitions, thresholds, and windows until the register-level bar clears.
8. Validate the register with the user section by section, hand the shipped signals to `mbfinotti/revops-skills@customer-health-score` for composite scoring, and book a re-backtest for when the next quarter of churn outcomes exists.
9. If your harness has persistent memory, memorize the approved register - signals, definitions, lift, lead times, coverage, verdicts, re-backtest date - so later recalibration runs start from it instead of re-interviewing.
## The Signal Register
Deliver every engagement as this artifact. Worked versions live in [references/examples.md](references/examples.md).
```
CHURN SIGNAL REGISTER - <company/book>, <date>, v<n>
Base rate : <B>% of accounts churn per <window>; churned sample = <N> (<period>)
Method : backtest + lift [+ retention curves | + WoE/IV] - in that order, stopping when
the register cleared; name any method deleted and why (sample, no refit, deadline)
Rows : ordered by value per unit of instrumentation effort
signal : <name>
category : usage | adoption | seats | relationship | support | billing | survey | crm | context
source : <system category, e.g. product event stream>
definition: event + threshold + observation window + baseline + normalization
window : <observation window>
baseline : <what normal is - account's own history and/or segment peers>
lift : of accounts showing this, <N>% churned within <window> vs base rate <B>% (<x.x>x)
lead time : median days between first fire and churn
coverage : % of past churned accounts that fired this signal
effort : near-zero | an hour | a week | a quarter | a standing job - to instrument and keep alive
confidence: evidence note - sample size, IV if computed, caveats
verdict : ship | watch | reject (confirmatory-flagged where applicable)
Register-level : % of past churns the shipped set would have caught at minimum lead time
Handoff : shipped signals -> mbfinotti/revops-skills@customer-health-score
```
## Pass Threshold
A signal earns "ship" only when all three clear. Treat the numbers as this skill's practical floor, not researched constants - tighten them from the user's own data when the sample allows:
- **Lift**: accounts showing the signal churn at **at least 2x the base rate** within the observation window. A thinner edge will not survive live noise.
- **Lead time**: median **at least 30 days** before churn - **90 for enterprise/annual-contract motions**, where intervention needs procurement-cycle room. Confirmatory tripwires are exempt but must carry the confirmatory flag.
- **Coverage**: fired for **at least 25% of past churned accounts** - a signal that catches almost none of the real churns is trivia however strong its lift.
Register-level: the shipped set combined must have fired on **at least 70% of past churned accounts** at the minimum lead time or earlier. Iterate - adjust thresholds and windows, add categories, split segments - until it clears; refuse to ship a register that would have missed most known churns.
## B2B vs B2C
- **B2B-only**: champion and executive-sponsor departure, and multi-stakeholder buying-committee dynamics - there is no internal committee to lose a member from in B2C. B2B contract terms also dampen and delay the visible churn event: an annual contract keeps a dissatisfied customer paying until expiry, so the behavioral signals run far ahead of the commercial one, which is exactly why leading signals matter more there, not less.
- **Shared, identical method**: usage/login decline, feature-adoption breadth, and support sentiment are tracked in both B2B and B2C. Only the unit changes: the account and its committee in B2B, the individual subscriber in B2C.
- **B2C-specific**: payment failure is a proportionally larger churn driver; cancellation is a one-person, one-tap decision, so lead times compress and intervention windows shrink; and "churn by indifference" - a forgotten low-cost subscription - has no B2B analogue.
## Common Failure Modes
| Trap | Why it burns | Fix |
| ---------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Lagging metric shipped as leading | Churn rate, lost MRR, NPS drops report the past | Classify every signal by measured lead time; tripwires get the confirmatory flag |
| Correlation without prediction | Metrics move before churn without forecasting it | Require backtested lift; sanity-check with churned-account interviews |
| Ranking by ratio starves the telemetry signals | Product usage and adoption breadth carry the highest lift and longest lead times and cost the most to instrument, so a register built purely on efficiency is all CRM-derived and lagging - the same register that survivorship bias produces from whatever the CRM already holds | Promote telemetry when events already exist, when the mandate is compounding, or when the cheap register misses the 70% floor; pull churned accounts' signal state 60-90 days pre-churn - still-green means blind spot |
| Support-ticket volume read one-way | Engaged customers file more tickets; silent disengagers file none | Pair volume with severity, age, and sentiment; treat silence-after-spike as its own signal |
| NPS/CSAT weighted as a strong predictor | Response bias - the quietly disengaging do not answer surveys | Rank below product engagement unless the user's own backtest says otherwise |
| Vendor-outreach artefacts | A CSM emailing an at-risk account makes "engagement" rise | Count customer-initiated activity only |
| Accuracy as the validation metric | Churners are rare; predicting "no churn" scores high | Precision/recall and PR-AUC once any model exists |
| Overfitting a small churned sample | A single train/test split misleads | K-fold rather than one split; widen the outcome period |
| Silent churn lost in aggregates | A slow per-account fade is invisible in book-level reporting | Trend each account against its own baseline, not the aggregate |
## Reference
- Read [references/signal-taxonomy.md](references/signal-taxonomy.md) when building the candidate list - the categories ranked by value per unit of instrumentation effort, the five-part operational template, lead-time expectations, framework credits, and the confirmatory-tripwire list.
- Read [references/ranking-methods.md](references/ranking-methods.md) when validating and ranking - the method order and what each rung buys, which methods a thin sample or an unmaintained model deletes, backtest mechanics, lift, retention-curve comparison, WoE/IV with interpretation bands, and PR-AUC.
- Read [references/examples.md](references/examples.md) when writing the register - a worked B2B register, a compact B2C one, worked lift and WoE/IV computations, and one plausible-looking bad signal decomposed.
- See `mbfinotti/revops-skills@lead-scoring` for the same backtest-and-lift validation logic applied pre-sale to leads instead of post-sale to accounts.