SKILL DETAIL
pipeline-stage-definition-audit
mbfinotti/revops-skills/pipeline-stage-definition-audit
Audit existing sales pipeline stage definitions against buyer-verifiable milestones and flag the ones built on rep activity instead. Core test - does each stage exit criterion name something the buyer did, said, or agreed to, checkable by two managers independently? Adds stage aging, conversion decay, stage-skip rate, and close-date push diagnostics. Use whenever the user mentions stage definitions, stage exit criteria, stage inflation, stages named after rep activity ("demo scheduled", "proposal sent"), "our stages don't mean anything", or "why is everything stuck in one stage" - even if they never say "audit". Covers enterprise and high-velocity pipelines, B2B and B2C. Do NOT use for designing a funnel stage set from scratch - use mbfinotti/revops-skills@revenue-funnel instead.
Installation
npx skills add https://github.com/mbfinotti/revops-skills --skill pipeline-stage-definition-audit
技能檔案
SKILL.md
最近同步 · 2026年9月15日
evals/evals.json›
{
"skill_name": "pipeline-stage-definition-audit",
"evals": [
{
"id": 1,
"prompt": "I run RevOps at Calderwell, B2B analytics, about $85k ACV, deals take roughly four months. Our seven opportunity stages are: 1 Intro Call Booked, 2 Discovery Complete, 3 Demo Delivered, 4 Champion Identified, 5 Proposal Sent, 6 Verbal Commit, 7 Closed Won/Lost. The written definitions are in a 2023 slide deck nobody has opened since. Everything piles up in Proposal Sent - 44% of open pipeline value sits there, median 68 days. Two of my managers argue about whether a deal belongs in 4 or 5 every single Thursday. Can you go through the stages and tell me what is actually wrong?",
"expected_output": "A stage-by-stage audit that classifies the pipeline type, applies a buyer-verifiability test to each exit criterion, maps stages to buyer-side buying jobs, flags uncovered jobs, assigns severities, and proposes rewritten criteria naming both the buyer act and where the evidence lives - without redesigning the stage set.",
"files": [],
"expectations": [
"Classifies the pipeline as a deal pipeline (as opposed to a development or transaction pipeline) before judging any individual stage",
"Judges each stage against a three-part test requiring the exit criterion to name a buyer action, be confirmable without the rep's opinion, and be recorded in a checkable place",
"States the operational form of the test as whether two managers looking at the same deal record would independently reach the same verdict",
"Marks Intro Call Booked, Demo Delivered and Proposal Sent as failing because they describe rep activity rather than buyer action",
"Marks Champion Identified as failing because it is rep-attested rather than independently checkable",
"Assigns Critical severity to at least one exit criterion that is pure rep activity or undocumented",
"Records the undocumented, unread stage definitions as a finding in their own right rather than silently reconstructing them",
"Maps stages to the six B2B buying jobs: Problem Identification, Solution Exploration, Requirements Building, Supplier Selection, Validation, Consensus Creation",
"Flags at least one buying job - Validation or Consensus Creation - as having no stage covering it",
"Identifies stages that map to no buying job as rep activity in disguise",
"Proposes at least one rewritten exit criterion that names both a buyer act and the artifact or field holding the evidence",
"Does not respond by designing a replacement stage set from scratch"
]
},
{
"id": 2,
"prompt": "Follow-up on the Calderwell stage cleanup. I have the admin free on Friday. She is going to rename 'Demo Delivered' to 'Solution Validation' and 'Proposal Sent' to 'Business Case Agreed' directly on the Opportunity Stage picklist, and delete 'Champion Identified' since we are dropping that stage. That Friday is 29 September, the last working day of our quarter, so everyone is heads-down on closing and nobody will notice the change. Anything I should tell her before she does it?",
"expected_output": "A refusal of the in-place rename and delete, with the deactivate-and-add sequence, a pre-change inventory of everything keyed to stage values, a forecast-category re-map for every stage, a rejection of the quarter-end cutover, and cutover-comparability mechanics.",
"files": [],
"expectations": [
"States that existing stage values must not be renamed or deleted, and that old values are deactivated while new values are added alongside",
"Explains that stage history in most CRMs is immutable and historical records keep referencing the old value, so a rename corrupts before/after reporting permanently",
"Warns that deleting the Champion Identified value risks destroying historical reporting rather than merely retiring the stage",
"Requires an inventory of everything keyed to stage values - automations, validation rules, workflows, dashboards, report filters, integrations - before any change",
"Requires the forecast-category mapping to be re-checked for every stage after the change, including stages that did not change",
"Rejects the last working day of the quarter as the cutover moment and asks for a low-traffic period instead",
"Recommends documenting a cutover date and annotating every dashboard that spans it",
"Recommends a bridge field such as legacy stage or definition version on open deals so cross-era cohorts stay comparable",
"States that old-definition and new-definition cohorts must be reported separately until the old cohort closes out, never blended into one conversion number",
"States that the new exit criteria should be written before any stage name is changed",
"Tells the user to capture the last old-definition forecast as a frozen baseline and warn finance that the series restarts at cutover"
]
},
{
"id": 3,
"prompt": "We finished the stage audit at Calderwell. Findings: four Critical (pure rep-activity criteria on stages 1, 3, 5 and 7), one High (a criterion that mentions the buyer but is recorded nowhere in the system), two Medium (ambiguous wording two managers could read differently), one Low (two stages that look almost identical). The board wants a sequenced plan by Monday. My plan is to attack the four Criticals first because they are Criticals, then the High, then the two Mediums, then merge the duplicate stages last. Sanity-check the order for me.",
"expected_output": "A re-sequenced plan ordered by accuracy recovered per unit of effort rather than by severity, with criterion rewrites first, the standing inspection second, evidence plumbing for the unrecorded criterion, the stage merge below it, and order-of-magnitude effort per rung.",
"files": [],
"expectations": [
"Rejects severity as the sequencing axis and states that severity is the value axis of a finding, not its queue position",
"Sequences fixes by accuracy recovered per unit of effort rather than by severity or cheapest-first",
"Places criterion rewrites first in the plan",
"Places the standing inspection - exit criteria on the recurring pipeline-review agenda plus a periodic inter-rater re-sample - ahead of evidence plumbing",
"States that the standing inspection recovers no accuracy on its own and earns its rank by stopping the rungs above it drifting back to opinion",
"Assigns the High finding whose criterion is recorded nowhere to the evidence-plumbing rung: a required field, validation rule, or evidence link",
"Places the stage cut or merge below evidence plumbing because it needs a deactivate-and-add migration, open-deal moves, and a forecast-category re-map",
"Gives each rung an order-of-magnitude effort such as near-zero, an hour, a week, or a quarter, rather than a precise estimate or a currency figure",
"Notes that reversibility dominates the ordering because a stage definition changed twice in a quarter destroys the trend data the definitions exist to produce",
"States that a Critical defect fixed by one wording change ships before a Medium one that needs a picklist migration",
"Does not recommend rebuilding the whole stage set as the first move"
]
},
{
"id": 4,
"prompt": "Quick one. Loopstack sells a $2,900/year dev tool; inside sales team of six, median deal closes 11 days after signup. Stages: 1 Trial Assigned, 2 Activated (buyer finished setup and invited a teammate), 3 Demo Booked, 4 Quote Accepted in portal, 5 Closed. Demo Booked gets skipped on 41% of deals and I checked - the deals that skip it close at a higher rate than the ones that go through it. My old enterprise playbook says every stage needs a mutual action plan and a security-review sign-off before a deal can advance. Should I add those here?",
"expected_output": "A refusal to import enterprise criteria into a high-velocity motion, classification of the pipeline type first, a verdict that Demo Booked is rep activity carrying no information, and a recommendation to cut rather than rewrite it - justified by pipeline type and team size, with aging thresholds derived in days from the pipeline's own distribution.",
"files": [],
"expectations": [
"Classifies the pipeline as a transaction pipeline and says so before judging any stage",
"Refuses to import mutual action plan and security-review criteria into this motion, on the grounds that a transaction pipeline judged by deal-pipeline expectations produces false findings",
"States that friction removal is this motion's design goal and that criteria must be weighted to the motion",
"Judges Demo Booked as failing the verifiability test because it names rep activity",
"Treats the 41% skip rate combined with better conversion among skippers as evidence the stage carries no information",
"Recommends cutting or merging Demo Booked rather than only rewriting its exit criterion",
"Justifies the cut over the add by pipeline type, stating that a transaction pipeline recovers more from the cut while a deal pipeline recovers more from the add",
"Treats Activated as passing because it rests on machine-recorded product events",
"Treats Quote Accepted in portal as passing because it records a buyer act in a checkable place",
"Sets aging flags in hours or days rather than weeks, derived from this pipeline's own distribution",
"Notes that a six-rep team on an 11-day cycle makes the migration small enough to promote the cut above its default rung position"
]
},
{
"id": 5,
"prompt": "The new CRO at Vantrix wants the pipeline to reflect our qualification methodology. Concretely: replace our six stages with MEDDPICC 1 through 8, one stage per letter, so a deal sitting in stage 4 means we have Metrics, Economic Buyer, Decision Criteria and Decision Process nailed down. Reps advance the deal when the next letter is filled in. He wants the picklist changed before our sales kickoff in three weeks. Write me the new stage definitions.",
"expected_output": "A refusal to build the stage list out of qualification-framework letters, an explanation that the stage records buyer position while each qualification dimension is its own continuously-updated field, and buyer-verifiable exit criteria offered instead - with the qualification framework retained only as red/yellow/green gating.",
"files": [],
"expectations": [
"Refuses to build a stage list out of the qualification framework's letters and states there is no such thing as MEDDIC or MEDDPICC stages numbered one per letter",
"States that the stage records the buyer's position in their purchase, never a qualification scorecard",
"Explains that each qualification dimension belongs in its own continuously-updated field, separate from the stage picklist",
"Warns that conflating the scorecard with the stage picklist degrades both constructs",
"Points out that a pipeline staged by scorecard letters can no longer answer where the buyer is in their purchase",
"Proposes exit criteria that name a buyer act and the artifact recording it, rather than a filled-in qualification field",
"Keeps the qualification framework as red/yellow/green scoring per exit criterion, requiring all-green to advance",
"States that yellow means the rep believes it and evidence is pending, and is never a passing state",
"Warns that any picklist change still requires the deactivate-and-add migration rather than an in-place edit, whatever the kickoff deadline",
"Does not deliver a stage set named after the qualification framework's letters"
]
},
{
"id": 6,
"prompt": "Building the exec deck for our stage audit at Prentwood. I want to open with '93% of sales teams cannot predict revenue within 5%' and 'CRM data is only 40-60% accurate' - those are everywhere online and they make great hooks. Then I will set our stale-deal threshold at 14 days because that is the industry standard, and say healthy pipelines carry 3.5x coverage. Draft the opening section with those numbers in it.",
"expected_output": "A refusal to state the circulating statistics as fact, the underlying mechanism stated in their place, a rejection of the borrowed 14-day and 3.5x constants, and every threshold that is given carrying a provenance tag alongside the number.",
"files": [],
"expectations": [
"Declines to state '93% of sales teams cannot predict revenue within 5%' as fact and notes the figure circulates unattributed",
"Declines to state 'CRM data is only 40-60% accurate' as fact on the same grounds",
"Replaces the statistics with the mechanism: when advancement is subjective, two reps stage identical deals differently, which corrupts stage conversion data and everything built on it",
"Rejects 14 days as an industry-standard stale threshold and requires the threshold be derived from this pipeline's own dwell-time distribution",
"Offers a pipeline-derived alternative such as a percentile of time-in-stage in place of the borrowed threshold",
"Labels each threshold it does give with a provenance tag such as published company practice, practitioner consensus, or derived from the user's own data",
"Keeps the provenance tag next to the number in the deliverable rather than relegating it to a separate methodology note",
"Never presents a vendor heuristic as an industry constant",
"Does not supply the 3.5x pipeline-coverage figure as an industry benchmark",
"Presents the 5-7 stage count and 2-4 criteria per stage bands as practitioner consensus rather than a law"
]
},
{
"id": 7,
"prompt": "Great news from Prentwood. Sixteen days after we rolled out the new exit criteria, required-field fill rate on opportunities went from 58% to 94% and rep logging volume is way up. I am putting 'stage definition project: success' on the QBR slide next week and closing the project out. Anything else worth measuring before I close it?",
"expected_output": "A rejection of compliance metrics as evidence of success, a statement that conversion data is meaningless this early, adoption checkpoints instead of a verdict, inter-rater agreement named as the earliest trustworthy signal, and the project kept open with a standing re-check.",
"files": [],
"expectations": [
"Rejects field-fill rate and logging volume as measures of success, stating they measure pressure rather than accuracy",
"Explains that under compliance pressure reps log more, faster and worse, so data quality need not have improved",
"States that conversion and distribution metrics are untrustworthy until a full sales cycle of deals has flowed through the new definitions",
"Treats the day-16 reading as an adoption checkpoint, not a verdict",
"Recommends 30/60/90-day adoption checkpoints rather than a single early declaration of success",
"Names inter-rater agreement on periodic deal samples as the earliest trustworthy signal, available long before conversion data matures",
"Names per-rep conversion variance as the sharpest single indicator that definitions are applied subjectively",
"Names stage-to-stage conversion stability and monotonic decay across the funnel as a tracked measure",
"Keeps a standing periodic inter-rater re-sample rather than closing the project, warning that criteria otherwise drift back to opinion within about two quarters",
"Recommends cross-checking self-reported fields against independent signals such as calendar attendees, attached artifacts, or product events",
"Schedules the real re-check at least one full sales cycle out"
]
},
{
"id": 8,
"prompt": "Two things on the Ashgrove audit. First, you keep asking for two managers scoring deals independently - we have exactly one sales manager, that is it, plus me, and I own the definitions so I am not neutral. Second, I pulled stage conversion from a report that counts how many open deals are sitting in each stage right now: Discovery 120, Evaluation 40, Negotiation 12. So that is 33% then 30%, right?",
"expected_output": "A correction that point-in-time counts are a stock measure and cannot yield conversion, the cohort-based flow computation instead, and a stated substitution for the missing inter-rater check rather than dropping it or claiming an agreement rate never measured.",
"files": [],
"expectations": [
"States that the point-in-time counts are a stock measure and cannot produce stage-to-stage conversion",
"Requires conversion to be computed on cohorts of deals that entered a stage in a period, measuring the share that ever reached the next stage",
"Explains that snapshot-based conversion overweights stuck deals",
"Replaces the missing inter-rater check with an automated-event spot audit against machine-recorded criteria rather than dropping the check",
"Requires the deliverable to state that substitution explicitly instead of reporting an agreement rate that was never measured",
"Gives the inter-rater check's shape where reviewers exist: two independent reviewers, a sample of 10-20 open deals, scored without conferring",
"States the agreement bar as at least 90% and labels it a working bar to tighten from the user's own data, not a published standard",
"Says disagreements point at the ambiguous criterion, which is then rewritten and re-sampled",
"Notes that any two independent reviewers who can judge a deal record without conferring can stand in for two managers",
"Warns that duration reporting may count only a stage's first occurrence, understating time for deals that regressed and re-entered",
"Tells the user to check how the export counts re-entries before trusting the aging numbers"
]
},
{
"id": 9,
"prompt": "Context before you write the plan for Mercerline. We are in week 5 of a 13-week quarter and the number is at risk, so nothing can be allowed to disturb this quarter. Our only CRM admin left in July and the backfill does not start until November. And honestly nobody owns stage definitions here - I asked, sales leadership says it is RevOps, RevOps is me, and I cannot approve my own changes. I still want the strongest plan you can give me. We have nine stages and I am fairly sure two of them are duplicates.",
"expected_output": "A plan with the ruled-out rungs deleted and named as deleted with the constraint that removed each one, leaving criterion rewrites plus the standing inspection, the governance gap recorded as a finding, and the stage count noted against the consensus band without acting on it this quarter.",
"files": [],
"expectations": [
"Strikes the stage cut and the stage add from the plan because no approver exists for a stage-definition change, and says explicitly that they are struck",
"Strikes the evidence-plumbing rung because no CRM admin is available, and says explicitly that it is struck",
"Strikes every migration rung for the duration of the live quarter because a stage change corrupts the quarter it lands in",
"Deletes each ruled-out rung from the plan rather than demoting it to a lower priority",
"Explains that a rung parked at the bottom of the plan returns later as scope nobody budgeted",
"Leaves criterion rewrites as the work that can proceed, since they need no picklist change, no migration and no forecast-category re-map",
"Adds the standing inspection, which costs roughly an hour a quarter and needs no admin",
"Records the absence of an owner and approval path for stage definitions as a governance finding",
"Notes the nine-stage count against the 5-7 practitioner-consensus band without treating that band as a rule",
"Does not propose renaming or merging the two suspected duplicate stages inside the live quarter",
"Says which constraint struck which rung rather than presenting the reduced plan without its reasons"
]
},
{
"id": 10,
"prompt": "Hartsook runs three motions through one Opportunity pipeline: self-serve signups that never touch a rep, an SMB inside-sales team, and a five-person enterprise team doing $200k-plus deals. Same six stages for all three. Marketing also tracks lifecycle stages in the marketing tool and wants to merge those into the opportunity stages so there is 'one funnel' everyone reads. Only 'Closed Won' and 'Contract Sent' have definitions anyone here can quote. Fix our stage definitions.",
"expected_output": "A scope call that declines to patch criteria and routes the work to funnel design, with the promotion conditions named, a refusal to merge marketing lifecycle stages into the opportunity stage list, and an explanation of why an efficiency ordering would otherwise defer this redesign indefinitely.",
"files": [],
"expectations": [
"Declines to patch the existing stage definitions and routes the work to a funnel-design redesign instead",
"Names mbfinotti/revops-skills@revenue-funnel as the destination for that redesign",
"Gives multiple motions sharing one stage list as a condition that promotes a full redefinition",
"Gives fewer than two stages surviving as usable anchors as a condition that promotes a full redefinition",
"Refuses to merge marketing lifecycle stages into the opportunity stage list, keeping funnel and pipeline as distinct reporting constructs",
"States that a self-serve motion with no rep-owned pipeline belongs in buyer-facing lifecycle or funnel stages owned by marketing or growth, not in the rep pipeline",
"Notes that a true pipeline construct reappears for the sales-assisted and enterprise motions, which need their own stage sets",
"States that a full redefinition ranks first on accuracy recovered and last on value per unit of effort, which is why an efficiency ordering defers it indefinitely",
"Explains that deferring the redefinition leaves well-written criteria sitting on a stage model that cannot carry them",
"Does not deliver rewritten exit criteria for the six shared stages as the main deliverable"
]
}
],
"trigger_queries": [
{ "query": "audit our pipeline stage definitions", "should_trigger": true },
{ "query": "our stages don't mean anything anymore", "should_trigger": true },
{ "query": "why is everything stuck in one stage", "should_trigger": true },
{ "query": "are our stage exit criteria any good", "should_trigger": true },
{ "query": "two of my managers stage the same deal differently", "should_trigger": true },
{ "query": "we have stages called Demo Scheduled and Proposal Sent, is that bad", "should_trigger": true },
{ "query": "review our opportunity stages and tell me which ones are just rep activity", "should_trigger": true },
{ "query": "stage inflation is wrecking our forecast, look at how the stages are defined", "should_trigger": true },
{ "query": "what should the exit criteria be for each of our deal stages", "should_trigger": true },
{ "query": "our reps advance deals on optimism, how do I write criteria that stop that", "should_trigger": true },
{ "query": "we have 11 stages in the CRM and half of them look the same", "should_trigger": true },
{ "query": "how do I tell whether a stage is real or just a rep task in disguise", "should_trigger": true },
{ "query": "can you check whether each of our stages proves the buyer actually did something", "should_trigger": true },
{ "query": "everything sits in Proposal Sent for two months", "should_trigger": true },
{ "query": "I want buyer-verifiable milestones instead of activity-based stages", "should_trigger": true },
{ "query": "grade our current stage definitions", "should_trigger": true },
{ "query": "our pipeline stages were written in 2022 and nobody follows them", "should_trigger": true },
{ "query": "which of our stages map to nothing the buyer actually does", "should_trigger": true },
{ "query": "deals skip stage 4 constantly, what does that tell me", "should_trigger": true },
{ "query": "is it safe to rename our opportunity stage picklist values", "should_trigger": true },
{ "query": "we want to replace our stages with MEDDPICC letters, thoughts", "should_trigger": true },
{ "query": "how many stages should a B2B pipeline have and what proves each one", "should_trigger": true },
{ "query": "our stage definitions live in people's heads, help", "should_trigger": true },
{ "query": "the same deal could sit in either of two stages depending on who you ask", "should_trigger": true },
{ "query": "write tighter advancement criteria for our sales stages", "should_trigger": true },
{ "query": "one of our stages converts at 97% and I don't trust it", "should_trigger": true },
{ "query": "how do I stop reps moving deals forward with no evidence", "should_trigger": true },
{ "query": "we're PLG going upmarket and sales-assisted deals now share a pipeline with self-serve signups, do our stages still mean anything", "should_trigger": true },
{ "query": "what evidence should a deal have before it is allowed to leave discovery", "should_trigger": true },
{ "query": "close dates keep slipping from the same stage every single time", "should_trigger": true },
{ "query": "audit our deal stages before we forecast off them", "should_trigger": true },
{ "query": "help me rewrite Demo Delivered into something the buyer actually does", "should_trigger": true },
{ "query": "should Champion Identified be a stage", "should_trigger": true },
{ "query": "I inherited a pipeline with 9 stages and zero documentation", "should_trigger": true },
{ "query": "how do I make stage advancement objective instead of a judgment call", "should_trigger": true },
{ "query": "check our advancement criteria against what the buyer has actually agreed to", "should_trigger": true },
{ "query": "our managers spend the whole pipeline review arguing about staging", "should_trigger": true },
{ "query": "is verbal commit a legitimate stage", "should_trigger": true },
{ "query": "what's wrong with naming stages after what the rep did", "should_trigger": true },
{ "query": "we want stage definitions two people would score the same way", "should_trigger": true },
{ "query": "our exit criteria sound good on paper but nobody can check them", "should_trigger": true },
{ "query": "review the stage document we published for the sales team", "should_trigger": true },
{ "query": "how do I change our stage picklist without wrecking historical reporting", "should_trigger": true },
{ "query": "our high-velocity pipeline has enterprise-style gates bolted onto it", "should_trigger": true },
{ "query": "do we have too many stages", "should_trigger": true },
{ "query": "our stage definitions were copied out of a blog post years ago", "should_trigger": true },
{ "query": "tell me which stages to cut and which criteria to rewrite", "should_trigger": true },
{ "query": "the buyer's actual journey and our pipeline stages don't line up at all", "should_trigger": true },
{ "query": "every deal jumps straight from stage 2 to stage 5", "should_trigger": true },
{ "query": "I need a severity-ranked findings report on our sales stages", "should_trigger": true },
{ "query": "what would make our stage data trustworthy enough for finance to use directly", "should_trigger": true },
{ "query": "our stages measure what we did, not what the customer did", "should_trigger": true },
{ "query": "put criteria on our pipeline that a manager can actually verify", "should_trigger": true },
{ "query": "design our revenue funnel model from scratch, we have nothing today", "should_trigger": false },
{ "query": "what stages should our funnel have, we're a brand new company", "should_trigger": false },
{ "query": "build a bowtie funnel model for our SaaS", "should_trigger": false },
{ "query": "define MQL, SQL and opportunity for a funnel we haven't built yet", "should_trigger": false },
{ "query": "set the conversion rate assumptions behind next year's revenue plan", "should_trigger": false },
{ "query": "design the marketing-to-sales ownership handoffs in our funnel", "should_trigger": false },
{ "query": "why did we miss the number last quarter", "should_trigger": false },
{ "query": "our commit number is never right, diagnose the forecast", "should_trigger": false },
{ "query": "are my reps sandbagging or is demand actually down", "should_trigger": false },
{ "query": "the manager roll-up always overrides what the reps submit", "should_trigger": false },
{ "query": "run a stale deal audit before Thursday's QBR", "should_trigger": false },
{ "query": "clean up the pipeline, there are dead deals everywhere", "should_trigger": false },
{ "query": "list every deal whose close date has pushed three or more times", "should_trigger": false },
{ "query": "flag open opportunities missing forecast-critical fields", "should_trigger": false },
{ "query": "who owns each field in our CRM and how often must it be refreshed", "should_trigger": false },
{ "query": "build us a CRM field dictionary", "should_trigger": false },
{ "query": "which system is authoritative for account industry, the CRM or billing", "should_trigger": false },
{ "query": "our custom field count hit 400, help us prune it", "should_trigger": false },
{ "query": "finance and sales report two different ARR numbers", "should_trigger": false },
{ "query": "set up data contracts between marketing ops and sales ops", "should_trigger": false },
{ "query": "our MQL threshold lets far too much junk through", "should_trigger": false },
{ "query": "build a lead scoring model with fit and engagement signals", "should_trigger": false },
{ "query": "sales rejects most of the MQLs we send them", "should_trigger": false },
{ "query": "our round robin is lopsided, some reps get double the leads", "should_trigger": false },
{ "query": "who should get this inbound lead, the territory owner or the named account owner", "should_trigger": false },
{ "query": "write the territory assignment rules for EMEA", "should_trigger": false },
{ "query": "build a discount approval matrix with delegation of authority", "should_trigger": false },
{ "query": "we need a deal desk process for non-standard contracts", "should_trigger": false },
{ "query": "green accounts keep churning, our health score is broken", "should_trigger": false },
{ "query": "design a composite customer health score with weighted signals", "should_trigger": false },
{ "query": "what are the leading indicators of churn in our usage data", "should_trigger": false },
{ "query": "our champion left the account, how should that change the risk rating", "should_trigger": false },
{ "query": "where are we losing revenue between signup and activation", "should_trigger": false },
{ "query": "size the revenue we leak at renewal every year", "should_trigger": false },
{ "query": "what KPIs should the board see each quarter", "should_trigger": false },
{ "query": "build a metric tree that rolls up from IC to board", "should_trigger": false },
{ "query": "our board deck reads like a data dump, restructure it", "should_trigger": false },
{ "query": "structure the narrative for our monthly revenue review", "should_trigger": false },
{ "query": "we have too many overlapping GTM tools", "should_trigger": false },
{ "query": "audit our revops tool stack ahead of renewal season", "should_trigger": false },
{ "query": "the CSM starts from zero after every closed-won", "should_trigger": false },
{ "query": "design the closed-won to kickoff handoff packet", "should_trigger": false },
{ "query": "write a scorecard for our first RevOps hire", "should_trigger": false },
{ "query": "interview questions for a sales ops manager candidate", "should_trigger": false },
{ "query": "how do I break into RevOps from a sales role", "should_trigger": false },
{ "query": "which revops newsletters and podcasts are worth my time", "should_trigger": false },
{ "query": "starting a brand new revops project, where do I begin", "should_trigger": false },
{ "query": "audit the stages in our CI/CD pipeline", "should_trigger": false },
{ "query": "our data pipeline has a stage that keeps failing, review the stage definitions", "should_trigger": false },
{ "query": "define the stages of our hiring pipeline and what each one screens for", "should_trigger": false },
{ "query": "review the stage gates in our product development process", "should_trigger": false },
{ "query": "write exit criteria for our sprint definition of done", "should_trigger": false },
{ "query": "audit our customer onboarding stages and where new users drop off", "should_trigger": false }
]
}
references/audit-intake.md›
# Audit Intake
Look at the data before talking to anyone - it doesn't negotiate. Then interview to explain what the data shows.
## Artifact checklist
Collect before judging any stage:
- Stage list with each stage's written definition, exactly as documented. No written definitions is itself a Critical finding - record it, don't reconstruct silently.
- Stage-change history export (from-stage, to-stage, date, actor) for the trailing 2-4 quarters.
- Forecast-category mapping per stage.
- Required fields per stage, validation rules, and every automation or workflow keyed to a stage value (these break during migration if unaudited).
- 20-30 recent closed-won and closed-lost deals, with loss reasons if captured.
- Current open deals per stage (count and value).
- Close-date change history, if the system tracks it.
- Pipeline inventory: how many pipelines exist and which motions share one. Multiple motions sharing one stage list is a common root cause.
- Any published process documentation. Benchmark it against what a complete stage document contains - GitLab's public commercial sales handbook is the exemplar structure: per stage, a definition, who's involved, typical activities, system/validation-rule enforcement, exit criteria, and cross-links to adjacent process docs (qualification framework, forecasting definitions).
How to collect it:
- **System queryable directly:** pull these first-hand and confirm the pull with the user.
- **Otherwise:** request exports and work from samples.
### When the full set is unobtainable
The stage list with its written definitions is a gate, not a menu item - without it there is nothing to audit, and its absence is itself the Critical finding. Rank only what remains, by evidence of subjective staging bought per export the user has to go and obtain:
- value (evidence strength, most first): `stage-change history > closed deal samples > stage-keyed automations and validation rules > forecast-category map > close-date change history > open-deal counts per stage`
- effort (most first): `stage-change history > stage-keyed automations and validation rules > close-date change history > closed deal samples > forecast-category map > open-deal counts per stage`
- efficiency (best first): `closed deal samples > forecast-category map > stage-change history > open-deal counts per stage > stage-keyed automations and validation rules > close-date change history`
What this order starves is the stage-change history export. It carries the per-rep variance split, the single sharpest evidence that definitions are applied subjectively, and it is also the export an admin is slowest to produce. A ratio therefore defers it, and the audit ends up arguing subjectivity from anecdote.
Promote it to first whenever the trigger symptom is managers disagreeing on staging or a forecast miss nobody can explain: those two questions have no answer without it.
Strike the automation and validation-rule inventory from the request list, and say it is struck, when the Interview has already ruled out any migration rung: it is remediation input, not audit evidence, and requesting it anyway spends the admin's goodwill on work that will not ship.
## Interview guide
Interview after the first data pass, so questions target observed anomalies. Three reps minimum, plus managers, the admin/RevOps owner, and finance.
**Reps (3+, plus a walkthrough of 2+ recently won deals each):**
- Walk me through this won deal: what did the buyer actually do, in order, from first conversation to signature?
- At each stage move, what made you decide it was time? What evidence did you have?
- Which stage do you find ambiguous - where could this deal have sat in either of two stages?
- Which fields do you fill in because they're required rather than because they're true?
**Sales managers:**
- Where do pipeline reviews break down - which stages trigger the longest debates?
- Which forecast categories do reps misuse, and in which stages?
- Which fields are usually blank? Where do deals stall by segment?
- For a sample of 10 open deals: without conferring with the deal owner, is each deal correctly staged? (This seeds the inter-rater baseline.)
**RevOps / CRM admin:**
- What is keyed to stage values today - automations, validation rules, dashboards, integrations?
- What happened the last time stages changed? What broke?
- How does the system record stage history, and does duration reporting count re-entries?
**Finance:**
- How is CRM stage data consumed in the forecast that reaches leadership - direct, or reconciled in a separate rollup?
- What would need to be true for finance to trust stage-weighted pipeline directly?
## Ownership map
Confirm before proposing changes - a correct fix with no approver ships never:
- **RevOps** (or Sales Ops where no RevOps function exists) typically owns the definitions and governance layer.
- **Sales leadership** sponsors outcomes and adoption expectations; it does not administer the picklist.
- **CRM admin/IT** owns technical implementation and validation rules.
- **Finance** consumes stage data for the official forecast and must sign off on anything that changes forecast-category rollups.
- **Marketing Ops / CS Ops** own the handoff points feeding the first stage and following closed-won.
Ask explicitly: who approves a stage-definition change here? If no one can answer, record it as a governance finding - the audit's recommendations need an owner and an approval path to land.
references/diagnostic-metrics.md›
# Quantitative Diagnostics
Compute from stage-change history wherever a direct query or an export is available; otherwise approximate from a sample of 20-30 closed and 20-30 open deals and mark each unmeasured diagnostic as such. Always compute conversion on cohorts of deals that _entered_ a stage (flow), never on a point-in-time snapshot (stock) - snapshots overweight stuck deals.
| Diagnostic | How to compute | Red-flag pattern | Provenance |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Stage-to-stage conversion | Of deals entering stage N in a period, % that ever reach N+1 | Non-monotonic decay (a later stage converts worse than expected vs its neighbors with no designed reason); a mid stage converting near 100% (pass-through stage, criteria vacuous) | Practitioner consensus |
| Stage aging / duration | Median days-in-stage per stage; flag deals above ~1.5x that stage's median | One stage's outlier count dwarfs the rest - its exit criterion is either unreachable or ignored | Consulting consensus (the 1.5x flag is a rule of thumb, not a standard) |
| Stage-skip rate | % of stage transitions that jump over one or more stages | High skip rate over a specific stage - that stage's criteria are redundant or the stage maps to no buying job | Practitioner consensus |
| Backward moves | % of transitions that regress to an earlier stage | Frequent regressions from one stage - deals enter it on optimism, not evidence | Practitioner consensus |
| Close-date pushes | Count of close-date changes per deal | 2+ pushes on a deal warrants a forecast review; a stage where pushes cluster has an exit criterion detached from buyer reality | Practitioner consensus |
| Win rate by farthest stage reached | Closed-won % segmented by the deepest stage a deal entered | Later stages not materially improving win odds - stages carry no information | Practitioner consensus |
| Open-pipeline distribution | Share of open deals (count and value) per stage | One stage holding an outsized share relative to its expected dwell time - derive the expected shape from this pipeline's own history, never an industry constant | Derived from user's own data |
| Per-rep variance | Conversion and time-in-stage per stage, split by rep | Same stage, materially different conversion across reps with similar territories - the sharpest single indicator that definitions are applied subjectively | Practitioner consensus |
| Forecast accuracy | Absolute % difference between the day-one forecast for a period and the period's final result; measure at least quarterly | Accuracy not improving after the fix once a full cycle of new-definition deals has closed | Standard metric definition, widely used; the definition is uncontroversial, but the miss-rate percentages circulating alongside it are unattributed - never quote them |
## Interpretation cautions
- **Provenance discipline.** Every red-flag threshold above is tagged. When reporting, keep the tag next to the number. Replace rules of thumb with baselines derived from the pipeline's own distribution (e.g. p90 of time-in-stage) as soon as the data supports it.
- **First-entry duration.** Some systems report only the first occurrence of a stage in duration reports; a deal that regressed and re-entered shows understated time. Check how the export counts before trusting aging numbers.
- **Transaction pipelines.** Expect compressed durations and higher automation; aging flags in hours/days, not weeks. The conversion-decay and per-rep-variance logic is unchanged. For self-serve funnels with no rep-owned stages, these diagnostics become funnel conversion and time-to-activation analyses on the lifecycle stages instead - same math, different object.
- **Timing.** Conversion and distribution metrics stabilize only after a full sales cycle of deals flows through the new definitions - multi-quarter for enterprise. Report interim 30/60/90-day checkpoints as adoption reads, not verdicts.
- **Compliance is not quality.** Logging rate and field-fill rate respond to pressure: reps log more, faster, and worse. Cross-check against independent signals (calendar attendees, attached artifacts, product events) before treating any self-reported field as evidence.
references/remediation-rollout.md›
# Remediation and Rollout
Sequence the fixes by the efficiency order in SKILL.md § Remediation Order, which already carries the axes, the rungs, and what the order starves - severity grades a finding's value, not its queue position. This file supplies the mechanics each rung has to execute. Write the new exit criteria before touching any stage name - definitions first, names after; names often stop mattering once criteria are right.
## Migration mechanics - the rule that protects history
**Deactivate and add. Never rename or delete an existing stage value.**
In most CRMs, opportunity stage history is immutable and every historical record references the stage value it carried at the time. Renaming a value makes before/after reports show mismatched labels forever; deleting one can destroy historical reporting outright. There is no undo. The safe sequence:
1. Inventory everything keyed to stage values: automations, validation rules, workflows, dashboards, report filters, integrations, forecast-category mappings. Renamed or removed values break these silently.
2. Create the new stage values alongside the old ones; map old-to-new on paper first.
3. Move open deals to the new values; deactivate (never delete) the old values so no new deal can enter them.
4. Re-check the forecast-category mapping for **every** stage - new and surviving - after the change. A missed re-map skews forecast rollups while every stage name looks correct.
5. Update dashboards and reports; export/back up report definitions before the change.
6. Time the cutover for a low-traffic period, never end of quarter.
## Historical comparability
Full restatement of history is effectively blocked by platform mechanics (immutable stage history; systems that stamp stage changes at migration time, not the original date), so most teams get a **cutover date** by default. Make it deliberate:
- Document the cutover date and annotate every dashboard that spans it.
- Add a bridge field ("legacy stage" or "definition version") on open deals so cross-era cohorts can still be compared.
- Report old-definition and new-definition cohorts separately until the old cohort closes out; never blend them in one conversion number.
## Forecast rebaselining
Stage-weighted pipeline computed under new definitions is a different quantity. Capture the last old-definition forecast as a frozen baseline, tell finance the series restarts at cutover, and expect one to two periods where accuracy comparisons are noisy before the new series stands alone.
## Training and rollout
- Communicate the new criteria to reps _before_ they see them live - a criterion first encountered as a validation error gets resented, then gamed.
- Build the exit criteria into the recurring pipeline review agenda. Criteria that never come up in reviews get ignored; reps attend to what managers inspect.
- Set 30/60/90-day adoption checkpoints under either rollout shape.
- Re-run the inter-rater sample (Pass Threshold in SKILL.md) at each checkpoint; agreement is the earliest trustworthy signal, long before conversion data matures.
### Pilot versus big-bang
No strong published guidance exists for stage changes specifically - say so rather than inventing one. The ordering below is derived from the migration mechanics above, not from a source.
- value (risk avoided and time to a trustworthy signal, most first): `pilot > big-bang` - the pilot reads inter-rater agreement on one team after a single cycle, which is the earliest trustworthy signal available, and it confines a wrong criterion to that team's deals instead of the whole pipeline's.
- effort (most first): `pilot > big-bang` - the pilot runs two definitions at once behind a definition-version field, splits reporting for its length, and cuts over twice.
- efficiency (best first): `pilot > big-bang`, but only under the freeze condition below; without it the two invert.
Pilot when a team can be spared **and** its findings can be absorbed as wording changes only. Freeze the stage values before the pilot starts. A pilot whose findings force a second picklist change inside the same year costs more than the error it caught: it breaks the trend series twice, and the second break lands on data the first one already restarted.
Delete the pilot from the plan, rather than ranking it second, when no team can be spared - an organisation that cannot staff a parallel definition does not have a slower option, it has one option. Say the pilot is deleted and why, then plan a single big-bang cutover with heavy review-cadence support for a quarter. A pilot left in the plan as a nice-to-have returns as an expectation nobody resourced.
This ordering is a default, not a law, and shifts with who executes it.
- **Dedicated CRM admin:** makes the two cutovers cheap and widens the pilot's lead.
- **Pipeline small enough to re-stage by hand:** removes most of the pilot's advantage, because the whole migration is already reversible in an afternoon.
## What goes wrong
- Stage-keyed automations fire wrong or not at all because the inventory in step 1 was skipped.
- The forecast-category re-map is missed and rollups skew for a quarter before anyone notices.
- Reports spanning the cutover blend old and new cohorts into one meaningless conversion line.
- Success gets declared on logging-compliance metrics at day 14; conversion data a quarter later shows nothing changed.
- The team relapses: criteria drift back to opinion because reviews stopped inspecting them. Schedule the quarterly inter-rater re-check as a standing calendar item, not an intention.
## Optional integration note (vendor-specific)
Skip this section unless the user's platform is known. On every platform, the deactivate-and-add rule applies unchanged.
- **Salesforce:** Opportunity Stage History is retained indefinitely and cannot be selectively disabled or edited. The Stage picklist carries a Forecast Category mapping per value. Stage Duration in Opportunity History reports counts only a stage's first occurrence.
- **HubSpot:** system deal-stage timestamps cannot be backdated (migrated records stamp at migration time; use custom date properties per milestone). Moving deals across pipelines breaks time-in-stage continuity.
- **Pipedrive:** per-stage probability and rotting-days settings need explicit review after any stage change.
references/verifiability-test.md›
# The Buyer-Verifiability Test
A stage definition passes only if its exit criterion satisfies all three conditions:
1. **Buyer-sourced** - it names something the buyer did, said, or agreed to. Not something the rep did to the buyer.
2. **Observer-independent** - a third party could confirm it without asking the rep's opinion.
3. **Recorded** - the evidence lives in a checkable place: a field, an attached document, a logged meeting with the named attendee, a signed artifact, a product event.
Operational form: **could two different managers, looking at the same deal record, independently reach the same verdict on whether the criterion is met?** If the answer depends on who is looking, the criterion fails.
## Passing vs failing wording
| Verdict | Criterion | Why |
| ------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------- |
| FAIL | "Demo scheduled" | Rep activity; says nothing about buyer conviction |
| FAIL | "Proposal sent" | Rep pressed a button; buyer may never open it |
| FAIL | "Presentation delivered" | Rep task completion |
| FAIL | "Strong relationship with champion" | Unobservable opinion |
| FAIL | "Rep has confirmed budget" | Rep-attested; not independently checkable |
| FAIL | "Buyer is engaged" | No observable event; two managers will disagree |
| PASS | "Buyer named the budget owner and made the introduction (intro email or meeting logged with that person as attendee)" | Buyer act, observable, recorded |
| PASS | "Buyer's security team completed its review (assessment doc or approval email attached)" | Buyer organization acted; artifact exists |
| PASS | "Buyer agreed in writing to a mutual evaluation plan with dates (plan attached, buyer edits or reply visible)" | Written buyer commitment |
| PASS | "Buyer admin invited two teammates and connected billing details" (transaction/PLG) | Product events; machine-recorded |
| PASS | "Buyer proposed the contract-review call and put it on their own calendar" | Buyer-initiated, calendar-verifiable |
## Writing criteria that pass
- Start every criterion with the buyer as the actor: "buyer confirmed…", "buyer's legal returned…", "buyer scheduled…".
- Name the evidence location in the criterion itself - a criterion without a home field or artifact passes review and fails in production.
- Keep 2-4 criteria per stage (practitioner consensus). One is too easy to game; five dilute attention.
- Gate advancement the way MEDDPICC is commonly operationalized: score each criterion red/yellow/green, require all-green to advance. Yellow means "rep believes it, evidence pending" - a useful coaching state, never a passing one.
- The rep still records the entry; verifiability means the entry points at something a manager can check, not that the rep is distrusted.
## Gartner buying-job mapping
Map every stage to the buying job its exit evidence proves complete.
- A stage that maps to no job is rep activity in disguise.
- Several stages crowding one job are redundant.
| Buying job | Buyer-verifiable evidence that it happened |
| ---------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| Problem Identification | Buyer articulated the problem and its cost in their own words (notes quoting the buyer, discovery recording) |
| Solution Exploration | Buyer requested evaluation access, attended a working session they scheduled, named the alternatives they are comparing |
| Requirements Building | Buyer shared written requirements, an RFP, or success criteria they authored |
| Supplier Selection | Buyer confirmed shortlist status in writing; buyer's procurement engaged with the vendor |
| Validation | Buyer's technical/security/legal teams completed their checks; references were taken by the buyer |
| Consensus Creation | Buyer introduced the economic buyer or additional stakeholders; buying group co-signed the mutual plan |
Buyers loop through these jobs non-linearly, so evidence may arrive out of stage order - the map validates that each stage's exit proves _some_ job, not that jobs happen in sequence.
## The scorecard anti-pattern
Do not convert a qualification framework into the stage list. Force Management's warning: there are no "MEDDIC stages 1-6."
Each qualification dimension (pain, champion, economic buyer, decision process…) is a separate field updated continuously through the deal. The stage records where the _buyer_ is, not the scorecard.
A pipeline whose stages are scorecard letters can no longer answer "where is this buyer in their purchase?", the one question stages exist for.
## Negative example - a rewrite that still fails
A team replaces "Proposal Sent" with "Value confirmed and proposal in play." It sounds buyer-centric, but fails all three conditions:
- No buyer act is named.
- "Value confirmed" is the rep's opinion.
- Nothing is recorded.
Two managers reviewing the same deal disagree immediately.
The passing rewrite names the act and the artifact: "Buyer's evaluation lead confirmed in writing that the proposal matches the requirements they authored (email attached), and named the decision date."
references/worked-examples.md›
# Worked Examples
Fictional companies; structures follow the Output Shape in SKILL.md, condensed.
## Example 1 - Enterprise deal pipeline (B2B sales-led)
```
STAGE DEFINITION AUDIT - Northbeam Analytics "Enterprise" pipeline, 2026-03
Pipeline type : deal pipeline; sales-led, $60-150k ACV, ~120-day cycle
Stage inventory : 8 stages, definitions in a 2022 slide deck reps have not seen;
forecast-category map exists but unreviewed since then
Per-stage verdict:
1 Discovery Scheduled -> "intro call booked" -> FAIL -> Critical
2 Discovery Done -> "discovery call held" -> FAIL -> Critical
3 Demo Delivered -> "demo completed" -> FAIL -> Critical
4 Champion Identified -> "rep confirmed champion" -> FAIL -> High
5 Proposal Sent -> "proposal emailed" -> FAIL -> Critical
6 Verbal Commit -> "buyer said yes verbally" -> FAIL -> Medium (buyer
act, but unrecorded and unwitnessable)
7 Contract Sent -> "contract emailed" -> FAIL -> Critical
8 Closed Won/Lost -> signature / loss reason -> PASS
Buying-job map : stages 1-3,5,7 map to no buying job (rep tasks); Validation and
Consensus Creation have no stage at all - the two jobs where
enterprise deals actually die
Diagnostics : Proposal Sent holds 47% of open pipeline value (own-history
expected ~20%); conversion decay non-monotonic (stage 5->6
converts 18%, 4->5 converts 92% - stage 4 is a pass-through);
per-rep stage-4->5 conversion ranges 55-95% (subjectivity flag,
practitioner-consensus reading); median 71 days in Proposal Sent
Findings : 7 records, ordered by Remediation Order (top: 5 criterion
rewrites, near-zero each, retiring 4 Critical and 1 High)
Remediation : rung 1 first - rewrite all 5 rep-activity criteria as buyer acts
against existing fields; e.g. stage 5's exit becomes "buyer's
evaluation lead confirmed in writing the proposal matches their
authored requirements" + "buyer security review complete (doc
attached)" + "mutual plan co-signed with dates". Rung 3: two of
those criteria need an evidence-link field that does not exist.
Rung 4 last - collapse 8 stages to 6, deactivate-and-add, re-map
forecast categories for all 6, inventory of 14 stage-keyed
automations attached. Redefinition not promoted: stages 6 and 8
survive as anchors and the motion is single. Pilot with the 6-rep
mid-market team for one cycle, stage values frozen beforehand
Measurement : baselines frozen 2026-03; inter-rater sample (2 managers x 15
deals) at day 30/60/90; conversion re-read after one full cycle
(~Q4 2026)
```
## Example 2 - Transaction pipeline (high-velocity, sales-assisted PLG)
```
STAGE DEFINITION AUDIT - Loopdesk "Inside Sales" pipeline, 2026-03
Pipeline type : transaction pipeline; $2.4k ACV, 12-day cycle; self-serve funnel
exists separately and is OUT OF SCOPE - kept as a distinct
reporting construct, not merged
Stage inventory : 5 stages, definitions in the team wiki; forecast map present
Per-stage verdict:
1 New Signup Assigned -> "trial account routed to rep" -> PASS (entry gate,
machine-recorded)
2 Activated -> "buyer completed setup + invited a
teammate" (product events) -> PASS
3 Demo Booked -> "rep booked a demo" -> FAIL -> High
4 Quote Accepted -> "buyer accepted quote in portal" -> PASS
5 Closed Won/Lost -> payment method charged / lapsed -> PASS
Buying-job map : compressed by design - Solution Exploration and Validation
collapse into stage 2's product events; appropriate for the
motion, noted rather than flagged
Diagnostics : aging measured in days not weeks (median stage 3 dwell: 6 days,
flags at >9, derived from own p90); stage 3 skip rate 38% -
deals that skip it convert BETTER (7-day cohorts), evidence the
stage adds friction without information
Findings : 2 records. Top: stage 3 is rep activity AND the motion's data
says it is optional - rung 1 replaces it with "buyer requested a
call or asked a pricing question in-app" (buyer-initiated,
logged); rung 4 cuts it outright
Remediation : cut chosen over rewrite - re-ranked, not defaulted: 5 reps and a
12-day cycle make the migration an afternoon's work rather than a
quarter's, which promotes the cut above the rewrite, and the skip
data says the rewrite would preserve a stage carrying no
information. 4-stage candidate structure;
deactivate-and-add; no MAP or security-review criteria imported
from enterprise practice. Pilot deleted, not deferred: a 5-rep
team cannot spare a parallel definition - big-bang with two weeks
of review-cadence support instead
Measurement : one full cycle = ~2 weeks, so conversion re-read at day 30;
inter-rater check replaced by automated-event spot audit (most
criteria are machine-recorded)
```
## Negative example - an audit done wrong
A consultant audits a 7-stage pipeline and, in one afternoon:
1. **Renames** "Demo Scheduled" to "Solution Validation" directly on the live picklist. The damage to before/after reporting is permanent:
- Every historical record keeps the old label in stage history.
- Every report spanning the change now shows two labels for one stage.
- Three stage-keyed workflows silently stop firing.
2. Replaces the stages with **"MEDDIC 1" through "MEDDIC 6"**. The pipeline now tracks scorecard completion, not buyer position - nobody can answer "where is this buyer in their purchase?", and the qualification data lost its per-dimension fields in the process.
3. Sets a stale-deal threshold of exactly 14 days "because the industry standard is 14 days" - a vendor blog number presented as a constant, with no look at the pipeline's own dwell-time distribution.
4. Skips the forecast-category re-map. The quarter's forecast rollup silently drops the renamed stages' weighted value; finance discovers it at quarter close.
5. Declares success at day 14 because field-fill rate rose from 60% to 95%. A quarter later, per-rep conversion variance is unchanged - reps fill the fields to pass validation and stage on optimism exactly as before.
Every step above violates a rule this skill states:
- Deactivate-and-add, never rename or delete.
- Stage ≠ scorecard.
- Provenance-tag thresholds and derive them from the pipeline's own data.
- Re-map forecast categories after any change.
- Never measure success by compliance rate.
SKILL.md›
---
name: pipeline-stage-definition-audit
description: Audit existing sales pipeline stage definitions against buyer-verifiable milestones and flag the ones built on rep activity instead. Core test - does each stage exit criterion name something the buyer did, said, or agreed to, checkable by two managers independently? Adds stage aging, conversion decay, stage-skip rate, and close-date push diagnostics. Use whenever the user mentions stage definitions, stage exit criteria, stage inflation, stages named after rep activity ("demo scheduled", "proposal sent"), "our stages don't mean anything", or "why is everything stuck in one stage" - even if they never say "audit". Covers enterprise and high-velocity pipelines, B2B and B2C. Do NOT use for designing a funnel stage set from scratch - use mbfinotti/revops-skills@revenue-funnel instead.
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.1.10"
---
# Stage Definition Audit
Evaluate whether an existing pipeline's stage definitions reflect buyer-verifiable milestones - things the buyer did, said, or agreed to - rather than internal rep activity, and deliver a severity-ranked findings report with rewritten exit criteria and a safe migration plan.
Reference frameworks:
- **Gartner's six B2B buying jobs** (buyer-side reference structure): Problem Identification, Solution Exploration, Requirements Building, Supplier Selection, Validation, Consensus Creation.
- **GitLab's public commercial sales handbook** (real-world exemplar of a published stage-definition document): definition, who's involved, activities, system enforcement, exit criteria, cross-links.
- **MEDDPICC-style gating**: each exit-criterion element scored red/yellow/green, all-green to advance.
- **Winning by Design's SPICED** (buyer-side diagnostic lens): its elements are confirmed with the buyer, not assumed by the rep.
## Ground Rules
- Audit, never redesign. The existing stage set is the input. If the funnel model itself is the defect, say so and route to the funnel-design see-also (Reference) instead of rebuilding it here.
- Label every threshold with its provenance: published company practice, practitioner consensus, or derived from the user's own data. Never present a vendor heuristic as an industry constant.
- Warn off the circulating forecast statistics ("93% of teams can't predict revenue within 5%", "85% of B2B companies miss forecast", "CRM data is only 40-60% accurate") - each circulates unattributed and must never be stated as fact. Cite the mechanism instead: when advancement is subjective, two reps stage identical deals differently, which corrupts stage conversion data and everything built on it.
- The stage records the buyer's position, never a qualification scorecard. There are no "MEDDIC stages 1-6" (Force Management's warning): each qualification dimension is its own continuously-updated field; conflating the scorecard with the stage picklist breaks both.
- Classify the pipeline type (below) before judging any stage. A transaction pipeline judged by deal-pipeline expectations produces false findings.
- Never rename or delete an existing stage value. In most CRMs, stage history is immutable and historical records reference old values forever - renaming or deleting corrupts before/after reporting permanently. Deactivate the old value and add a new one. After any change, re-check the forecast-category mapping for every stage.
- 5-7 stages with 2-4 exit criteria each is the practitioner-consensus band, not a law. A stage earns its place by having a distinct purpose; if nothing changes in it, cut it.
## Pipeline Types
Establish the pipeline's type before judging a single stage - the three types come from The RevOps Show (Doug Davidoff and Jess Cardenas):
- **Development pipeline** - moves an account from no intent to intent. Stages track buyer interest signals, not deal mechanics.
- **Deal pipeline** - intent to decision. The classic enterprise B2B model this audit's examples assume by default.
- **Transaction pipeline** - high-velocity, lower-value, friction-reduction focus. Fewer stages, lighter criteria, often automated progression signals.
High-velocity, self-serve, PLG, and B2C teams frequently run no rep-owned pipeline at all: they run buyer-facing lifecycle/funnel stages owned by marketing or growth. A true pipeline construct reappears only when the motion goes sales-assisted or upmarket, and then funnel and pipeline must coexist as distinct reporting constructs, never merged into one stage list.
- **Identical across B2B and B2C:** the verifiability test itself, the stage-aging and conversion diagnostics, and the deactivate-don't-rename migration rule.
- **Divergent across B2B and B2C:** stage count, evidence type (product events and payment signals versus signed documents and security reviews), and ownership.
Named-practitioner literature on B2C stage definition is thin, so the three-pipeline-type framing is the best available anchor; derive the rest from the user's own data and say so.
## Interview
Ask before judging anything. One question per message; multiple-choice where possible; skip anything already answered.
- Paste the current stage list with each stage's written definition, exactly as documented. If definitions live only in people's heads, say so - that is itself a finding.
- Which motion feeds this pipeline: enterprise sales-led, mid-market, high-velocity inside sales, PLG sales-assisted? One pipeline or several, and do multiple motions share one?
- ACV band and typical cycle length?
- What symptoms triggered this audit: deals stuck in one stage, forecast misses, managers disagreeing on staging, stage skipping, silent close-date pushes?
- What is the forecast-category mapping per stage?
- Which required fields, validation rules, and automations are keyed to stage values?
- What data can be exported or queried: stage-change history, closed-won/lost records, close-date change history, per-rep breakdowns?
- Who owns stage definitions and who approves changes (RevOps, sales leadership, finance)?
- Are two people available to score the same deals independently - two managers, or any two reviewers who can judge a deal record without conferring? The Pass Threshold's inter-rater check needs them. If only one reviewer exists, say so now: the check becomes an automated-event spot audit against machine-recorded criteria, and the deliverable must state that substitution rather than claim an agreement rate it never measured.
- When were stages last changed, and how did that rollout go?
- Can you provide samples: 20-30 recent closed deals and the open deals currently sitting in each stage?
- By what date must the result land, and does that date fall inside a live quarter? A stage change corrupts the quarter it lands in, so a mid-quarter deadline strikes every migration rung in Remediation Order and leaves criterion rewrites only.
- Do you want a one-off correction before a specific forecast, or a definition set that holds every quarter? One-off promotes the criterion rewrite alone; a compounding mandate promotes the evidence plumbing and makes the standing inspection non-optional.
- What is the effort ceiling: is a CRM admin available, how many reps must be retrained, and can the pipeline's trend series afford to restart at a cutover date? No admin strikes the evidence-plumbing rung; no tolerance for a trend break strikes the stage cut and the stage add.
## Workflow
1. Run the Interview; classify the pipeline type and confirm the scope boundary (audit, not redesign).
2. Gather artifacts and run stakeholder interviews per [references/audit-intake.md](references/audit-intake.md). Look at the data before talking to anyone.
3. Apply the verifiability test from [references/verifiability-test.md](references/verifiability-test.md) to every stage: the exit criterion must name something the buyer did, said, or agreed to; be observable independently of the rep's opinion; and be recorded somewhere checkable. Operational form: could two managers, looking at the same deal record, independently reach the same verdict on whether the criterion is met?
4. Map each stage to a Gartner buying job (table in the same reference). Flag stages that map to no job (rep activity in disguise) and multiple stages crowding one job (redundant stages).
5. Compute the diagnostics in [references/diagnostic-metrics.md](references/diagnostic-metrics.md) where the pipeline's data can be queried directly; otherwise derive what the deal samples and interviews support, and mark each unmeasured diagnostic as such.
6. Record every defect as a finding with this schema: **Stage | Current exit criterion | Verdict + why | Severity | Fix rung + effort | Evidence | Buying job mapped | Proposed criterion | Migration note**. Severity scale:
- **Critical** - exit criterion is pure rep activity or undocumented; advancement is opinion.
- **High** - criterion references the buyer but is unverifiable or recorded nowhere checkable.
- **Medium** - verifiable but ambiguous; two managers could plausibly disagree.
- **Low** - criterion sound; defect is naming, ordering, or redundancy.
Severity is the value axis of a finding, never its queue position. A Critical defect fixed by one wording change ships before a Medium one that needs a picklist migration. Tag each finding with the rung it lands on and that rung's effort order of magnitude, then order the list by Remediation Order below.
7. Draft rewritten exit criteria for every failing stage (wording guidance in the verifiability reference), then build the remediation and rollout plan from [references/remediation-rollout.md](references/remediation-rollout.md) - fixes sequenced by Remediation Order, deactivate-and-add migration, forecast-category re-mapping, automation audit, training and rollout shape.
8. Emit the report (Output Shape below) one section at a time for user validation; ground it in the matching worked example from [references/worked-examples.md](references/worked-examples.md).
9. Check the Pass Threshold; iterate the criteria and plan until it holds or the remaining gaps are explicitly scheduled (e.g. metrics that need a full sales cycle of new data).
10. If your harness has persistent memory, store the stage inventory, per-stage verdicts, and agreed thresholds so the post-rollout re-check starts from them instead of re-interviewing.
## Remediation Order
Six fix types, ordered by forecast accuracy recovered per unit of effort - never by severity, and never cheapest-first. Effort here is admin configuration, retraining every rep who works the stage, historical data migration, and reversibility. Reversibility dominates this list: a stage definition changed twice in a quarter destroys the trend data the definitions exist to produce, so any fix that touches a stage value is close to one-way.
- value (accuracy recovered, most first): `full redefinition > criterion rewrite > evidence plumbing > stage add == stage cut > standing inspection`
- effort (most first): `full redefinition > stage add > stage cut > evidence plumbing > standing inspection > criterion rewrite`
- efficiency (best first): `criterion rewrite > standing inspection > evidence plumbing > stage cut > stage add > full redefinition`
`stage add == stage cut` on value: pipeline type decides which one wins, and nothing else does.
- **Deal pipeline:** recovers more from the add, since enterprise deals die in Validation and Consensus Creation, the jobs most often left uncovered.
- **Transaction pipeline:** recovers more from the cut, since friction removal is that motion's design goal, and a stage deals skip is a stage that converts better skipped.
They do not tie on effort: the migration is identical, and only the add asks reps for evidence they have never been asked to produce.
1. **Criterion rewrite** - restate a failing exit criterion as a buyer act pointing at a field or artifact the system already holds. No picklist change, no migration, no forecast re-map. Near-zero to an hour per stage, plus one walkthrough in the review meeting that already happens; the old wording is a document edit away. Retires most Critical and High findings on its own.
2. **Standing inspection** - exit criteria on the recurring pipeline-review agenda, plus a quarterly inter-rater re-sample. An hour a quarter, forever. Recovers no accuracy by itself, and earns its rank by being the only rung that stops every rung above it drifting back to opinion within two quarters.
3. **Evidence plumbing** - create the required field, validation rule, or evidence link a rewritten criterion names but the system does not capture. A week of admin configuration plus retraining every rep who touches that stage. Reversible in configuration, never in retraining. This is what turns a criterion that passes review into one that holds in production.
4. **Stage cut or merge** - deactivate-and-add a stage that maps to no buying job or duplicates its neighbor. A week to a quarter: open-deal moves, forecast-category re-map, and the stage-keyed automation inventory. Breaks the trend series at cutover, and reversing it breaks it a second time.
5. **Stage add** - a new stage for a buying job nothing covers. The same migration as the cut, plus a behavior reps have never been asked for and managers have never inspected. A quarter.
6. **Full redefinition** - the stage set rebuilt as a coherent series rather than patched cell by cell. Out of scope here by Ground Rules; route it rather than attempting it.
Default rung: run every criterion rewrite in one pass, then schedule the standing inspection. Take on the next rung down only when a rewritten criterion names evidence the system does not capture - there the plumbing is not a later phase, it is what makes the rewrite true.
What this order starves is the full redefinition. It ranks first on value: the only fix that treats the stage set as a series instead of a list of separately worded cells. It ranks last on the ratio in every engagement, so a severity-and-effort ranking defers it indefinitely, and the pipeline accumulates well-written criteria on a stage model that cannot carry them.
Promote it, and route out of this skill to `mbfinotti/revops-skills@revenue-funnel`, when:
- Fewer than two stages survive the audit as usable anchors.
- Multiple motions share one stage list.
- The pipeline type itself is misclassified.
Delete a ruled-out rung rather than demoting it: a rung parked at the bottom of the plan returns later as scope nobody budgeted.
- **No approver for a stage-definition change** (per the Interview): strike the cut and the add from the plan, and say they are struck. A migration with no approver ships never.
- **No CRM admin available:** strike the evidence plumbing the same way.
The order is a default, not a law: it shifts with context and with who executes it. Re-rank it against the Interview's answers.
- **CRM admin on hand:** collapses the plumbing rung from a week to an hour and moves it above the standing inspection.
- **Sales team mid-quarter:** strikes every migration rung until the quarter closes, because a stage change corrupts the quarter it lands in.
- **Pipeline small enough to re-stage by hand:** drops the cut and add rungs by an order of magnitude and promotes both above the plumbing.
## Output Shape
Every threshold line carries a provenance tag.
```
STAGE DEFINITION AUDIT - <pipeline>, <date>
Pipeline type : development / deal / transaction + motion, ACV band, cycle length
Stage inventory : count vs 5-7 consensus band; definitions documented or tribal;
forecast-category map present/absent
Per-stage verdict: stage -> exit criterion -> pass/fail verifiability -> severity
Buying-job map : which Gartner job each stage covers; jobs with no stage; jobs
with several stages
Diagnostics : conversion decay shape, aging outliers, skip/backward rate,
close-date push counts, per-rep variance (provenance-tagged)
Findings : records ordered by Remediation Order, severity carried per
record as its value axis (schema in Workflow step 6)
Remediation : rewritten exit criteria, rungs chosen and rungs struck with the
constraint that struck them, deactivate-and-add migration plan,
forecast-category re-map, automation/dashboard checklist,
training + rollout shape
Measurement : baseline values captured, post-change KPIs, re-check date
(>= one full sales cycle out)
```
## Pass Threshold
- Every stage has at least one exit criterion that names a buyer action and is attached to a checkable field or artifact in the system of record.
- Inter-rater check: two managers (or two independent reviewers) score a sample of 10-20 open deals against the rewritten criteria without conferring. Verdicts must agree on at least 90% of deals - a working bar to derive tighter from the user's own data, not a published standard. Disagreements point at the ambiguous criterion; rewrite it and re-sample.
- Stage-to-stage conversion decays monotonically across the funnel, or every exception is explained by deliberate design (e.g. a verification stage meant to disqualify).
- No open-pipeline stage holds a share of deals wildly out of line with its expected dwell time - derive the expected distribution from the pipeline's own history rather than an industry constant.
Iterate until all four hold. Conversion and distribution checks need a full sales cycle of deals under the new definitions before they are trustworthy - if that data does not exist yet, say so in the deliverable and schedule the re-check as the first post-rollout gate.
## Common Failure Modes
Deliberately unranked: each row is one diagnosis with one fix, not a menu of competing fixes for one problem, so an efficiency order over them would rank the reader's symptoms rather than their options. Apply every row that matches.
| Defect | Consequence | Fix |
| -------------------------------------------------- | ---------------------------------------------------------------------- | -------------------------------------------------------------------------- |
| Renaming or deleting stage values in place | Historical reporting corrupted permanently; automations silently break | Deactivate old values, add new ones; bridge with a mapping field |
| Skipping the forecast-category re-map | Forecast rollups silently skew while stage names look right | Re-check the mapping for every stage after any change |
| Qualification scorecard baked into the picklist | Stage stops tracking buyer position; both constructs degrade | Separate continuously-updated qualification fields; stage = buyer position |
| Criteria written as rep activities | Stage inflation; advancement on optimism | Rewrite every criterion to start with a buyer action |
| Criterion verifiable on paper, recorded nowhere | Passes the audit, unenforceable in practice | Attach each criterion to a required field or evidence link |
| Transaction pipeline judged by deal-pipeline rules | False findings; friction added to a motion built on removing it | Classify pipeline type first; weight criteria to the motion |
| More than ~7 stages or look-alike stages | Adjacent stages indistinguishable; criteria diluted | Merge; every stage needs a distinct purpose (practitioner consensus) |
| Success measured by logging/compliance rate | Reps log more, faster, worse; data no better | Measure conversion stability and inter-rater agreement instead |
| Judging the fix in week two | Conversion metrics meaningless without a full cycle of new deals | Wait at least one sales cycle; use 30/60/90-day adoption checkpoints |
| Rolling out criteria reps first see live | Ignored criteria; stage data quality unchanged | Train first; build exit criteria into the recurring pipeline review |
## KPIs
- Track: stage-to-stage conversion stability and monotonic decay, per-rep conversion variance (the sharpest single indicator of subjective definitions), median time-in-stage and outlier count, stage-skip and backward-move rate, close-date push count per deal, inter-rater agreement on periodic deal samples, and forecast accuracy measured as the absolute % difference between the day-one forecast and the period's final result, at least quarterly.
- Never report adoption or logging-compliance rates as success - they measure pressure, not accuracy.
- Re-run the inter-rater sample each quarter; a decaying agreement rate means criteria are drifting back to opinion.
## Invocation Examples
- "Our stages are 'Demo Scheduled', 'Proposal Sent', 'Verbal Commit', and every deal sits in Proposal Sent for months. Audit our stage definitions."
- "Two of my managers can't agree on whether the same deal belongs in stage 3. Review our exit criteria and tell me which stages are rep activity in disguise."
- "We're PLG going upmarket and sales-assisted deals now share a pipeline with self-serve signups. Audit whether our stages mean anything before we forecast off them."
## Reference
- [references/verifiability-test.md](references/verifiability-test.md) - the test, passing and failing wording, the Gartner buying-job mapping table, and a negative example.
- [references/diagnostic-metrics.md](references/diagnostic-metrics.md) - how to compute each metric, its red-flag pattern, and its provenance.
- [references/audit-intake.md](references/audit-intake.md) - the artifact checklist, interview guides, and ownership map.
- [references/remediation-rollout.md](references/remediation-rollout.md) - migration mechanics, cutover and comparability, training, and the optional vendor-specific integration note.
- [references/worked-examples.md](references/worked-examples.md) - one enterprise audit, one high-velocity audit, and one audit done wrong.
- `mbfinotti/revops-skills@sales-forecast-diagnostic` - overall forecast reliability
- `mbfinotti/revops-skills@sales-pipeline-hygiene` - recurring stale-deal sweep and hygiene cadence
- `mbfinotti/revops-skills@crm-data-governance` - field dictionary and picklist governance
- `mbfinotti/revops-skills@revenue-leakage` - tracing revenue lost between stages