返回 Skills 目錄
mbfinotti/revops-skills已通過檢查

SKILL DETAIL

sales-forecast-diagnostic

mbfinotti/revops-skills/sales-forecast-diagnostic

Diagnose why an existing sales forecast is unreliable and recommend fixes - stage inflation and happy ears, sandbagging, deals without verifiable buyer evidence, stale or mass-pushed close dates, wrong forecast-category assignment, zombie and duplicate opportunities, late-created deals, manager roll-up overrides, and comp incentives that reward bias - separating data-quality from behavioral from genuine demand problems. Use whenever the user mentions forecast accuracy, a forecast miss, reps sandbagging, deals that keep slipping, "why did we miss the number", or "our commit is never right" - even if they never say "forecast". Covers B2B deal-based and B2C/high-volume. Do NOT use for pipeline coverage modeling - use mbfinotti/sales-skills@sales-pipeline-coverage-modeling instead.

安裝量 · 171查看來源

Installation

npx skills add https://github.com/mbfinotti/revops-skills --skill sales-forecast-diagnostic

技能檔案

SKILL.md

最近同步 · 2026年9月15日

evals/evals.json
{
  "skill_name": "sales-forecast-diagnostic",
  "evals": [
    {
      "id": 1,
      "prompt": "I run RevOps at Corviq, B2B SaaS, 34 AEs, quarterly forecast. Q3 quarter-start commit was $6.2M and we closed $5.1M. When I pulled the won list I noticed $310K of it came from four deals that were in no forecast category at quarter start - two of them from the same AE, Priya, who has beaten her commit by a wide margin three quarters running. Our CRO is going to tell the board 'we missed by $1.1M' and he wants one number and one cause on the slide. I have quarter-start snapshots and stage, close-date and category field history going back six quarters. What do I give him?",
      "expected_output": "A decomposition that reports the $1.41M gross over-forecast on committed deals separately from the $310K of hidden upside, refuses the single-number and single-cause framing, reports signed and absolute error with a fixed denominator, and sets up a snapshot-based reconstruction with per-deal attribution.",
      "files": [],
      "expectations": [
        "States the gross over-forecast on committed deals as $1.41M, distinct from the $1.1M net gap.",
        "Attributes the $310K of wins from outside the forecast as its own offsetting error rather than netting it against the over-forecast.",
        "Refuses to deliver a single blended gap figure as the answer, stating that netting understates both problems.",
        "Rejects the 'one cause' framing and requires the gap decomposed across named deals instead.",
        "Reports both signed error (bias) and absolute error rather than a single accuracy percentage.",
        "Names which error denominator is in use (miss over actual, or actual over forecast) and commits to holding it fixed across periods.",
        "Requires at least 90% of gap dollars attributed to named deals, with the residual capped at 10% and explicitly labeled unexplained.",
        "Assigns each attributed deal exactly one primary root cause and exactly one class among data, behavior, and demand.",
        "Treats Priya's three-quarter beat pattern as a sandbagging candidate rather than as high performance, and names wins closing from outside the forecast as the sandbagging signature.",
        "Requires two independent signals per confirmed cause, at least one of them from recorded history such as a snapshot, field history, or activity log.",
        "Specifies reconstructing the period from the quarter-start snapshot rather than from the current record."
      ]
    },
    {
      "id": 2,
      "prompt": "Our VP of Sales has already decided the diagnosis and wants me to build the deck around it. His case has two parts: our win rate on forecast-weighted deals last quarter was 22% when our late stages imply about 70%, and he remembers one deal, Halvorsen Group at $480K, that the rep swore was closing and never did. His fix is mandatory qualification-framework retraining for all 28 AEs plus a new certification. We're Northvane, we sell HR software, quarterly cycle. One thing I noticed: that 22% was computed on the end-of-quarter opportunity list. I do have quarter-start snapshots. Can you help me put the deck together?",
      "expected_output": "A refusal to build the deck as framed, identifying the win rate as hindsight-contaminated, the anecdote as a non-signal, the two-signal rule as unmet, and retraining as a behavior-class fix aimed at a problem not yet shown to be behavioral.",
      "files": [],
      "expectations": [
        "Declines to build the deck around the retraining conclusion and states the verdict is unconfirmed.",
        "Identifies the 22% win rate as computed on end-of-period data and requires recomputing it on the quarter-start snapshot.",
        "Names zombie or already-dead deals padding the denominator as a data-class explanation for an artificially low win rate.",
        "States that the Halvorsen Group anecdote is not a signal and cannot support a cause.",
        "Applies the two-signal evidence rule and concludes the stage-inflation verdict fails it on the evidence presented.",
        "Identifies retraining as a behavior-class fix and warns it is being aimed at a problem not yet shown to be behavioral.",
        "Names hindsight bias explicitly: the period-start call must be judged on what was knowable at period start.",
        "Warns that a diagnosis delivered as blame drives reps to under-call, flipping the sign of the error instead of shrinking it.",
        "Requires the discriminating rep interview question - whether the rep knew the deal was dead or real before the record said so - to split data from behavior.",
        "Requires pulling stage field history and activity logs before issuing any inflation verdict.",
        "Keeps the accuracy investigation separate from performance reviews and compensation conversations."
      ]
    },
    {
      "id": 3,
      "prompt": "Sales ops at Tessaly Labs, B2B, quarterly cycle. We missed Q2 by about $740K against a $3.4M commit. Our current quarter closes in six weeks and the CEO wants the Q3 call to actually hold. He has told me plainly that anything which does not improve this quarter's number is a distraction. I have a CRM admin who can ship a validation rule or a new field this week, roughly 20 analyst hours, and sales leadership has agreed to nothing yet. Give me the fix plan.",
      "expected_output": "A plan that deletes every week-tier and quarter-tier fix from this cycle, names what was deleted, ships only the hour-tier data-class fixes in order, and states plainly that the behavior causes go untouched this cycle.",
      "files": [],
      "expectations": [
        "Deletes every week-tier and quarter-tier fix from this cycle's plan rather than listing them lower in the same plan.",
        "Names explicitly which fixes were deleted and states that the deadline is what removed them.",
        "Delivers a surviving fix set ordered zombies first, then stale or mass-pushed close dates, then wrong category assignment.",
        "States plainly that the behavior-class causes are left untouched this cycle.",
        "Treats the available CRM admin as keeping the hour-tier fixes at hour-tier effort rather than turning them into a week-tier request in someone else's queue.",
        "Notes that the close-date reason requirement only bites at the next edit, so it needs a full cycle before dates are trustworthy.",
        "Notes that the zombie purge makes coverage honest within the current cycle.",
        "Pairs the zombie purge with an entry-standard fix or a standing sweep rather than treating the purge as finished work.",
        "Names for each surviving fix its cause class, its owner, its effort tier, and the single metric that shows within one period whether it worked.",
        "Does not recommend a comp-plan change or a rep calibration scorecard as part of this cycle's plan.",
        "States which period each recommendation is aimed at rather than leaving the timing implicit."
      ]
    },
    {
      "id": 4,
      "prompt": "Brightmoor Systems, 19 AEs. We migrated off a homegrown CRM onto Salesforce five months ago and nobody turned on field history tracking. There is no such thing as a period-start snapshot here either - the forecast was a spreadsheet somebody emailed round and then overwrote each week. Our CEO is convinced two named reps are sandbagging and that one regional manager pads his roll-up, and he wants me to prove it before the comp cycle closes. Where do I start?",
      "expected_output": "A plan that removes the sandbagging and manager-override causes from this pass for lack of history, starts capturing history immediately, proceeds with current-state fingerprints plus interviews at lower confidence, and refuses to name individuals into a comp decision.",
      "files": [],
      "expectations": [
        "States that the sandbagging and manager-override causes cannot be evidenced in this pass and removes them from the fix plan rather than ranking them low.",
        "Refuses to name individual reps as sandbaggers on the available evidence.",
        "Schedules the sandbagging and override analyses for a later re-run once history exists, rather than dropping them permanently.",
        "Instructs enabling field history tracking now on stage, close date, amount, and forecast category.",
        "States that field history tracking is opt-in per field and not retroactive, so enabling it does not recover the missed period.",
        "Proceeds with a current-state fingerprint diagnosis plus interviews rather than refusing to diagnose at all.",
        "Labels every conclusion reached without recorded history as lower-confidence.",
        "States that interview testimony alone never confirms a cause.",
        "Requires a repeated pattern across at least two periods of the same person's own history before sandbagging can be named.",
        "Flags the CRM migration as a confounder for the period under investigation.",
        "Warns against feeding the accuracy finding into the comp cycle, because a diagnostic that feeds punishment teaches reps to sandbag."
      ]
    },
    {
      "id": 5,
      "prompt": "Here is where we landed at Ardenway - B2B SaaS, quarterly, $5.0M commit against $3.9M actual, $1.1M gap, 91% of it attributed. Confirmed causes with their attributed dollars: comp-driven bias $420K, stage inflation $310K, stale and mass-pushed close dates $240K, zombie opportunities $130K. We have a RevOps admin with change authority, sales leadership is bought in, and the CFO will look at the comp plan at the annual refresh in five months. Give me the fix list in the order we should ship it.",
      "expected_output": "Three orderings delivered together: the dollar impact order, the leverage order led by stale dates, and an explicit account of every cause that moved between them and the effort that moved it.",
      "files": [],
      "expectations": [
        "Produces three distinct outputs: an impact order, a leverage order, and an explicit list of what moved between them.",
        "Gives the impact order as the straight dollar ranking: comp-driven bias, stage inflation, stale dates, zombies.",
        "Leads the leverage order with stale and mass-pushed close dates, not with zombies and not with comp-driven bias.",
        "Places zombies ahead of stage inflation in the leverage order despite the smaller dollar figure, on the hour-versus-week effort difference.",
        "Places comp-driven bias last in the leverage order, on the quarter-tier plan-amendment cost.",
        "Names, for every cause that changed position between the two orders, the specific effort that moved it.",
        "States that the leverage denominator is effort - analyst hours, admin work, cross-team negotiation, forecast cycles to effect - and never a currency amount.",
        "States that leverage promotes cheap fixes but not cheapness: a larger attributed cause outranks a smaller one at the same effort tier.",
        "Delivers in leverage order with the impact order printed beside it rather than shipping one order alone.",
        "Names comp-driven bias as the root behind the behavior-class causes and states that the behavior fixes are temporary while the plan is unchanged.",
        "Attaches to each fix its cause class, owner, effort tier, and the single metric that proves it worked within one period."
      ]
    },
    {
      "id": 6,
      "prompt": "Second pass at Velmara. Last quarter we did exactly what the previous analysis said: purged 214 zombie opportunities, shipped reason-required close-date pushes, and started logging category changes. Forecast accuracy this quarter is unchanged, still about 19% over on the commit call. The new fingerprint run shows two things. Slips cluster in the fortnight right after each AE crosses 100% of quota, and the same four AEs under-call when they are ahead of plan and over-call when they are behind. Our comp plan has a hard accelerator cliff at 100%. Our CRO has told me the plan is locked until January and he will not reopen it. What now?",
      "expected_output": "Comp-driven bias promoted to the top of the narrative on all three promotion conditions, the cheap data fixes declared spent rather than re-listed, and the comp fix deleted from the plan with the regrowing behavior causes named.",
      "files": [],
      "expectations": [
        "Promotes comp-driven bias to the top of the narrative rather than ranking it last by leverage ratio.",
        "Cites the slip clustering at the quota cliff as a promotion condition.",
        "Cites the reps' bias flipping with their attainment position as a second promotion condition.",
        "Cites the shipped data-class fixes producing no improvement in the behavior class as a third promotion condition.",
        "States that the cheap data-class fixes have no remaining value here and does not re-recommend them.",
        "Deletes the comp-plan fix from the actionable plan given the CRO's refusal, rather than parking it at the bottom as a demoted option.",
        "Names the specific behavior-class causes that will regrow because the comp fix was ruled out.",
        "Does not recommend punishing forecast error through compensation or performance review.",
        "Notes that a comp-plan amendment lands only at the next plan period, needs finance, HR and legal sign-off, and is barely reversible once issued.",
        "Flags predictable quarter-end discounting as a pricing decision for leadership rather than a rep-coaching item.",
        "Frames the sandbagging as a rational response to the cliff rather than a character flaw.",
        "Recommends measuring forecast accuracy separately from attainment."
      ]
    },
    {
      "id": 7,
      "prompt": "Prepping a QBR slide at Kestrelane and I want a sanity check on four things. One, I want to open with 'only 18% of companies forecast within 5%' - is that the right stat or is there a better one? Two, we sit at 2.1x pipeline coverage and everyone says 3x, so that is my gap slide. Three, our stage probabilities are the CRM defaults, 90/70/50/30/10, and our weighted forecast comes straight out of them. Four, our commit list converted at 100% last quarter, every single committed deal closed, and I want to show that as proof the commit category is healthy.",
      "expected_output": "A refusal of the unsourced statistics, a rejection of 3x as a benchmark in favour of the team's own win rate, replacement of default stage probabilities with measured conversion, and an identification of 100% commit conversion as a sandbagging signal.",
      "files": [],
      "expectations": [
        "Refuses to cite the '18% of companies forecast within 5%' figure and states that it circulates without primary attribution.",
        "Rejects 3x coverage as a benchmark and identifies it as folklore with a disputed origin.",
        "States that the underlying coverage arithmetic is roughly the inverse of the win rate and instructs deriving the target from the team's own win rate.",
        "States that default stage probabilities are tool conventions rather than measured conversion, and must be replaced with the team's own conversion by stage before any weighted forecast is trusted.",
        "Identifies the 100% commit conversion as a sandbagging signal rather than evidence of a healthy commit category.",
        "Directs every threshold to be derived from the organization's own trailing history rather than from published tables.",
        "States that any circulated number that must be mentioned carries an explicit provenance label in the deliverable.",
        "Rejects industry MAPE bands as imported from supply-chain forecasting and not transferable to sales pipelines.",
        "Specifies deriving the staleness threshold from median time-in-stage per stage, flagging deals past roughly twice their stage median.",
        "Specifies deriving the push-count threshold from where the lost-or-slipped population separates from the won population in the team's own push-count distribution.",
        "Recommends dollar-weighted error when deal sizes vary widely, rather than averaging per-deal percentages.",
        "Recommends recomputing derived thresholds every few periods as the team, product and market drift."
      ]
    },
    {
      "id": 8,
      "prompt": "Sifting through Q1 at Oakmere Digital and I want your read on six findings before I write the verdicts up. One, nineteen commit deals have the Economic Buyer field filled in with 'TBD' or 'unknown'. Two, AE Danilo can describe exactly what the buyer last did on his two biggest deals - dates, names, the security review they ran - but none of it is in the CRM. Three, AE Ruth's $290K deal has no buyer-side evidence anywhere and she cannot name one thing the buyer did. Four, field history shows 31 deals jumping from Discovery straight to Negotiation on the same afternoon in February. Five, our Omitted category grew from 4 deals to 61 over the quarter. Six, when I asked six AEs what Commit means I got six different answers. Our field completeness report says we are at 96% on required fields. Which class is each one?",
      "expected_output": "Per-finding class assignments using the data-versus-behavior discriminators, the 96% completeness figure dismissed because placeholders count as complete, the Omitted growth flagged as disguised losses, and the conflicting Commit definitions reported as a methodology finding at the scope boundary.",
      "files": [],
      "expectations": [
        "Classifies the placeholder-filled Economic Buyer fields as a missing or unverifiable evidence problem and states that a filled field is not evidence.",
        "States that a completeness report checking only for non-blank values counts 'unknown' and 'TBD' as complete, so the 96% figure evidences nothing.",
        "Instructs spot-checking field values rather than fill rates.",
        "Classifies Danilo's case as data, a capture burden, because the evidence exists but was never recorded.",
        "Classifies Ruth's case as behavior, because the evidence does not exist anywhere and the deal was never qualified.",
        "Treats the 31 same-afternoon stage jumps as a candidate bulk edit or workflow auto-advance and requires establishing which before any inflation verdict.",
        "Classifies a workflow-driven stage advance as data rather than behavior.",
        "Flags the Omitted category growth as a possible soft delete disguising closed-lost deals and requires auditing its contents.",
        "Treats the six conflicting definitions of Commit as a methodology finding reported at the scope boundary, not as a rep failure.",
        "States that redesigning the forecast category definitions is out of scope for this diagnostic.",
        "States that ambiguous category definitions make honest reps look dishonest, so the definition check precedes any behavior verdict.",
        "Uses the discriminating question - what the buyer did last, not what the rep did - to separate inflation from data.",
        "Requires the two-signal evidence rule to be satisfied before any of these becomes a confirmed cause."
      ]
    },
    {
      "id": 9,
      "prompt": "Different shape of problem. Lumeo is a self-serve subscription product, roughly 40,000 paid accounts at an average $79 a month, no AEs on the core motion. Our monthly revenue forecast is a model: trailing trial-to-paid conversion, cohort retention curves, and a run rate on new trials. On top of that our inside-sales lead Marguerite adds a manual overlay for the assisted-upgrade deals her four SDRs work. We have over-forecast six months running by 8 to 14%, and Marguerite's overlay has come in under her actual every single month. Everyone here tells me this kind of forecast investigation is for enterprise deal teams and does not apply to us. Does it?",
      "expected_output": "A translation of the deal-level method onto the model layer and the human overlay: the diagnostic applies, overlay treated as the human layer, stale assumptions mapped to stale dates, over-inclusive funnel definitions to stage inflation, the conservative overlay to sandbagging, with attribution to named model assumptions.",
      "files": [],
      "expectations": [
        "States that the diagnostic applies here and does not concede that a model-based forecast is out of scope.",
        "Applies the deal-level fingerprints to Marguerite's manual overlay as the human layer of the forecast.",
        "Maps stale conversion assumptions in the model onto the stale-close-date cause.",
        "Maps over-inclusive funnel definitions onto the stage-inflation cause.",
        "Identifies Marguerite's chronically conservative overlay as the sandbagging equivalent.",
        "Recommends tracking overlay-versus-model accuracy the same way manager-override-versus-roll-up accuracy is tracked.",
        "Attributes gap dollars to named model assumptions where no named deal exists, rather than skipping attribution for this motion.",
        "Keeps the same attribution coverage rule: at least 90% attributed, with at most 10% residual labeled unexplained.",
        "Keeps the three-class verdict system of data, behavior and demand unchanged for this motion.",
        "Reports signed error and absolute error separately for the model layer and the overlay layer.",
        "States that six months of consistent direction is bias rather than noise, and keeps bias separate from magnitude.",
        "Does not simply recommend removing the overlay; evaluates whether it beats or underperforms the raw model first."
      ]
    },
    {
      "id": 10,
      "prompt": "Two regional managers at Portwell. Ines submits a roll-up consistently about 12% below the sum of her reps' commits; Anselm submits about 9% above his. Over the last five quarters Ines's submitted number landed within 4% of actual while her reps' raw sum was off by 15%, and Anselm's submitted number was off by 18% while his reps' raw sum was off by 6%. Our new VP wants to ban manager overrides outright and just take the system roll-up. Separately, I have built a set of flag rules out of last quarter's slipped deals and I would like them in production next week. Thoughts?",
      "expected_output": "A rejection of the blanket ban, with Ines's override kept as informed judgment and Anselm's treated as quantified bias on the accuracy comparison, overrides logged and scored per manager, and the new flag rules backtested against an earlier closed period before deployment.",
      "files": [],
      "expectations": [
        "Rejects the blanket override ban rather than endorsing it.",
        "Identifies Ines's override as informed judgment because it beats the raw roll-up, and recommends keeping it with a logged reason.",
        "Identifies Anselm's override as quantified bias because it underperforms the raw roll-up, and treats that divergence as a cause carrying attributable dollars.",
        "States that the discriminating test is override accuracy compared against raw roll-up accuracy over trailing periods, not the size or direction of the divergence.",
        "Requires overrides to be logged with a reason and scored per manager per period.",
        "States that the override log is about an hour of admin to ship but needs roughly a quarter of trailing overrides before the score means anything.",
        "States that managers must accept being scored for the measure to be honest.",
        "Requires backtesting the new flag rules on an earlier closed period before putting them into production.",
        "Sets the backtest pass threshold at flagging at least 80% of the deals that actually slipped or were lost from commit in that period.",
        "States that below that threshold the rules are fitted to the story rather than the data, and must be revised and re-run.",
        "Classifies the manager override as behavior by construction, leaving only judgment-versus-padding open.",
        "Requires a logged reason on the surviving override rather than removing the manager's latitude."
      ]
    }
  ],
  "trigger_queries": [
    { "query": "why did we miss the number last quarter", "should_trigger": true },
    { "query": "our commit is never right", "should_trigger": true },
    { "query": "reps keep sandbagging and I cannot prove it", "should_trigger": true },
    { "query": "deals keep slipping out of the quarter, every quarter", "should_trigger": true },
    { "query": "our forecast accuracy is terrible and I need to know why", "should_trigger": true },
    { "query": "we called $4M and closed $3.1M, what happened", "should_trigger": true },
    { "query": "I do not trust the forecast anymore", "should_trigger": true },
    { "query": "every quarter we come in about 20% under the call", "should_trigger": true },
    { "query": "how do I tell whether my reps have happy ears", "should_trigger": true },
    { "query": "our commit list converts at 40%, is that a rep problem or a data problem", "should_trigger": true },
    { "query": "the VP's roll-up is always higher than the sum of his reps", "should_trigger": true },
    { "query": "we beat the number by 30% and nobody saw it coming", "should_trigger": true },
    { "query": "run a root cause analysis on last quarter's revenue miss", "should_trigger": true },
    { "query": "diagnose our forecast", "should_trigger": true },
    { "query": "all the close dates got pushed the night before the forecast call", "should_trigger": true },
    { "query": "half the deals we committed had no real buyer contact logged", "should_trigger": true },
    { "query": "leadership wants an explanation for why the committed deals did not land", "should_trigger": true },
    { "query": "our quarter-start call and our quarter-end actual never match", "should_trigger": true },
    { "query": "how do I figure out whether we missed because of bad data or genuinely weak demand", "should_trigger": true },
    { "query": "the board asked why we missed and I need a real answer", "should_trigger": true },
    { "query": "we keep finding deals that closed and were never in the forecast", "should_trigger": true },
    { "query": "I think my managers are padding their submitted numbers", "should_trigger": true },
    { "query": "wins keep coming from outside commit", "should_trigger": true },
    { "query": "is 21% error normal for a team our size", "should_trigger": true },
    { "query": "our pipeline looks healthy on paper and we miss anyway", "should_trigger": true },
    { "query": "quarter after quarter the same two reps under-call", "should_trigger": true },
    { "query": "why is our forecast always wrong in the same direction", "should_trigger": true },
    { "query": "what is causing our commit category to be so unreliable", "should_trigger": true },
    { "query": "one rep has beaten commit by 40% three quarters running", "should_trigger": true },
    { "query": "help me build the case for what actually went wrong in Q2", "should_trigger": true },
    { "query": "should I retrain the team or fix the CRM", "should_trigger": true },
    { "query": "we shipped a pipeline cleanup last quarter and the forecast is still wrong", "should_trigger": true },
    { "query": "my CRO wants to know if this was a demand problem or a discipline problem", "should_trigger": true },
    { "query": "we have deals sitting in negotiation with no economic buyer on record, how bad is that", "should_trigger": true },
    { "query": "our forecast is soft mid-quarter and back-loaded every single time", "should_trigger": true },
    { "query": "I need to explain a $900K gap to the board next week", "should_trigger": true },
    { "query": "what tells you a committed deal was never real in the first place", "should_trigger": true },
    { "query": "how do I separate data quality problems from rep behaviour in the forecast", "should_trigger": true },
    { "query": "our self-serve model over-predicts every month and the inside sales overlay is always low", "should_trigger": true },
    { "query": "post-mortem on the revenue miss", "should_trigger": true },
    { "query": "my call was 95% accurate but the absolute error was awful, which do I report", "should_trigger": true },
    { "query": "reps close deals out of the pipeline category that were never in best case", "should_trigger": true },
    { "query": "we have zombie opportunities inflating coverage and the call still missed", "should_trigger": true },
    { "query": "the same deal has been sitting in commit for four quarters", "should_trigger": true },
    { "query": "why does my team beat the call every quarter and still miss the annual plan", "should_trigger": true },
    { "query": "break down last period's forecast versus actual by category", "should_trigger": true },
    { "query": "figure out whether our comp plan is what is causing the forecast bias", "should_trigger": true },
    { "query": "our quarter-end discounts have trained buyers to wait and the mid-quarter call is useless", "should_trigger": true },
    { "query": "I want to stop guessing and attribute the miss to named deals", "should_trigger": true },
    { "query": "what is the right way to score whether my managers' overrides help or hurt", "should_trigger": true },
    { "query": "we missed again, is it the process, the people, or the market", "should_trigger": true },
    { "query": "our deals all close in the last three days of the quarter and we never see them coming", "should_trigger": true },
    { "query": "design our pipeline coverage model, how much pipeline do we need to hit $12M", "should_trigger": false },
    { "query": "rewrite our stage exit criteria so each one names something the buyer did", "should_trigger": false },
    { "query": "our stages are all named after rep activity like demo scheduled, audit them", "should_trigger": false },
    { "query": "run the monthly stale deal sweep before the QBR", "should_trigger": false },
    { "query": "set up a recurring pipeline hygiene checklist with a disposition per flagged deal", "should_trigger": false },
    { "query": "which system should be source of truth for the close date field", "should_trigger": false },
    { "query": "who owns the amount field in our CRM and how often must it be refreshed", "should_trigger": false },
    { "query": "finance and sales report two different ARR numbers, sort out the governance", "should_trigger": false },
    { "query": "design the funnel stage set for our new PLG motion from scratch", "should_trigger": false },
    { "query": "build a bowtie funnel model with post-sale stages", "should_trigger": false },
    { "query": "where are we losing deals between MQL and SQL", "should_trigger": false },
    { "query": "size how much recoverable revenue our leaky funnel is costing us", "should_trigger": false },
    { "query": "write the narrative for the board revenue deck", "should_trigger": false },
    { "query": "how do I present a revenue miss to the board without surprising them", "should_trigger": false },
    { "query": "build our KPI tree from board level down to individual contributor", "should_trigger": false },
    { "query": "help me pick a north star metric for the revenue org", "should_trigger": false },
    { "query": "score our inbound leads on fit and engagement", "should_trigger": false },
    { "query": "our MQL threshold lets too much junk through, recalibrate it", "should_trigger": false },
    { "query": "set up round-robin routing with territory and capacity rules", "should_trigger": false },
    { "query": "leads are sitting unworked in a queue, fix the assignment logic", "should_trigger": false },
    { "query": "design a tiered discount approval matrix with delegation of authority", "should_trigger": false },
    { "query": "our reps discount heavily at quarter end, build the approval tiers to stop it", "should_trigger": false },
    { "query": "build a customer health score with weights, decay and bands", "should_trigger": false },
    { "query": "which product usage signals actually predict churn", "should_trigger": false },
    { "query": "design the closed-won to CSM handoff packet and its SLA", "should_trigger": false },
    { "query": "consolidate our GTM tool stack, we have too many overlapping tools", "should_trigger": false },
    { "query": "we pay for four sales engagement tools, which do we cut", "should_trigger": false },
    { "query": "am I ready to step up to a RevOps manager role", "should_trigger": false },
    { "query": "write the interview scorecard for our first sales ops hire", "should_trigger": false },
    { "query": "which RevOps newsletters and podcasts are worth my time", "should_trigger": false },
    { "query": "which revops skill do I need for this project", "should_trigger": false },
    { "query": "who should I follow in revenue operations on LinkedIn", "should_trigger": false },
    { "query": "generate a weighted forecast with best case, likely and worst case scenarios", "should_trigger": false },
    { "query": "build my commit versus upside breakdown from this pipeline CSV", "should_trigger": false },
    { "query": "run our weekly pipeline review and tell me which deals to push on this week", "should_trigger": false },
    { "query": "assess our gap to quota from the current pipeline", "should_trigger": false },
    { "query": "forecast our cloud infrastructure spend for next fiscal year", "should_trigger": false },
    { "query": "build a demand forecasting model for our inventory replenishment", "should_trigger": false },
    { "query": "how do I calculate MAPE for a demand planning model", "should_trigger": false },
    { "query": "write next year's sales compensation plan", "should_trigger": false },
    { "query": "allocate quotas across our four territories for next fiscal year", "should_trigger": false },
    { "query": "coach my AE on running better discovery calls", "should_trigger": false },
    { "query": "write a MEDDIC qualification checklist for our AEs to fill in", "should_trigger": false },
    { "query": "set up a win/loss interview program with our lost prospects", "should_trigger": false },
    { "query": "why is our win rate dropping quarter over quarter", "should_trigger": false },
    { "query": "build a weekly forecast submission process from scratch, we have none", "should_trigger": false },
    { "query": "define our forecast categories and what qualifies as commit", "should_trigger": false },
    { "query": "should we use commit, best case and pipeline, or a different category set", "should_trigger": false },
    { "query": "decide who calls the final number and what the roll-up cadence should be", "should_trigger": false },
    { "query": "evaluate forecasting tools to replace our spreadsheet", "should_trigger": false },
    { "query": "model ARR growth for our Series B deck", "should_trigger": false },
    { "query": "our sales cycle keeps getting longer, what is driving that", "should_trigger": false }
  ]
}
references/evidence-and-metrics.md
# Error metrics, calibration, and evidence discipline

Read this before quantifying a miss or quoting any number to stakeholders. The forecast-accuracy literature is dominated by vendors selling forecasting software; their mechanics are reliable, their statistics are marketing.

## Choosing the error metric

- Always compute two numbers per level per period: **signed error** (bias - are we consistently high or low?) and **absolute error** (magnitude). A symmetric average alone lets a rep who over-calls and a rep who under-calls cancel to a healthy-looking zero. These two are not competing options to be ranked and chosen between - neither is readable without the other, so both ship every time.
- Fix one denominator before reporting anything. The same quarter - forecast $10M, closed $9.5M - can be presented as "95% accuracy" (actual/forecast) or a "5.3% error rate" (miss/actual): both are defensible alone, but mixing them across periods or teams makes trends meaningless. Pick one definition, write it in the report, and keep it.
- When deal sizes vary widely, weight the error by dollars (a WMAPE-style sum of absolute errors over sum of actuals) rather than averaging per-deal percentages - small deals otherwise dominate the number while the money is elsewhere.
- Benchmark against the team's own trailing periods, not published tables. Industry MAPE bands come from supply-chain forecasting and do not transfer to sales pipelines.

## Rep and manager calibration scorecard

Per person, trailing 4+ periods, compared to their own history first and the team second. Compute every row rather than picking among them: each exposes a different bias and they are all one query against the same history, so there is nothing to rank.

| Metric                                                                   | What it exposes                                                                          |
| ------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------- |
| Commit conversion rate (won-as-called / committed)                       | Inflation when chronically low; sandbagging when chronically near-perfect with big beats |
| Share of wins from outside the forecast                                  | Hidden pipeline - the sandbagging signature                                              |
| Average close-date push count on their deals                             | Date discipline                                                                          |
| Category-change latency (evidence event to category update)              | Whether the record tracks reality or the calendar                                        |
| Signed forecast error, per period                                        | Direction and consistency of personal bias                                               |
| (Managers) override delta vs. rep sum, and override accuracy vs. raw sum | Whether the override adds judgment or bias                                               |

One outlier period is noise. Industry lore - a Forrester line that circulates only as a vendor-blog quotation, with no primary attached - puts the sandbagger bar at the same person doing it more than twice. Treat that as directional support for "require a repeated pattern," not as a citable standard.

## Which circulating statistics are safe to repeat

Almost none as benchmarks. Status of the ones most likely to come up:

| Claim                                                                                                                                         | Status                                                                                                                                                                                                                         |
| --------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| "93% of sales leaders can't forecast within 5%" / "only 18% of firms within 5%"                                                               | Vendor-repeated from an unnamed 2021 survey, with no primary attribution. Do not cite.                                                                                                                                         |
| MEDDPICC teams forecast within 10%, +18% win rates, faster cycles                                                                             | Vendor-blog cluster, no primary study anywhere in the chain. Do not cite.                                                                                                                                                      |
| "Happy ears causes 40-60% slip rates"                                                                                                         | Vendor glossary, no methodology. Do not cite.                                                                                                                                                                                  |
| 3x pipeline coverage (SMB 2.5-3x, enterprise 4-5x variants)                                                                                   | Folklore with a disputed origin story; the granular variants trace to an aggregator citing an unnamed dataset. The underlying arithmetic - coverage ≈ 1 / win rate - is sound; derive the target from the team's own win rate. |
| Default stage probabilities (80/60/40/20...)                                                                                                  | Tool conventions, not measurements. Replace with the team's own conversion by stage before trusting any weighted forecast.                                                                                                     |
| Forecast-category conversion odds (pipeline ~25%, best case ~a third to half)                                                                 | Practitioner rules of thumb repeated by consultants, not published by any platform. Directional at best.                                                                                                                       |
| Vendor slip-rate findings (most teams see >10% of committed deals slip; committed-deal conversion ranging roughly 60-80% by performance tier) | One vendor's own analysis of its customer base, methodology and sample undisclosed. Cite only with that caveat, never as an industry benchmark.                                                                                |
| "Nobody has a 100% conversion rate on committed deals"                                                                                        | Qualitative, safe, and useful: a commit list converting at 100% is itself a sandbagging signal, not excellence.                                                                                                                |

Standing rule: derive every threshold from the organization's own history; when a circulated number must be mentioned, label its provenance in the deliverable rather than dropping it into a slide as fact.

## Frameworks usable as evidence checklists

Three inspection practices, competing for the same review hours. Recommend one; the reader cannot run all three at once.

- value (most first): `qualification-framework scoring > early-quarter champion verification > buyer-side progress question`
- effort (most first): `qualification-framework scoring > early-quarter champion verification > buyer-side progress question`
- efficiency (best first): `buyer-side progress question > early-quarter champion verification > qualification-framework scoring`

Start at the buyer-side progress question: near-zero effort, one sentence added to a review that already happens, and it unwinds most stage inflation on its own. Move up to early-quarter champion verification when the miss came from deals that were never real rather than deals that slipped - it costs a leader an hour per rep at period start and buys the whole period to act.

The full framework rollout is starved by that order and stays starved: highest value, a quarter to embed, cross-team. It is the right call only when reps genuinely disagree about what qualified means, which is a methodology finding at this skill's scope boundary, not a fix it ships.

- **Buyer-side progress as the unit of evidence**: recurring practitioner principle - real progress is something the buyer did, not something the seller did. It is the single question that unwinds most stage inflation, and it costs one sentence in a review that already happens.
- **Early-quarter champion verification** (associated with John McMahon's writing on sales leadership, via published summaries): have leaders directly verify the champions of forecasted deals in the first weeks of the period, and push unqualified deals off the forecast early rather than letting them sit uninspected until they miss. A practical inspection habit; adopt the practice without attaching numbers to it. An hour per rep at period start, and it buys the rest of the period to act.
- **MEDDICC / MEDDPICC** (or whatever qualification framework the team already runs): legitimate, widely used structures for scoring deal evidence element by element. The diagnostic value is the evidence standard, not the acronym - score against buyer-confirmed facts, and remember a filled CRM field is not evidence: "unknown" counts as complete in any system that only checks non-blank. The accuracy statistics vendors attach to these frameworks are the unsourced claims flagged above, and it takes a quarter to embed across a team, which is why the order above puts it last despite ranking first on value.
references/root-cause-fingerprints.md
# Root-cause fingerprints

One block per cause: the signals to pull, the fingerprint pattern, a confirming second signal, the data-vs-behavior discriminator, the fix, and what the fix costs. The two-signal evidence rule applies to every block - one signal raises a suspicion, two confirm a cause.

Blocks run in the register order from SKILL.md - fix effort ascending, which is the efficiency order only while the attributed dollars are unknown. Detection order is not delivery order: run every fingerprint, then rank every confirmed cause by dollars divided by the fix effort tier recorded here, and show the adjustment. There is exactly one ordering across this skill; this file does not carry a second one.

## Zombie / duplicate opportunities

- Pull: last-activity date and age per open deal; duplicate candidates by account plus similar amount/contact.
- Fingerprint: open forecast-weighted deals with no activity beyond the team's derived staleness threshold; the same buying intent recorded as two opportunities and counted twice.
- Second signal: reps privately concede the deal is dead ("keeping it open just in case"); win rate looks artificially low because the denominator is padded with corpses.
- Discriminator: almost always data - the record outlived reality. It becomes behavior only when zombies are kept deliberately to fake coverage.
- Fix: close or archive with reason codes; then hand recurring prevention to the sales-pipeline-hygiene audit - a one-time purge without an entry-standard fix regrows.
- Fix effort: an hour of bulk admin, then a standing sweep. Nobody's judgment is constrained and the purge is reversible, so it needs no negotiation; coverage stops being fiction in the same cycle.

## Stale / mass-pushed close dates

- Pull: close-date field history per deal - number of pushes, edit timestamps, edit sizes.
- Fingerprint: push counts accumulating without a logged reason; edits clustering on a single day near period end - the bulk-update fingerprint of a rep or manager sweeping dates the night before the forecast call.
- Second signal: close dates piling up on the last day of the period; deals whose date has been pushed more times than the team's derived threshold (below).
- Discriminator: a push with a buyer-side reason logged is a forecast update; repeated silent pushes are a data-integrity failure; mass same-day pushes are process theater.
- Fix: no silent pushes - every date move carries a reason; repeated pushes trigger a category review, not just a new date.
- Fix effort: an hour of CRM admin for the reason requirement, then a standing review of what the reasons say. It only bites at the next push, so allow one full cycle before the dates mean anything.

## Wrong forecast-category assignment

- Pull: category field history; category vs. stage vs. evidence cross-tab; contents of the omitted/excluded category.
- Fingerprint: category contradicts stage and evidence (commit with no activity for weeks); the omitted/excluded category used as a soft delete to dodge a formal closed-lost; categories set at creation and never updated.
- Second signal: conversion by category diverges wildly from the team's own trailing baseline per category.
- Discriminator: if reps articulate different definitions of "commit", the definition is ambiguous - a methodology finding to report at the boundary, not a rep failing. If definitions are shared and the assignment still contradicts evidence, it is behavior.
- Fix: category changes logged with reasons; omitted entries audited for disguised losses; per-category conversion tracked as the honesty check.
- Fix effort: an hour of admin to log changes, plus one conversation with the people whose calls the logging exposes. Jumps to a week when the ambiguity finding is real and the definitions must be re-agreed - and at that point it is a methodology question at the scope boundary, not a fix this skill ships.

## Stage inflation / happy ears

- Pull: field history of stage changes; activity log; contact roles; win rate by stage from trailing history.
- Fingerprint: late-stage deals with no buyer-side evidence - no economic-buyer contact logged, no next meeting booked, single-threaded; or stage skips (discovery straight to negotiation) in field history.
- Second signal: actual conversion from that stage runs far below whatever probability the stage implies; deals carrying a close date within days while sitting in an early stage.
- Discriminator: ask the rep what the buyer did last, not what the rep did. A confident seller-side answer with no buyer-side action is inflation; "the field was auto-advanced by a workflow" is data.
- Fix: buyer-verifiable evidence required to enter commit-eligible stages; review questions shift from "when will it close" to "what did the buyer do this week". If the stage definitions themselves have no verifiable exit criteria, hand off to the stage-definition audit.
- Fix effort: a week to define the gate and rewrite the review questions, sales leadership's agreement to enforce it, then a standing inspection. The forecast call gets honest one cycle later; the win-rate evidence takes two.

## Missing / unverifiable evidence

- Pull: required-field completeness on forecast-category deals; the field values themselves, not just non-blank counts.
- Fingerprint: fields filled with placeholders - "unknown", "TBD", the rep's own name as champion. A filled field is not evidence; systems that only check non-blank count "unknown" as complete.
- Second signal: nobody in the deal review can state the buyer's last verifiable action; qualification-framework scores (MEDDICC/MEDDPICC or whatever the team uses) all green with no artifacts behind them.
- Discriminator: if the rep can supply the evidence verbally but never recorded it, it is data (capture burden); if the evidence does not exist anywhere, it is behavior (the deal was never qualified).
- Fix: evidence-based entry criteria for commit categories; spot-check field values, not fill rates.
- Fix effort: the same week, the same gate, and the same owner as stage inflation - one intervention answers both, so never bill them as two efforts or ship one without the other.

## Late-created deals

- Pull: created date vs. close date per won deal; creation stage.
- Fingerprint: deals created and closed within days, or created directly in a late stage. The revenue is real; the forecast never saw it coming - coverage math and early-quarter calls were fiction.
- Second signal: a large share of the period's won revenue carries a creation date inside the period's final weeks.
- Discriminator: data/process - reps work deals outside the system and book them at signature. Not dishonesty; invisibility.
- Fix: log opportunities at first qualified conversation; measure created-to-closed interval per rep and set a floor consistent with the real sales cycle.
- Fix effort: a week to set the rule and the measurement, but the change itself is a rep habit - expect a cycle or two of partial compliance before coverage math can be trusted again.

## Manager roll-up override

- Pull: submitted number at each roll-up level vs. the sum of the level below, across periods; override reasons if logged.
- Fingerprint: the manager number diverges from the rep sum in a consistent direction with no logged reason.
- Second signal: compare override accuracy vs. raw roll-up accuracy over trailing periods. A manager whose overrides beat the raw sum is adding judgment - keep it, require a logged reason. A manager whose overrides underperform the sum is adding bias - that divergence is itself a quantified cause.
- Discriminator: the override is behavior by construction; the question is only whether it is informed judgment or padding/haircutting habit. The accuracy comparison answers it with data.
- Fix: overrides stay allowed but logged with a reason and scored per manager per period.
- Fix effort: an hour to capture the reason, but the score is meaningless until a quarter of overrides has accumulated, and managers have to accept being scored before any of it is honest. Cheap to ship, slow to pay.

## Rep sandbagging

- Pull: rep-level forecast submissions vs. actuals across 4+ periods; origin category of every won deal.
- Fingerprint: the same rep beats commit by a wide margin repeatedly, with wins closing from deals never placed above the lowest category despite late-stage activity visible in history.
- Second signal: large deals surfacing into commit only in the final days of the period, already fully evidenced; quarter-start commit consistently a fixed fraction of quarter-end actual.
- Discriminator: one strong quarter is noise. Require the pattern across at least two consecutive periods for the same person before naming it - and check the comp plan first, because sandbagging is usually a rational response to cliffs, not a character flaw.
- Fix: score rep calibration openly against their own history - per-rep forecast-vs-actual across the last 4+ periods, published to the team; adjust incentives that reward hiding upside; never punish accuracy misses in comp.
- Fix effort: a week to build the scorecard, and 2+ periods of history before the cause can even be named. Naming an individual is the one verdict here with employment consequences, so it carries a review the other fixes do not. Anything touching the incentive itself is the comp-driven bias fix, at that fix's cost - do not price it as part of this one.

## Comp-driven bias

- Pull: comp plan mechanics (cliffs, accelerators, quarter-end discount authority); timing distribution of closes and slips around period boundaries.
- Fingerprint: slips clustering just past a quota cliff once the number is made; buyers trained by predictable quarter-end discounts to wait, making every quarter back-loaded and every mid-quarter forecast soft.
- Second signal: the same reps' behavior flips with their attainment position - sandbagging when ahead, inflating when behind.
- Discriminator: this is the root behind most behavior-class causes. Fixing surface behavior while the incentive stays is temporary.
- Fix: report the incentive-to-bias link explicitly; propose measuring forecast accuracy separately from attainment; flag discount-timing predictability as a pricing decision for leadership.
- Fix effort: a quarter at best, since plan amendments usually land only at the next plan period, need finance, HR and legal sign-off, and are barely reversible once issued to reps. This is the cause the efficiency order starves - biggest impact, worst ratio, and it stays unfixed forever if the register is read by ratio alone. Promote it anyway on the conditions named in SKILL.md; when it is ruled out, say which behavior causes will regrow rather than quietly listing it last.

## Deriving thresholds from the team's own history

No universal day-count or push-count exists - published ones are tool defaults or folklore. Derive each threshold from the team's own trailing data (4+ periods where available). These are not ranked and should not be: each one serves exactly one fingerprint, so deriving a threshold for a fingerprint that never flagged is wasted work regardless of its cost - derive only what the flagged causes need.

1. Staleness: compute median time-in-stage per stage; flag deals beyond 2x their stage median.
2. Push count: compute the push-count distribution for won vs. lost/slipped deals; set the flag where the lost/slipped population separates from the won one (often the top decile).
3. Category conversion baselines: trailing conversion per category per segment; a period deviating far from its own baseline is the anomaly to explain.
4. Late-creation floor: compare created-to-closed intervals against the measured sales cycle; flag intervals implausibly short for the motion.
5. Recompute every few periods - thresholds drift as the team, product, and market change.
references/worked-diagnosis-example.md
# Report shape and worked examples

## Report shape

Present in this order, section by section, each gated on user approval:

1. **The miss, quantified** - forecast vs. actual per level and category, signed bias and absolute error, the error definition used.
2. **Attribution table** - every gap dollar to a named deal (or model assumption), one primary cause, one class; residual labeled unexplained.
3. **Verdicts** - causes in impact order (attributed dollars), each with its two supporting signals and its class. This section is the arithmetic of the miss and stays in dollar order.
4. **Demand isolation** - the hindsight-corrected forecast and the residual genuine shortfall, if any.
5. **Fixes** - in leverage order (dollars divided by fix effort), one per cause, each with class, owner, effort tier, and the metric that proves it worked within one period. Print the impact order beside it and name every cause that moved between the two, with the effort that moved it.
6. **Threshold check** - attribution coverage, evidence rule, backtest result.

## Worked example (positive)

B2B SaaS, 40 reps, quarterly forecast. Quarter-start commit $4.0M; actual $3.15M; gap $850K (21% signed miss at company level). Snapshots and field history available for 6 quarters.

Attribution table (all figures from the period-start snapshot reconstruction):

| Gap component                        | Deals | $               | Primary cause                   | Class    | Signals                                                                                                                        |
| ------------------------------------ | ----- | --------------- | ------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------ |
| Commit deals slipped to next quarter | 5     | $370K           | Stale / mass-pushed close dates | Data     | 3+ pushes each, no reasons; all 5 edited the same evening before the month-end forecast call                                   |
| Commit deals lost                    | 2     | $260K           | Stage inflation                 | Behavior | Both single-threaded with no economic-buyer contact ever logged; both skipped from discovery to negotiation in field history   |
| Commit deals dead at quarter start   | 3     | $140K           | Zombie opportunities            | Data     | No activity for 70+ days at snapshot date; reps confirmed in interview they knew the deals were dead                           |
| Wins from outside the forecast       | 4     | +$120K (offset) | Rep sandbagging                 | Behavior | One rep, third consecutive quarter beating commit by >30%; all 4 deals showed late-stage activity weeks before entering commit |
| Unexplained residual                 | -     | $80K            | -                               | -        | 9.4% of gap, within the 10% cap, labeled as such                                                                               |

Note the offset is reported, not netted: the true over-forecast was $970K, partially masked by $120K of hidden upside. Netting the two would have understated both problems.

Demand isolation: rebuilding the quarter-start forecast with zombies removed, the two unevidenced deals downgraded, and evidence-backed dates gives a corrected commit of $3.28M vs. $3.15M actual - a $130K residual attributable to genuinely thin coverage in one segment. That part goes to pipeline generation, not to forecast-process fixes.

Backtest: the same fingerprints run on the prior quarter flagged 9 of the 11 deals that actually slipped or lost from commit (82% - passes).

### The ranking adjustment, shown

Impact order, straight from the table: stale dates $370K > stage inflation $260K > zombies $140K > sandbagging $120K.

Leverage order, the same causes divided by their fix effort tier: stale dates ($370K, an hour) > zombies ($140K, an hour) > stage inflation ($260K, a week plus a cycle) > sandbagging ($120K, a week plus two periods of history).

What moved and why:

- **Zombies up, #3 to #2.** Same admin hour as the leader for a third of the dollars, and the purge makes coverage honest in the current cycle rather than the next one.
- **Stage inflation down, #2 to #3.** The largest behavioral number in the diagnosis, but the evidence gate takes a week to define, needs sales leadership to enforce it, and does not move commit conversion until the following period.
- **Nothing moved into first place.** The leader is the largest attributed cause that also happens to be an hour of admin - leverage promotes cheap fixes, it does not promote cheapness. A $20K zombie purge would still rank below a $370K one at identical effort.
- **Comp-driven bias is absent, and that is a finding.** Nothing in this period's fingerprints pointed at a quota cliff, so it was never confirmed. Had it been, it would have ranked last on leverage and first in the narrative, with the behavior fixes flagged as temporary until the plan changed.

Fixes shipped, in leverage order:

| Fix                                       | Class    | Owner            | Effort                              | Metric                      |
| ----------------------------------------- | -------- | ---------------- | ----------------------------------- | --------------------------- |
| Reason-required close-date pushes         | Data     | RevOps           | An hour                             | Silent-push count           |
| Zombie purge plus recurring hygiene audit | Data     | RevOps           | An hour, then a standing job        | Stale-deal count            |
| Buyer-evidence gate on commit entry       | Behavior | Sales leadership | A week, then a standing inspection  | Commit conversion rate      |
| Open rep-calibration scorecard            | Behavior | Sales leadership | A week, plus two periods of history | Rep commit beat/miss spread |

## Negative example (a wrong diagnosis, and why)

Same company, first pass - before this method was applied. Leadership concluded: "reps have happy ears - mandate qualification retraining."

Evidence offered: (a) the team's win rate on forecast-weighted deals was 19%, far below the ~60% the late stages implied; (b) the VP recalled one vivid deal a rep had sworn was closing that never did.

Why the diagnosis was wrong:

- **Single-signal reasoning.** The win-rate gap was one signal, and the anecdote was not a signal at all. No field history was pulled; the two-signal rule would have blocked the verdict.
- **Hindsight data.** The 19% was computed on the end-of-quarter deal list, whose denominator was padded with zombie deals nobody had closed out - a data problem inflating the apparent optimism. On the period-start snapshot with zombies excluded, evidenced late-stage deals actually converted at 54%.
- **Wrong class, wrong fix.** Retraining (a behavior fix) was applied to what was mostly a data problem. Stale dates and zombies were untouched, so the next quarter missed again.
- **Second-order damage.** Because the retraining arrived framed as blame, reps protected themselves the following quarter by under-calling - the error flipped from over-forecast to under-forecast, and leadership concluded the training "worked" until the annual plan, built on sandbagged numbers, came in short.

The corrected diagnosis (the positive example above) attributed only $260K of a $970K gross miss to actual inflation - the retraining had targeted roughly a quarter of the problem with a fix that made the rest worse.
SKILL.md
---
name: sales-forecast-diagnostic
description: Diagnose why an existing sales forecast is unreliable and recommend fixes - stage inflation and happy ears, sandbagging, deals without verifiable buyer evidence, stale or mass-pushed close dates, wrong forecast-category assignment, zombie and duplicate opportunities, late-created deals, manager roll-up overrides, and comp incentives that reward bias - separating data-quality from behavioral from genuine demand problems. Use whenever the user mentions forecast accuracy, a forecast miss, reps sandbagging, deals that keep slipping, "why did we miss the number", or "our commit is never right" - even if they never say "forecast". Covers B2B deal-based and B2C/high-volume. Do NOT use for pipeline coverage modeling - use mbfinotti/sales-skills@sales-pipeline-coverage-modeling instead.
license: MIT
metadata:
  author: Maya-Beth Finotti
  version: "1.3.4"
---

# Forecast Diagnostic

Investigate a forecast that missed, or that nobody trusts, and name the root cause with evidence. A forecast call asks "what's the number?" This diagnostic asks "was the number ever real?"

Practitioners converge on the view that most misses are inspection failures, not market failures: the evidence that a committed deal was fiction usually existed in the record before the miss.

Scope boundary, stated up front: this skill consumes whatever forecasting methodology already exists - the category set, roll-up rules, commit criteria, and coverage targets are inputs, not deliverables. Designing or redesigning that methodology is out of scope. When the investigation shows the methodology itself is undefined or ambiguous ("commit means different things to different reps"), report that as a finding and stop at the boundary.

Every confirmed root cause lands in exactly one of three classes, because each class has a different owner and a different fix:

- **Data**: the record was wrong, a human knew better.
- **Behavior**: the record faithfully reflects a biased human call.
- **Demand**: honest, evidenced pipeline still fell short.

The most common diagnostic error is issuing a behavior verdict for a data problem.

## Interview

Ask before analyzing:

- One question per message; offer multiple-choice options when possible.
- Skip anything already answered.
- Ask the first three questions below before any of the rest. Their fixes diverge by two orders of magnitude in effort and in how long they take to reach a forecast, and the answers decide the order the whole deliverable is ranked in.

Questions:

- By when must the number be more reliable - this period's call, the next one, or the next planning cycle? A fix that lands after the period closes does not help this forecast: a mid-period deadline deletes every fix slower than one cycle from the plan and leaves the hour-tier CRM-admin fixes carrying the whole load.
- One-off correction of this miss, or a standing accuracy program? One-off promotes the reconstruction, the zombie purge, and the date cleanup; standing promotes evidence gates, calibration scorecards, and threshold derivation.
- Effort ceiling: how many analyst hours, whether a CRM admin can ship a field or validation change this week, and who has authority to reopen the comp plan or change a review cadence. No admin turns every hour-tier fix into a week-tier request; no comp authority deletes the comp fix outright.
- What triggered this - a specific missed period, chronic inaccuracy, or a suspicion (sandbagging, inflated pipeline)? Which period(s)?
- Which direction is the error - over-forecast (missed the call), under-forecast (beat it by a lot), or erratic both ways?
- At which level does the number go wrong - individual rep, manager roll-up, or company-wide? Where in the roll-up can a human override the sum?
- What does the process look like today - forecast categories in use, submission cadence, who calls the final number? (Recorded as-is; not redesigned.)
- What history exists - opportunity snapshots, field history (stage, close date, category, amount changes), activity logs? How far back?
- What is the motion - B2B deal-based, B2C/high-volume/self-serve model-based, or mixed? Roughly how many deals per period?
- How is the team paid - quota cliffs, accelerators, quarter-end discount authority, anything rewarded or punished based on the forecast number itself?
- Any confounders in the period - CRM migration, stage redefinition, territory change, a bulk hygiene cleanup?

## Workflow

1. Run the Interview; establish the period, direction, and level of the error before touching deal data.
2. Quantify the miss before explaining it: forecast vs. actual per period, per roll-up level, per forecast category. Compute both signed error (bias) and absolute error - a symmetric average hides offsetting errors, since the same miss reads very differently depending on the chosen denominator, so fix one error definition first. Both numbers always ship together, as a correctness rule rather than a menu to rank and pick from - see [references/evidence-and-metrics.md](references/evidence-and-metrics.md).
3. Inventory the evidence base: period-start snapshot, field history, activity logs. If no snapshots or field history exist, start capturing them now, diagnose from current-state fingerprints plus interviews, and label every conclusion lower-confidence - without history, behavior verdicts are hard to defend.
4. Reconstruct the missed period from the period-start snapshot. Classify every deal that was in a forecast category by outcome (won as called, slipped, lost, removed, still open) and every deal actually won by origin (in the forecast? which category? when created?). If your harness can execute code, script this from a two-file export (period-start forecast, period-end actuals); otherwise ask the user for both lists and reconcile them by hand.
5. Run the root-cause fingerprints in [references/root-cause-fingerprints.md](references/root-cause-fingerprints.md) over the reconstruction. Require two independent signals per suspected cause, at least one from recorded history.
6. Interview reps and managers on flagged deals to split data from behavior. The discriminating question: "did you know this deal was dead (or real) before the CRM said so?" Yes means the record lagged reality - a data problem. No means behavior or demand.
7. Isolate demand last: rebuild the period-start forecast with hindsight-corrected data - zombies removed, honest categories, evidence-backed close dates. Whatever gap survives the correction is a genuine demand or coverage shortfall that no forecasting-process fix will recover; say so plainly.
8. Attribute the gap: every dollar of miss maps to a named deal (or a named model assumption in high-volume motions), each with one primary root cause and one class. Offsetting errors (sandbagged wins masking inflated losses) are attributed separately, never netted silently.
9. Check the pass thresholds in Measurement below; tune the fingerprints and re-run until every threshold holds.
10. Rank the confirmed causes for delivery using Fix leverage below: attributed dollars first, then divided by fix effort, with the adjustment shown rather than the two orders silently merged. One fix per cause matched to its class, plus the metric that will prove each fix worked within one period.
11. Deliver section by section for user approval; report shape and worked examples in [references/worked-diagnosis-example.md](references/worked-diagnosis-example.md).
12. If your harness has persistent memory, memorize the confirmed causes, the thresholds derived from this team's history, the effort tiers as they turned out for this team, and rep-level calibration baselines - the next period's re-run starts from them.

## Root causes in scope

The compact map: full detection fingerprints, second signals, and fixes live in [references/root-cause-fingerprints.md](references/root-cause-fingerprints.md), in this same order.

The value axis cannot be pre-ranked here and this skill does not pretend to: which cause is worth the most dollars is the output of the attribution table, and it differs with every engagement. What is stable is the fix - what it costs to ship and how long before a forecast reflects it - so the register is ordered by that, and the diagnosis then re-ranks it against the dollars it actually found (see Fix leverage).

- fix effort (most first): `comp-driven bias > rep sandbagging > stage inflation == missing evidence > late-created deals > manager override == wrong category > stale dates == zombies`
- time-to-effect (slowest first): `comp-driven bias > rep sandbagging > manager override > late-created deals > stage inflation == missing evidence > stale dates == wrong category > zombies`
- compliance cost (most first): only two carry any - `comp-plan change > a behavior verdict naming an individual rep`; every other fix is a CRM setting or a review habit with no external sign-off, so the axis is not worth a full ordering.
- efficiency at equal dollars (best first): `zombies > stale dates == wrong category > stage inflation == missing evidence > late-created deals > manager override > rep sandbagging > comp-driven bias`

Ties, each earned rather than dodged:

- **Stage inflation == missing evidence**: both fixes are literally one intervention, a buyer-evidence gate on commit entry, same owner, same week, same standing inspection.
- **Manager override == wrong category**: both are an hour of CRM admin (a reason field, category field history) plus exactly one negotiation with the people whose judgment the change constrains.
- **Stale dates == zombies**: both are one admin change plus a recurring sweep, constrain nobody's judgment, and are reversible the day they ship.
- **Stale dates == wrong category**, on time-to-effect only: both bite at the next edit, so one full cycle either way.

| Root cause                       | Usual class                   | Signature                                                                                                                               | Fix effort                                                                                                                                                             |
| -------------------------------- | ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Zombie / duplicate opportunities | Data                          | Long-dead deals still open and counted; the same buying intent counted twice                                                            | An hour to purge with reason codes, then a standing hygiene sweep; coverage is honest in the same cycle                                                                |
| Stale / mass-pushed close dates  | Data                          | Push counts pile up; date edits cluster on a single day near period end                                                                 | An hour of admin for reason-required date moves, then a standing job; bites at the next push, so one cycle                                                             |
| Wrong category assignment        | Data or behavior              | Category contradicts stage and evidence; an omitted/excluded category used as a soft delete; category set at creation and never touched | An hour of admin for logged category changes - a week if the definitions themselves must be re-agreed across the team                                                  |
| Stage inflation / happy ears     | Behavior                      | Late-stage deals with no buyer-side evidence; stage skips in field history; wins far below what the stage implies                       | A week to define the buyer-evidence gate and change the review questions, then a standing inspection; the call gets honest one cycle later, win rates two              |
| Missing / unverifiable evidence  | Data or behavior              | Required fields filled with placeholders; nobody can say what the buyer last did                                                        | Same week, same gate, same owner as stage inflation - shipping one ships the other                                                                                     |
| Late-created deals               | Data                          | Deals created and closed within days, or created directly in a late stage - pipeline visibility is bookkeeping after the fact           | A week to set the logging rule, then a rep-habit change; coverage math is untrustworthy for a cycle or two after                                                       |
| Manager roll-up override         | Behavior                      | Manager number diverges from the rep sum in a consistent direction with no logged reason                                                | An hour to capture reasons, but a quarter of trailing overrides before the per-manager score means anything, and managers must accept being scored                     |
| Rep sandbagging                  | Behavior                      | Rep beats commit by a wide margin repeatedly; wins close from outside the forecast                                                      | A week to build the open calibration scorecard, 2+ periods of history before the pattern can be named at all; anything touching the incentive belongs to the row below |
| Comp-driven bias                 | Behavior (root of the others) | Slips cluster just past quota cliffs; predictable quarter-end discounts teach buyers to wait                                            | A quarter at best - plan amendments land at the next plan period, need finance/HR/legal sign-off, and are barely reversible once issued                                |

The efficiency order starves comp-driven bias, the one cause that drives the others: top of the impact axis in most diagnoses, bottom of the ratio in every one. A register read by ratio alone recommends the same CRM hygiene every period while the incentive that produces the behavior stays untouched, and the behavioral causes regrow.

Promote comp-driven bias above everything else when:

- Fingerprints show slips clustering at a quota cliff.
- The same rep's bias flips with their attainment position.
- A previous cycle's data fixes shipped and the behavior class did not improve.

At that point, the cheap fixes have no remaining value, and the deliverable says so instead of relisting them.

Boundary with neighbors: this skill detects deals sitting in a stage without evidence, but it does not redesign the stage set (see `mbfinotti/revops-skills@pipeline-stage-definition-audit` when the definitions themselves are the problem). It is a root-cause investigation of a specific miss, not the recurring stale-deal/missing-field sweep (see `mbfinotti/revops-skills@sales-pipeline-hygiene` for prevention once causes are fixed).

## Fix leverage

The register above is the tiebreak, not the answer. Once the attribution table exists, rank the confirmed causes for delivery in three passes, and put all three in the deliverable:

1. **Impact order** - causes sorted by attributed gap dollars, straight off the attribution table. This is the arithmetic of the miss and it never gets adjusted away; it is what makes the size of each problem auditable.
2. **Leverage order** - the same causes divided by their fix effort tier from the register. Never divide by a currency amount: the denominator is analyst hours, admin or engineering work, cross-team negotiation, and the number of forecast cycles before the fix reaches a call.
3. **The adjustment, shown** - every cause that moved between the two orders, with one line naming the effort that moved it. Presenting only the leverage order looks like the impact arithmetic was silently overruled, and the first executive question is always why the biggest number is not first.

Deliver in leverage order, with the impact order printed beside it. Lead with the best ratio, which is frequently not the cheapest fix - a large attributed cause with an hour-tier fix outranks a small one with the same hour-tier fix, and outranks everything week-tier. A worked adjustment is in [references/worked-diagnosis-example.md](references/worked-diagnosis-example.md).

## Re-rank before delivering

Every ordering here is a default, not a law. It shifts with the team and with who executes it, so re-rank against what the Interview and the reconstruction revealed:

- Deadline inside the current period - delete every week-tier and quarter-tier fix from this cycle's plan and say which, rather than listing them at the bottom where they come back as scope. What survives is `zombies > stale dates > wrong category`, and the deliverable states plainly that the behavior causes are untouched.
- A CRM admin on hand with authority to ship a validation rule this week - the hour-tier fixes stay an hour and lead. Without one, they become a week-tier request in someone else's queue and lose their ratio advantage to the evidence gate, which at least ships in the same week.
- A sales leader who will not reopen the comp plan - delete the comp fix from the plan, and name the specific behavior causes that will regrow because of it. Do not park it at the bottom as a demoted option.
- No snapshots or field history - the manager-override and sandbagging causes cannot be evidenced this pass at all. Delete them from the fix plan, start capturing history now, and schedule them for the re-run.
- Mid-quarter, forecast cycle already running - a fix shipped now changes the next call, not this one. Say which of the two periods each fix is aimed at, and stop offering the ones that reach neither.
- Anything the team already owns that the default order assumes away - a working evidence gate, a scored override log, a hygiene automation - drops that cause out of the register instead of being re-recommended.

## Data vs. behavior vs. demand

- **Data**: someone knew the truth; the record did not. Fix with hygiene automation and field ownership rules (`mbfinotti/revops-skills@crm-data-governance`), not with rep coaching.
- **Behavior**: the record faithfully reflects a biased judgment - optimism, self-protection, or a rational response to incentives. Fix with evidence standards for category entry, incentive changes, and review discipline. A dashboard never fixes behavior.
- **Demand**: honest, evidenced pipeline still fell short - coverage was genuinely thin or win rates genuinely moved. The fix belongs to pipeline generation and market strategy, not the forecast process. The coverage rule of thumb (target ≈ inverse of win rate, hence the popular "3x") is a widely repeated heuristic with no traceable primary study - derive the real target from the team's own conversion history.

Verdict discipline: a behavior verdict is an accusation. Make one only with the two-signal evidence rule satisfied plus the rep interview done, and keep the accuracy conversation separate from the performance/comp conversation - a diagnostic that feeds punishment teaches reps to sandbag, and the next miss becomes the diagnostic's own doing.

## B2B vs. B2C / high-volume / self-serve

- **B2B deal-based**: the workflow above applies literally - deal-by-deal reconstruction, deal-level fingerprints.
- **B2C / high-volume / self-serve**: the forecast is usually a model (run rate, cohort conversion, funnel math) with a human overlay. Apply deal-level fingerprints to whatever human layer exists (an inside-sales commit list, a manual adjustment). Translate the rest to the model layer:
  - Stale conversion assumptions are the stale close dates.
  - Over-inclusive funnel definitions are the stage inflation.
  - A chronically conservative manual overlay is the sandbagging.

  Track overlay-vs-model accuracy exactly like manager-override-vs-roll-up accuracy.

- **Identical for both motions**: the miss-decomposition method, the three-class verdict system, the measurement thresholds, and the statistics discipline do not change.

## Failure modes

| Trap                                            | Why it is wrong                                                                                               | Fix                                                                                                                                                               |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Diagnosing sandbagging from one strong quarter  | One beat is noise; sandbagging is a repeated pattern by the same person                                       | Require the pattern across 2+ periods against the rep's own history                                                                                               |
| Behavior verdict on a data problem              | Ambiguous category definitions make honest reps look dishonest                                                | Run the discriminating interview question first; check definition consistency across reps                                                                         |
| Judging period-start calls with period-end data | Hindsight bias - the call must be judged on what was knowable then                                            | Reconstruct from the period-start snapshot, never from today's record                                                                                             |
| Anchoring on the biggest slipped deal           | One vivid deal rarely explains the gap; the long tail usually does                                            | Decompose all gap dollars before concluding anything                                                                                                              |
| Symmetric accuracy metric as the goal           | Over- and under-forecast have different causes and costs, and they net out                                    | Always report signed bias and absolute error separately, per level. Not a menu to rank and pick one from - both numbers are required for the other to be readable |
| Ranking the fixes by attributed dollars alone   | The comp-plan redesign outranks the field change that would have landed this cycle, and nothing ships in time | Divide dollars by the fix effort tier and show the adjustment (Fix leverage), then re-rank against the Interview answers                                          |
| Deleting zombies and declaring victory          | The evidence standard that admitted them is untouched; they regrow                                            | Pair every cleanup with the entry-standard fix, and hand recurring prevention to the hygiene audit                                                                |
| Punishing forecast error in comp                | Reps respond by sandbagging; error flips sign instead of shrinking                                            | Measure accuracy openly, separate it from performance reviews                                                                                                     |
| Trusting circulating benchmark statistics       | Most quoted forecast-accuracy numbers have no traceable primary study                                         | Derive every threshold from the team's own history; see the evidence reference                                                                                    |

## Measurement and pass thresholds

The diagnostic must meet every threshold below before it ships; iterate until it does:

- **Attribution coverage**: at least 90% of the gap dollars attributed to named deals (or named model assumptions), each with exactly one primary root cause and one class; the residual, at most 10%, is explicitly labeled unexplained - never silently absorbed.
- **Evidence rule**: every confirmed cause rests on at least two independent signals, at least one from recorded history (snapshot, field history, activity log). Interview testimony alone never confirms a cause.
- **Backtest**: the same fingerprint rules, run on an earlier closed period, flag at least 80% of the deals that actually slipped or were lost from commit in that period. Below 80%, the rules fit the story rather than the data - revise and re-run.
- **Actionability**: every shipped fix names its cause class, its owner, its effort tier, and the single metric that will show within one period whether it worked. The fix list ships in leverage order with the impact order printed beside it and every move between them explained; a list carrying only one of the two orders fails this threshold.

After fixes ship, keep scoring each period: signed bias and absolute error per level, rep calibration against their own baseline. Improvement in the named metrics - not a quieter forecast call - is what closes the investigation.

## Optional integration note

Skip this section unless the user names one of these platforms.

**Salesforce:**

- Forecast categories (Commit / Best Case / Pipeline / Omitted) are distinct from stage and independently editable; diagnose both fields.
- Field history tracking is opt-in per field and not retroactive. Enable it on stage, close date, amount, and forecast category immediately, even if this period must be diagnosed without it.
- The Pipeline Inspection feature surfaces week-over-week deal changes on supporting editions.

**HubSpot:**

- Deal-stage probabilities are editable defaults, not measured conversion. Replace them with the team's own history before trusting any weighted number.
- Manual forecast submissions are recorded and can be compared against the computed roll-up.

## Reference

- See [references/root-cause-fingerprints.md](references/root-cause-fingerprints.md) for per-cause detection signals, discriminators, and fixes, plus how to derive thresholds from the team's own history.
- See [references/worked-diagnosis-example.md](references/worked-diagnosis-example.md) for the report shape, a worked diagnosis, and a negative example (a wrong diagnosis and why it failed).
- See [references/evidence-and-metrics.md](references/evidence-and-metrics.md) for error-metric choices, the rep calibration scorecard, and which circulating statistics are safe to repeat.
- See `mbfinotti/revops-skills@pipeline-stage-definition-audit` to judge whether stage definitions are buyer-verifiable when the fingerprints implicate the definitions themselves.
- See `mbfinotti/revops-skills@sales-pipeline-hygiene` for the recurring audit that prevents the data-class causes from regrowing.
- See `mbfinotti/revops-skills@crm-data-governance` for field ownership and source-of-truth rules behind data-class fixes.
- See `mbfinotti/revops-skills@revenue-reporting` for turning a finished diagnosis into the executive or board narrative.
sales-forecast-diagnostic · 熱門 Agent Skills | Mengbi