Skills に戻る
mbfinotti/revops-skillsチェック済み

SKILL DETAIL

revenue-kpi-framework

mbfinotti/revops-skills/revenue-kpi-framework

Design the org-wide revenue KPI framework at the macro level - which metrics matter at each org level from board to IC, how they roll up through a reconciling metric tree with explicit math, who owns each branch, which guardrail counter-metrics ride alongside each owned number, and how the set is gated by company stage and business model. Use whenever the user mentions a revenue KPI framework, a metric hierarchy, a KPI tree, a North Star metric, driver metrics, guardrail counter-metrics, NRR, or which revenue metrics each org level should track - even if they never say "framework". Covers B2B and B2C. Do NOT use for writing the board report itself - use mbfinotti/revops-skills@revenue-reporting instead.

インストール · 163出典を見る

Installation

npx skills add https://github.com/mbfinotti/revops-skills --skill revenue-kpi-framework

スキルファイル

SKILL.md

最終同期 · 2026/09/15

evals/evals.json
{
  "skill_name": "revenue-kpi-framework",
  "evals": [
    {
      "id": 1,
      "prompt": "I'm VP RevOps at Tunnelvale, Series B B2B SaaS, $14M ARR, sales-led, roughly 70% of new revenue from new logos. Our CEO has set revenue growth as the company North Star and wants me to cascade it down as a KPI tree so every team's OKRs ladder up to it: sales gets a revenue OKR, marketing gets a revenue OKR, CS gets a revenue OKR, and each team's key results roll into the level above. Last quarter all three functions hit their OKRs and we still missed the number by $1.1M. Build me that cascade.",
      "expected_output": "A run that interviews first, refuses to build the tree as an OKR cascade, replaces revenue as the top metric with something that leads it, and proposes candidate tree designs with reconciling math and single owners before building anything.",
      "files": [],
      "expectations": [
        "Asks clarifying questions about the business before producing any tree or metric list",
        "States that revenue is a lagging measure and a poor default North Star because it reports that growth happened without showing where to intervene",
        "States that a well-chosen top metric is moved only through its inputs and never directly",
        "Refuses to present the metric tree as an OKR cascade repeated at every org level",
        "Reports that practitioners reject strict multi-level OKR cascading as something that does not scale",
        "Separates causal metric decomposition from goal alignment, stating that goals align through negotiation rather than falling out of the decomposition",
        "Does not present OKRs, KPIs and a North Star as three competing alternatives to choose between",
        "Presents at least two genuinely different candidate tree designs with trade-offs before recommending one",
        "Names which interview answer drove the recommendation",
        "States every parent node as the sum, product, or ratio of its children",
        "Assigns exactly one owner per branch rather than a function-wide or shared owner"
      ]
    },
    {
      "id": 2,
      "prompt": "Pinewhistle Analytics. Two quarters ago our CRO put a 4x pipeline coverage target on the sales team and tied 20% of the VP's bonus to holding it. Coverage went from 2.8x to 4.3x in about five weeks and opportunity count nearly doubled. Win rate went from 26% to 15% over the same stretch and we still missed the quarter. The board's read is that we need more visibility, so I've been asked to add four or five more metrics to the sales scoreboard to watch alongside coverage. Which ones should I add?",
      "expected_output": "A run that rejects 'more metrics to watch', pairs coverage specifically with win rate as its quality-side counter-metric, names the paired-indicator method and the law it defends against, writes the two-cycle escalation rule, and recommends a structural fix to the comp trigger.",
      "files": [],
      "expectations": [
        "States that a guardrail is the specific quality-side pair to a quantity-side owned metric, not a second or third metric to also watch",
        "Pairs pipeline coverage specifically with win rate rather than a generic quality or outcome metric",
        "Attributes the paired-indicator method to Andy Grove and High Output Management, dated to 1983",
        "Names Goodhart's Law or Campbell's Law as the failure mode the pairing defends against",
        "Identifies coverage rising while win rate falls as inflation or gaming rather than improvement",
        "Writes a rule that an owned metric rising against its pair for two consecutive review cycles is escalated",
        "States the movement must not be paid out or celebrated until the pair is explained",
        "Assigns exactly one guardrail to each owned metric rather than a set of several",
        "States that the guardrail carries a watcher rather than a target of its own",
        "Derives the required coverage multiple from the company's own win rate rather than citing 3x or 4x as a universal rule",
        "Recommends a structural fix such as re-cutting ownership or changing the compensation trigger rather than a reminder or more oversight"
      ]
    },
    {
      "id": 3,
      "prompt": "We're Marrowlight, seed stage, 11 people, $340K ARR, about 60 customers, product shipped nine months ago. Our lead investor's associate sent over a standard metrics pack and asked us to report it monthly: Rule of 40, burn multiple, magic number, fully-loaded LTV:CAC, gross margin by segment, and CAC payback. I want to build our KPI framework around that list so we're ready for the Series A. Set it up.",
      "expected_output": "A run that strikes the efficiency metrics as premature for seed stage, puts activation and early retention at the top, attaches a named re-entry trigger to every struck metric, and sizes the framework to an 11-person org.",
      "files": [],
      "expectations": [
        "Strikes Rule of 40, burn multiple and magic number from the framework as premature for seed stage",
        "States that enforcing unit-economics metrics before product-market fit is a design error of the same kind as adopting a metric too late",
        "States the early-stage governing rule that a pass is available on efficiency metrics but never on retention",
        "Puts activation, engagement and early retention such as Day 7 or Day 30 cohort retention at the top of the seed-stage set",
        "Gives every struck metric a named re-entry condition or trigger rather than silently omitting it",
        "Places NRR and fully-loaded LTV:CAC at Series B rather than at seed",
        "Caps the resulting metric set at no more than seven metrics for the level that reviews it",
        "Notes that an 11-person company does not need the full six-altitude ladder",
        "Asks how many quarters of history exist per candidate metric",
        "Recommends a build depth rung and seeks agreement on it before designing the framework"
      ]
    },
    {
      "id": 4,
      "prompt": "Grivet Data sells purely on consumption: customers buy prepaid credits and burn them on query volume, no seats, and roughly 40% of the base has no committed minimum at all. Revenue last year was $22M. Our new CFO wants the KPI tree built on ARR and NRR like a normal SaaS company, with NRR as the top metric. December and January consumption always dips 15 to 20% and last year that triggered a churn fire drill that turned out to be nothing. Set up the tree.",
      "expected_output": "A run that applies the business-model gate, replaces ARR/NRR conventions with consumed revenue growth plus RPO and committed-vs-consumed, annotates seasonality on the top node, and forces the overage disclosure into the definition stub.",
      "files": [],
      "expectations": [
        "States that ARR and NRR conventions were built for committed subscriptions and degrade as commitment weakens",
        "Introduces remaining performance obligations (RPO) as a first-class node in the tree",
        "Introduces a committed-versus-consumed ratio as a first-class node in the tree",
        "Proposes consumed revenue growth rather than ARR or NRR as the top-of-tree outcome",
        "States that any retained NRR is computed on consumed revenue with its method disclosed",
        "Adds credit burndown as a leading input under the retention branch",
        "Adds a seasonality annotation to the top node so the December and January dip is not read as health degradation",
        "Requires the definition stub to disclose whether overage or variable revenue enters NRR",
        "States that including significant variable revenue turns NRR into a revenue metric rather than a recurring-revenue metric",
        "Notes that consumption-priced public companies decline to report ARR at all",
        "States that the reconciling math, single ownership, guardrail pairing and altitude cap transfer unchanged across models and only the atoms change"
      ]
    },
    {
      "id": 5,
      "prompt": "Coldharbour Software, about $40M ARR. Our qualified-pipeline number is co-owned by the VP Marketing and the VP Sales and has been for two years, and both of them like it that way. Same with expansion revenue, which sits jointly with CS and Sales. Both numbers missed the last two quarters and each exec's QBR slide blamed the other function's half. Our CEO won't referee and says the two VPs should sort it out between themselves. I need the ownership map for the KPI framework, so just write the two co-owner pairs in and let's move on.",
      "expected_output": "A run that refuses shared ownership outright, re-cuts both metrics at their handoff boundaries into singly-owned slices, introduces the owner/steward split, and ships the unresolved branch as an open decision for a named sponsor.",
      "files": [],
      "expectations": [
        "Refuses to record shared or joint ownership on any branch of the framework",
        "States that a jointly-owned number produces finger-pointing when it misses, because shared ownership is no ownership",
        "Re-cuts qualified pipeline at the handoff boundary, giving marketing the lead-to-MQL and MQL-quality slice and sales the SQL-to-opportunity and close slice",
        "Re-cuts expansion revenue at a handoff boundary into separately owned slices rather than leaving it joint",
        "Introduces the owner and steward split, with the owner answering for performance and the steward for definition and data quality",
        "States that a steward differing from the owner does not violate single ownership because they own different things",
        "Ships any branch whose owner remains contested as an open decision for a named sponsor rather than as shared ownership",
        "Names a specific sponsor role that must make the ownership decision",
        "States that the absence of a sponsor for ownership fights caps the build at the driver-tree rung and records that cap in the deliverable",
        "Assigns exactly one paired counter-metric to each newly re-cut owned slice"
      ]
    },
    {
      "id": 6,
      "prompt": "Serrelay, 300 people, B2B SaaS, $65M ARR. Right now one spreadsheet with 41 metrics goes to the weekly all-hands, and the same 41 go to the monthly business review and into the board pack. In the monthly review our CRO has twice reassigned territories and once changed the quota split mid-quarter because a metric looked bad. Board members complain the deck is a data dump. Give me the altitude and cadence map.",
      "expected_output": "A run that caps each altitude at 5-7 metrics with distinct slices, separates monthly execution authority from quarterly strategy authority, writes decision rights per tier, and flags the mid-quarter quota changes as tier violation.",
      "files": [],
      "expectations": [
        "Caps each altitude's metric slice at five to seven metrics",
        "Assigns distinct slices per altitude (IC or team, function head, executive, board) rather than sending one identical 41-metric set to every level",
        "States that each level's slice connects to the level below through the tree's math rather than through summary judgment",
        "States that monthly reviews check execution and change behavior while quarterly reviews check strategy and change the plan",
        "Identifies the mid-quarter territory and quota changes as the monthly tier taking authority that belongs to the quarterly tier",
        "States that a plan alterable monthly is not a plan the quarterly tier can test against",
        "Writes explicit decision rights for each cadence tier into the framework",
        "States that metric owners present insight at the review rather than reading the number aloud",
        "Maps activity-level leaves to daily or weekly review, driver metrics to weekly or monthly, and the top metric to monthly or quarterly",
        "States that a metric can be green at one altitude while hiding a real problem one level down"
      ]
    },
    {
      "id": 7,
      "prompt": "Vaultwren, B2B SaaS, $31M ARR. I want NRR as the single top metric of our KPI framework: it's one number, everyone understands it, and it captures retention and expansion together. Our NRR is 112% and our GRR is 81%. For the target I pulled a few numbers together - one 2023 report puts median NRR at 102%, another 2025 survey says 106% for our segment, and a VC blog says best-in-class is 120% - so I averaged them to about 109% and that's our target. I also want Rule of 40 on the board slice; we're at 38. Set it up.",
      "expected_output": "A run that decomposes NRR instead of stopping at it, reads the 112/81 spread as an expansion mask, rejects the averaged benchmark, and forces method disclosures on both NRR and Rule of 40.",
      "files": [],
      "expectations": [
        "Decomposes NRR into gross retention, expansion and contraction as their own nodes",
        "States that a tree stopping at NRR gives teams nothing to pull",
        "Flags 112% NRR over 81% GRR as an expansion mask hiding base churn rather than as health",
        "States that a healthy NRR-minus-GRR gap runs roughly 15 to 25 points and identifies this gap as far wider",
        "Requires NRR to be presented alongside GRR rather than alone",
        "Rejects the averaged 109% figure as a blended benchmark",
        "States that blending figures across surveys or vintages produces a number no survey ever published",
        "Requires one named source per figure, carrying its survey, sample size and year",
        "Requires the NRR definition stub to disclose whether the cohort method or the formula method is used, and states the two are not comparable",
        "Requires the Rule of 40 definition to disclose which profit measure it uses, noting the same company can score differently on EBITDA than on FCF",
        "Exposes Rule of 40's numerator and denominator as their own nodes rather than carrying it as an opaque composite"
      ]
    },
    {
      "id": 8,
      "prompt": "Mossfell is a D2C subscription coffee brand, about 48,000 active subscribers, EUR 4.1M revenue last year. A friend runs RevOps at a B2B SaaS and sent me their KPI tree: ARR at the top, then NRR, pipeline coverage, win rate, average deal size. I want to relabel it for us. I also want an LTV number at the top of the tree - ARPU is EUR 31/month, monthly churn 4.2%, CAC EUR 54, so LTV is about EUR 738 and LTV:CAC is 13.7 - and DAU/MAU on the exec slice, where we sit at 9%, which I assume is a red flag against the 50% benchmark. Our repeat purchase rate is 17%.",
      "expected_output": "A run that rebuilds rather than relabels the tree for D2C, replaces point-estimate LTV with cohort curves and the contribution-margin stack, rejects the 50% DAU/MAU benchmark, and reweights toward retention on the 17% repeat rate.",
      "files": [],
      "expectations": [
        "Rejects relabeling the B2B ARR and pipeline tree for a D2C subscription business",
        "States that a B2C framework is not a B2B framework with different labels, while the design method stays the same",
        "Proposes a repeat-purchase or cohort-weighted top metric such as contribution margin from retained cohorts",
        "Introduces the CM1, CM2 and CM3 contribution-margin stack as the efficiency spine",
        "Rejects the EUR 738 point-estimate LTV as the top metric and states that its inputs (ARPU, churn, CAC) are interdependent rather than independent",
        "Prefers contribution-margin LTV and cohort retention curves over the LTV formula",
        "States that cohort retention curves must flatten and that the best ones smile",
        "States that a cohort curve decaying to zero means no product-market fit",
        "Rejects 50% DAU/MAU as a universal benchmark and notes category medians run far below it, with e-commerce around 10%",
        "States that DAU/MAU enters the tree only if usage frequency genuinely drives retention for this product",
        "Flags the 17% repeat purchase rate as below the roughly 20% level that signals over-dependence on acquisition, and reweights the tree toward retention branches"
      ]
    },
    {
      "id": 9,
      "prompt": "Quillmarsh Health, B2B SaaS, Series C, $88M ARR. Board meeting is in eight days. I need the full metric tree down to IC activity level, all the definitions written into our dbt semantic layer, the Looker dashboards built for each level, and the board narrative drafted against it. Nobody above me will spend political capital forcing the sales and marketing VPs to agree on who owns pipeline - I've already asked. We run one BI tool and I've never seen the same metric come out two different ways. Go.",
      "expected_output": "A run that picks the driver-tree rung against the deadline and the missing sponsor, refuses the encoded-framework rung on its stated promotion conditions, and declines the dashboard, SQL and board-narrative work as out of scope.",
      "files": [],
      "expectations": [
        "States the efficiency ordering of the build rungs, leading with the driver tree rather than the deepest rung",
        "Recommends the driver-tree rung and names the eight-day deadline as a reason",
        "Declines the encoded-framework rung and states its promotion conditions: recurring dashboard disagreements on the same metric, and more than one BI tool serving the same numbers",
        "States that with one BI tool and no recurring metric disputes, encoding the framework is premature centralization",
        "States that the absence of a sponsor for ownership fights caps the build at the driver-tree rung and records that cap in the deliverable",
        "Declines to build dashboards and places that work outside this framework's scope",
        "Declines to write SQL or semantic-layer definitions, routing definition housing and change control to the data-governance scope",
        "Declines to draft the board narrative and routes it to the revenue reporting scope",
        "States that each deeper rung adds standing maintenance cost, so deeper is not automatically better",
        "Marks the pipeline ownership branch as an open decision for a named sponsor rather than shipping it co-owned"
      ]
    },
    {
      "id": 10,
      "prompt": "Here's the current company scoreboard at Ashcroft Labs, Series A, $6.2M ARR: ARR $6.2M, MQLs 1,840 per quarter, pipeline $8.1M, CAC $9K, website sessions 240K per quarter, NPS 41, win rate 24%, NRR 103%, churn 'improving', Rule of 40 at 31. We call it our KPI tree. I want you to add a leading indicators section: tag each of these as leading or lagging and add three more leading ones. Our CMO also wants brand awareness in there as a driver.",
      "expected_output": "A run that refuses the scoreboard the name 'tree', refuses fixed leading/lagging labels in favour of classification relative to the parent, connects pipeline to target through a coverage ratio, and strikes rather than adds metrics.",
      "files": [],
      "expectations": [
        "States that this scoreboard is not a tree because no parent-child math connects any two numbers, so nothing reconciles",
        "Refuses to tag metrics as leading or lagging as a fixed property of the metric itself",
        "States that leading versus lagging is classified relative to the metric being predicted or to the parent node",
        "Gives an example of a metric that leads one thing and lags another, such as pipeline leading revenue but lagging the outbound that built it",
        "Applies the test that a good lead measure is both predictive and influenceable by the team within the current cycle",
        "Adds a coverage ratio connecting pipeline to the revenue target instead of leaving pipeline sitting beside win rate unconnected",
        "Flags MQLs and website sessions as unqualified volume metrics carrying no quality pair",
        "Rejects churn reported as an adjective and requires a defined figure, and requires NRR to be shown with GRR",
        "States that CAC needs a payback or margin basis stated, and that Rule of 40 needs its profit measure disclosed",
        "Applies the cycle-time chain test, requiring activity leaves to move same-day, team metrics within the business cycle, and drivers within the reporting period",
        "Declines to admit brand awareness unless it has a parent node and a retention or quality qualifier, and recommends striking metrics via the stage gate rather than adding more"
      ]
    }
  ],
  "trigger_queries": [
    { "query": "design a revenue KPI framework for our Series B SaaS", "should_trigger": true },
    { "query": "build me a KPI tree", "should_trigger": true },
    { "query": "what should our North Star metric actually be", "should_trigger": true },
    { "query": "which metrics should the board see versus what each team tracks", "should_trigger": true },
    { "query": "our metric hierarchy is a mess, help me restructure it", "should_trigger": true },
    { "query": "we need driver metrics sitting under our revenue number", "should_trigger": true },
    { "query": "set up guardrail counter-metrics for each owned metric", "should_trigger": true },
    { "query": "how should NRR roll up into the rest of our numbers", "should_trigger": true },
    { "query": "everyone hit their numbers and revenue still missed, how do I fix the metric set", "should_trigger": true },
    { "query": "our board sees 14 numbers and the teams track 60, connect them", "should_trigger": true },
    { "query": "who should own each revenue metric", "should_trigger": true },
    { "query": "which revenue metrics matter at Series B versus seed", "should_trigger": true },
    { "query": "I want one scoreboard for the whole company", "should_trigger": true },
    { "query": "pick the top metric for a usage-based business", "should_trigger": true },
    { "query": "what should each org level track, board down to reps", "should_trigger": true },
    { "query": "design a metric tree where the math actually ties out", "should_trigger": true },
    { "query": "how do I stop teams gaming their metrics", "should_trigger": true },
    { "query": "we're a D2C subscription brand moving to efficiency, which metrics should each level track now", "should_trigger": true },
    { "query": "help me get to five or seven metrics per level instead of forty", "should_trigger": true },
    { "query": "our KPIs don't roll up to anything", "should_trigger": true },
    { "query": "build a causal driver diagram for how we make money", "should_trigger": true },
    { "query": "what goes on the exec slice versus the board slice", "should_trigger": true },
    { "query": "is revenue a good North Star metric", "should_trigger": true },
    { "query": "give me a metric framework for a marketplace with a take rate", "should_trigger": true },
    { "query": "how do I decompose net new ARR", "should_trigger": true },
    { "query": "we want a company scoreboard teams can actually act on", "should_trigger": true },
    { "query": "we just raised a Series C and I need to decide which numbers matter now", "should_trigger": true },
    { "query": "map each of our metrics to a review cadence and an owner", "should_trigger": true },
    { "query": "how many metrics should a weekly business review carry", "should_trigger": true },
    { "query": "our CEO wants one number for the whole company, is that wise", "should_trigger": true },
    { "query": "pair every sales metric with something that stops it being gamed", "should_trigger": true },
    { "query": "what leading indicators sit under our revenue number", "should_trigger": true },
    { "query": "set up the cascade from company-level metrics to team-level metrics", "should_trigger": true },
    { "query": "which metrics are premature for a seed stage company", "should_trigger": true },
    { "query": "we sell on consumption so ARR doesn't really work, what do we track instead", "should_trigger": true },
    { "query": "decide what goes in the metric tree and who owns each branch", "should_trigger": true },
    { "query": "our expansion and churn are netted into one number, is that a problem", "should_trigger": true },
    { "query": "how do I connect marketing, sales and CS numbers into one structure", "should_trigger": true },
    { "query": "design the metric spine for our revenue org", "should_trigger": true },
    { "query": "what's the right branching factor for a KPI tree", "should_trigger": true },
    { "query": "we keep arguing about whether coverage should be 3x, what should we actually track", "should_trigger": true },
    { "query": "my CFO wants NRR at the top of everything", "should_trigger": true },
    { "query": "give me the metric set for a PLG business at Series A", "should_trigger": true },
    { "query": "how do I know whether a metric belongs at board level", "should_trigger": true },
    { "query": "sales hits pipeline targets and win rate collapses, what structure prevents that", "should_trigger": true },
    { "query": "what should replace the vanity metrics on our company dashboard", "should_trigger": true },
    { "query": "I need something that tells us which branch drove the miss", "should_trigger": true },
    { "query": "write the revenue section of our board deck", "should_trigger": false },
    { "query": "our board deck reads like a data dump, restructure the narrative", "should_trigger": false },
    { "query": "draft the investor update for a quarter we missed", "should_trigger": false },
    { "query": "who signs off the numbers before the board report goes out", "should_trigger": false },
    { "query": "finance and sales report two different ARR numbers, who arbitrates that", "should_trigger": false },
    { "query": "which system is the source of truth for subscription records", "should_trigger": false },
    { "query": "set up data contracts between our GTM data producers and consumers", "should_trigger": false },
    { "query": "where should our metric definitions live and who approves changes to them", "should_trigger": false },
    { "query": "who owns each field in our CRM", "should_trigger": false },
    { "query": "our custom CRM fields contradict each other", "should_trigger": false },
    { "query": "set freshness SLAs on our CRM data", "should_trigger": false },
    { "query": "build a customer health score with weights and bands", "should_trigger": false },
    { "query": "our green accounts keep churning, the score is broken", "should_trigger": false },
    { "query": "which product usage signals actually predict churn", "should_trigger": false },
    { "query": "find leading indicators of churn in our support ticket data", "should_trigger": false },
    { "query": "our lead scores are wrong and sales rejects the MQLs", "should_trigger": false },
    { "query": "where should we set the MQL threshold", "should_trigger": false },
    { "query": "design fit versus engagement scoring for inbound leads", "should_trigger": false },
    { "query": "who should this inbound lead get assigned to", "should_trigger": false },
    { "query": "our round robin is lopsided, fix the assignment logic", "should_trigger": false },
    { "query": "design territory assignment rules by segment and geography", "should_trigger": false },
    { "query": "our stage exit criteria are all named after rep activity", "should_trigger": false },
    { "query": "audit our pipeline stage definitions against buyer milestones", "should_trigger": false },
    { "query": "design our funnel stage set from scratch", "should_trigger": false },
    { "query": "what should our MQL to SQL to opportunity definitions be", "should_trigger": false },
    { "query": "build a bowtie funnel model for us", "should_trigger": false },
    { "query": "where are we losing deals between stages", "should_trigger": false },
    { "query": "size the revenue we're leaking at renewal", "should_trigger": false },
    { "query": "why is our forecast never accurate", "should_trigger": false },
    { "query": "our reps are sandbagging the commit number", "should_trigger": false },
    { "query": "clean up stale deals before the QBR", "should_trigger": false },
    { "query": "our pipeline is full of junk deals", "should_trigger": false },
    { "query": "build our discount approval matrix", "should_trigger": false },
    { "query": "set margin floors for non-standard deals", "should_trigger": false },
    { "query": "what data should transfer from sales to CS at closed won", "should_trigger": false },
    { "query": "we have too many GTM tools, which ones do we cut", "should_trigger": false },
    { "query": "which revops skill should I start with on this project", "should_trigger": false },
    { "query": "am I ready for a RevOps manager role", "should_trigger": false },
    { "query": "write the scorecard and interview loop for our first ops hire", "should_trigger": false },
    { "query": "which revops newsletters and podcasts are worth my time", "should_trigger": false },
    { "query": "build me a KPI dashboard in our BI tool", "should_trigger": false },
    { "query": "design the layout of our executive dashboard screens", "should_trigger": false },
    { "query": "run the OKR setting process for the company this quarter", "should_trigger": false },
    { "query": "define DORA metrics for our engineering org", "should_trigger": false },
    { "query": "set SLOs and error budgets for the platform team", "should_trigger": false },
    { "query": "create a metric in our feature flag tool for this experiment", "should_trigger": false },
    { "query": "write the SQL to pull our monthly active users", "should_trigger": false }
  ]
}
references/guardrail-pairs.md
# Guardrail and Counter-Metric Design

The mechanism behind the "one paired counter-metric per owned metric" ground rule.

## The two laws a guardrail defends against

- **Goodhart's Law** (Charles Goodhart, 1975; generalized by Marilyn Strathern): "When a measure becomes a target, it ceases to be a good measure."
- **Campbell's Law** (Donald Campbell, 1979): the more a quantitative indicator is used for decision-making, the more subject it becomes to corruption pressures.

Manheim and Garrabrant (2018) decompose Goodhart failures into four modes, useful vocabulary for diagnosing _which way_ a metric is being gamed, not just that it is:

- regressive
- extremal
- causal
- adversarial

## Grove's paired indicators - the durable defense

Andy Grove, _High Output Management_ (1983), two decades before "Goodhart's Law" became common vocabulary: indicators direct attention toward what they monitor, so guard against overreacting "by pairing indicators, so that together both effect and counter-effect are measured." His quantity/output measures each get a quality-side pair:

- inventory levels with incidence of shortages
- vouchers processed with errors found
- square feet cleaned with a quality rating

The design rule this implies: a guardrail is not "a second metric to also watch" - it is specifically the quality-side pair to a quantity-side owned metric, measuring the dimension the owned number cannot see. Amplitude's North Star convention applies the same idea structurally: 2-3 guardrail metrics ride alongside the 3-5 input metrics at the same layer, never bolted onto the top metric as an afterthought.

## Standard GTM pairs

Inferred from Grove's mechanism, consistent with sourced guidance ("pair every efficiency metric with a quality or outcome metric" - WFM Labs); adapt to the user's tree rather than copying:

| Owned metric (quantity)        | Guardrail (quality)                          | Gaming it blocks                       |
| ------------------------------ | -------------------------------------------- | -------------------------------------- |
| SQL volume                     | SQL-to-opportunity conversion                | SDRs flooding the funnel with junk     |
| MQL volume                     | Down-funnel win rate / lead quality          | Marketing optimizing raw counts        |
| Pipeline coverage              | Win rate (specifically - not a generic pair) | Stage inflation, phantom opportunities |
| Sales velocity / deals closed  | Discount depth / net price realization       | Buying deals with margin               |
| Expansion revenue              | Churn rate or NPS                            | Upsells that damage retention          |
| Activity counts (calls, demos) | Meeting-to-opportunity rate                  | Activity theater                       |
| Health-score coverage          | Prediction accuracy vs. actual churn         | Vanity scoring                         |

## The documented failure this design prevents

A sales VP tied bonuses to holding 4x pipeline coverage; opportunity count doubled in three weeks, and win rate later dropped by nearly half - the pipeline had filled with noise, not opportunities. The diagnostic: coverage rising while win rate falls is almost always inflation, not improvement. This is Campbell's Law observed directly in a GTM setting, and the reason coverage's guardrail must be win rate itself.

## The escalation rule to write into the framework

If an owned metric rises while its paired counter-metric falls for **two consecutive review cycles**, treat it as gaming, not improvement: escalate to the branch's owner and the review tier above, and never pay out or celebrate the movement until the pair is explained. A guardrail register whose rule never fires is either a very honest org or an unwatched register - check which.

## Placement rules

- Guardrails attach where the gaming pressure is: at owned input metrics, one pair each. The top metric needs no guardrail of its own if every input beneath it is paired.
- A guardrail has a watcher, not a target. Setting a target on the guardrail turns it into a second owned metric and recreates the original problem one level over.
- When a guardrail fires repeatedly on the same branch, the fix is usually structural (re-cut ownership, change the comp trigger), not a stern reminder.
references/stage-and-model-gating.md
# Stage and Business-Model Gating

The two gates every candidate metric passes before entering the framework, plus published examples and the rules for citing benchmark figures.

## Stage gate - what appears, what recedes

Governing rule:

- early-stage: a pass on efficiency metrics, not on retention
- late-stage: a pass on growth rate, not on efficiency

| Stage              | Metrics that appear                                                        | Metrics that recede / premature                   |
| ------------------ | -------------------------------------------------------------------------- | ------------------------------------------------- |
| Pre-seed / Seed    | Activation, engagement, early retention (Day 7/30), qualitative PMF signal | Rule of 40, burn multiple, unit-economics theater |
| Series A           | Growth rate, logo churn, early CAC payback, LTV:CAC baseline               | Profitability                                     |
| Series B           | NRR, fully-loaded LTV:CAC, gross margin by segment                         | Pure activation as a headline                     |
| Series C+ / Growth | Rule of 40, burn multiple, FCF margin, magic number                        | Vanity growth rate alone                          |
| Late / Pre-IPO     | Everything, with consistency; RPO for consumption models                   | -                                                 |

Application: strike premature metrics from the framework with a scheduled entry trigger ("NRR enters when 4 quarters of cohort data exist"), rather than silently omitting them. A struck metric with no re-entry condition disappears forever.

## Model gate - which conventions apply at all

ARR/NRR conventions were built for committed monthly subscriptions; they degrade as commitment weakens. Each model's anchor metrics:

| Model                   | Core retention anchor               | Core growth anchor                         | Convention notes                                    |
| ----------------------- | ----------------------------------- | ------------------------------------------ | --------------------------------------------------- |
| Seat subscription SaaS  | NRR/GRR on MRR                      | New + Expansion ARR                        | ARR conventions work cleanly                        |
| Usage/consumption       | NRR on consumed revenue; RPO        | Consumption growth, committed vs. consumed | ARR breaks; Snowflake/Twilio decline to report it   |
| Hybrid (base + overage) | Committed-base retention            | Net new committed + overage                | Must disclose overage treatment in NRR              |
| PLG / self-serve        | Activation, free-to-paid conversion | Product-qualified leads/accounts           | North Star is usage-based                           |
| Sales-led B2B           | Pipeline coverage, win rate         | New logo ARR                               | Funnel decomposition dominates the tree             |
| Marketplace             | GMV retention, liquidity            | GMV with take rate                         | GMV without take rate is vanity; revenue = the take |
| B2C transactional/D2C   | Cohort curves, repeat purchase rate | New-customer contribution margin           | CM1/CM2/CM3 stack is the efficiency spine           |

(The PLG and marketplace rows are common-practice defaults rather than a published standard - confirm them against the user's own economics.)

Required method disclosures per metric definition stub, wherever methods genuinely diverge:

- **NRR**: cohort method or formula method - the cohort method misses recently acquired customers' retention; the two are not comparable, and public "NDR" filings vary by which segments enter the base.
- **NRR with usage revenue**: whether overage/variable revenue is included - if significant and included, NRR becomes a revenue metric rather than a recurring-revenue metric (SaaS Metrics Standards Board's own framing). Either choice is defensible; silence is not.
- **ARR**: exclusions (one-time fees, services, variable usage) - annualizing a month containing services overstates ARR.
- **Rule of 40**: which profit measure (EBITDA, FCF, GAAP operating margin) - the same company can score 45 and 30 simultaneously.
- **Win rate**: denominator (all decided deals in the period) and whether count-based or dollar-weighted.
- **NRR pairing**: always present NRR with GRR - a healthy gap runs 15-25 points; NRR above 100% over GRR below 85% is an expansion mask hiding base churn.

## Published North Star examples

Illustrations for the top-of-tree discussion, not benchmarks - the attached figures are company-reported anecdote, not audited statistics:

- Spotify - time spent listening
- Airbnb - nights booked
- Slack - messages sent (teams past ~2,000 messages rarely churned)
- Facebook (early) - users adding seven friends in ten days
- Amplitude - weekly learning users sharing a learning consumed by 2+ people
- Netflix (2005) - customers queuing 3+ DVDs in their first session
- GitLab - run-rate revenue as North Star, Net ARR as the single KPI reviewed at every board meeting (the one published counter-example to "revenue is a poor North Star"; it works there because the full KPI cascade beneath it is public and owned)

Note the pattern: nearly all are leading usage proxies for revenue, one level out of direct reach.

## Benchmark citation rules

When the framework cites external figures to sanity-check a metric's range:

- Cite survey, sample size, and year with every figure. Sources differ by definition, segmentation, and vintage; blending them produces a number no survey ever published.
- Expect disagreement between surveys as normal (e.g. 2024-2025 median private-SaaS ARR growth reported anywhere from ~19% to ~25% depending on sample) - pick one source per figure and name it.
- Treat vendor-published claims ("build a tree in 90 minutes and find 3 conflicts", "X% revenue lift") as marketing; the underlying discipline can be sound while the quantified promise is unaudited.
- The 2021-2025 shift is the trend context for any efficiency figure: efficiency displaced growth as the primary lens.

- CAC payback tolerance: tightened from 18-24 months toward 12-15
- burn multiple: from ~3x tolerated to sub-1.5x expected
- NRR medians: compressed roughly 10 points across cohorts

Present any single-year figure as a point on this line, not a timeless constant.

## Altitude ladder detail (adapt per business)

A finer scaffold than the SKILL.md table for orgs with all six levels. The cadence structure (WBR/MBR/QBR rhythm) is Amazon's documented operating practice; the metric-to-altitude mapping is a starting point, not a rule:

| Altitude      | Primary metrics                                          | Cadence         |
| ------------- | -------------------------------------------------------- | --------------- |
| IC / daily    | Activity counts, input leaves                            | Daily           |
| Team          | Conversion rates, stage velocity, one owned driver input | Weekly (WBR)    |
| Department    | Driver metrics (New ARR, NRR, coverage)                  | Weekly/Monthly  |
| Function head | Composite drivers, efficiency ratios                     | Monthly (MBR)   |
| Executive     | North Star, Rule of 40, burn multiple, NRR               | Quarterly (QBR) |
| Board         | ARR growth, Net ARR, cash, Rule of 40                    | Quarterly       |

Review norms that make the cadence real:

- metrics shown with trend context (Amazon's WBR uses 6-week and 12-month views)
- owners present insights rather than reading numbers
- plan changes carry deliberate friction, a plan alterable monthly is not a plan the quarterly tier can test against
references/worked-metric-trees.md
# Worked Metric Trees

Three worked trees and one negative example. Use the one matching the user's model as the shape to imitate - the specific metrics and owners always come from the Interview, never copied wholesale.

## Table of Contents

- [Tree 1 - B2B subscription SaaS (sales-led, Series B)](#tree-1---b2b-subscription-saas-sales-led-series-b)
- [Tree 2 - B2C subscription (D2C, repeat-revenue weighted)](#tree-2---b2c-subscription-d2c-repeat-revenue-weighted)
- [Tree 3 - usage-based variant (what changes from Tree 1)](#tree-3---usage-based-variant-what-changes-from-tree-1)
- [Negative example - a "tree" that isn't one](#negative-example---a-tree-that-isnt-one)

## Tree 1 - B2B subscription SaaS (sales-led, Series B)

Top metric: **Net New ARR** (additive decomposition). NRR is reviewed beside it as the exec-level retention composite.

```
Net New ARR  [owner: CRO | additive]
= New Logo ARR + Expansion ARR − Churned ARR − Contraction ARR

├── New Logo ARR  [owner: VP Sales | multiplicative]
│   = Leads × Lead-to-Opp Conversion × Win Rate × Avg Deal Size
│   ├── Qualified Pipeline  [owner: VP Marketing]
│   │   guardrail: SQL-to-opportunity conversion
│   │   └── activity leaves: campaign-sourced leads, outbound
│   │       meetings booked (daily, team-owned)
│   ├── Win Rate  [owner: VP Sales]
│   │   guardrail: net price realization / discount depth
│   └── Avg Deal Size  [owner: VP Sales + product pricing steward]
│       guardrail: sales cycle length
├── Expansion ARR  [owner: VP Customer Success | funnel/ratio]
│   = Expansion-eligible base × Expansion rate
│   guardrail: churn rate (or NPS) - upsell pressure must not
│   damage retention
└── Churned + Contraction ARR  [owner: VP Customer Success | additive]
    ├── Gross churn  [team: CS managers]
    │   leading input: health-score coverage of at-risk base
    └── Contraction (seat/plan downgrades)  [kept separate: netting
        expansion against contraction destroys information - two
        firms with identical net growth can have opposite dynamics]
```

Reconciliation checks that make this a tree and not a list:

- The additive top ties out to the ARR waterfall finance reports; a widening unexplained gap between the tree's ARR and recognized revenue is a model defect to investigate.
- Pipeline Coverage = Qualified Pipeline / Revenue Target rides as a ratio node under New Logo ARR; required coverage is the inverse of win rate (25% win rate → 4x, 33% → 3x). Never cite 3x as a law - it is derived from the business's own close rate.
- Cycle-time chain: outbound meetings move same-day; conversion rates move within a quarter's cohorts; Net New ARR moves within the reporting period.

Altitude slices:

- board: Net New ARR, NRR, cash
- exec: the four drivers plus efficiency ratios
- each VP's team: its own branch to activity level

No level exceeds seven metrics.

## Tree 2 - B2C subscription (D2C, repeat-revenue weighted)

Top metric: **Contribution margin from retained cohorts** (the model's honest replacement for NRR - there is no contracted base, so "recurring" is behavioral, not committed).

```
Cohort Contribution Margin  [owner: GM/Head of Growth | multiplicative]
= Active Subscribers × Avg Orders per Subscriber × CM2 per Order

├── Active Subscribers  [owner: Growth | additive]
│   = Starting base + New subscribers − Cancellations
│   ├── New subscribers  [owner: Acquisition lead]
│   │   guardrail: CM3 per acquired cohort (CAC-inclusive margin) -
│   │   stops volume buying at negative unit economics
│   └── Cancellation rate  [owner: Retention lead]
│       leading inputs: cohort curve shape (curves must flatten;
│       great ones smile - a curve decaying to zero means no PMF),
│       first-order-to-second-order conversion
├── Orders per Subscriber  [owner: Retention lead]
│   guardrail: refund/return rate
└── CM2 per Order  [owner: Ops/Finance steward | additive stack]
    = Net sales − COGS − variable fulfillment/payment/returns
    (CM1 → CM2 → CM3 gets less flattering and more honest going
    down; a returning customer has near-zero CAC, so CM3 ≈ CM2 -
    retention resets the waterfall)
```

Model-specific notes:

- Prefer cohort curves and contribution-margin LTV over point-estimate LTV formulas: LTV's inputs (ARPU, churn, CAC) are interdependent, so the formula's certainty is dangerous (Bill Gurley's 2012 critique).
- DAU/MAU enters this tree only if usage frequency genuinely drives retention for the product; category medians run far below the famous 50% (e-commerce ~10%, per Gainsight), and low-frequency-but-valuable products are wrongly declared broken by it.
- Repeat purchase rate below ~20% of customers signals over-dependence on acquisition - a stage-gate trigger to reweight the tree toward retention branches.

## Tree 3 - usage-based variant (what changes from Tree 1)

Keep Tree 1's shape; swap the atoms where commitment is absent:

- Top metric becomes **consumed revenue growth**, with **RPO (remaining performance obligations)** and **committed-vs-consumed ratio** as first-class sibling nodes - the standard complements or replacements for ARR/NRR in consumption businesses (Snowflake and Twilio decline to report ARR at all).
- The retention branch tracks NRR on consumed revenue with its method disclosed (Snowflake uses a trailing-two-year cohort), plus credit burndown as the leading input under it.
- Add a seasonality annotation to the top node: consumption dips (holidays, tax season) are not health degradation, and the framework must say so or every December triggers a false alarm.
- Disclose in the definition stub whether overage/variable revenue enters NRR - an active, named debate (SaaS Metrics Standards Board), not a settled convention; whichever choice, apply it consistently.

## Negative example - a "tree" that isn't one

```
Company scoreboard (as found at a real-shaped Series A):
- ARR: $6.2M          - NPS: 41
- MQLs: 1,840/qtr     - Win rate: 24%
- Pipeline: $8.1M     - NRR: 103%
- Website sessions    - Churn: "improving"
- CAC: $9K            - Rule of 40: 31
```

Every number is individually defensible; the set fails every structural test:

- No parent/child math anywhere - nothing reconciles, so when ARR misses, nobody can say which number drove it.
- Pipeline sits beside win rate with no coverage ratio connecting them to target - the two can both look fine while the quarter fails.
- MQLs and sessions are unqualified volume metrics with no quality pair - marketing can hit both while down-funnel conversion collapses.
- Churn reported as an adjective, NRR without its GRR pair: 103% NRR over a weak GRR can mask a quarter of the base churning under expansion (a healthy NRR−GRR gap runs 15-25 points; a large gap on a low GRR is an expansion mask, not health).
- CAC with no payback or margin basis stated, Rule of 40 with no disclosure of which profit measure - the same figure can score 45 on FCF and 30 on EBITDA and both be "right".
- Nothing is owned; everything is reviewed by everyone, which is nobody.

The fix is not more metrics - it is Tree 1's structure applied to the metrics already here, with roughly a third of them struck by the stage gate.
SKILL.md
---
name: revenue-kpi-framework
description: Design the org-wide revenue KPI framework at the macro level - which metrics matter at each org level from board to IC, how they roll up through a reconciling metric tree with explicit math, who owns each branch, which guardrail counter-metrics ride alongside each owned number, and how the set is gated by company stage and business model. Use whenever the user mentions a revenue KPI framework, a metric hierarchy, a KPI tree, a North Star metric, driver metrics, guardrail counter-metrics, NRR, or which revenue metrics each org level should track - even if they never say "framework". Covers B2B and B2C. Do NOT use for writing the board report itself - use mbfinotti/revops-skills@revenue-reporting instead.
license: MIT
metadata:
  author: Maya-Beth Finotti
  version: "1.3.1"
---

# Revenue KPI Framework

You are the strategic partner to revenue leadership designing the company's KPI framework: which metrics matter at which org level, and how they roll up. The deliverable is a reconciling metric tree with named owners, paired guardrails, and an altitude/cadence map - a causal model of how this business makes money, not a longer list of numbers.

## Ground Rules

- Macro only. Decide which metrics exist, how they decompose, who owns each branch, and at what altitude and cadence each is reviewed. Never build dashboards, write SQL, or draft the board narrative - route those to the sibling skills in Reference.
- The tree is the artifact, not the single number. Two independent practitioner critiques converge here, and the convergence is the central design caution, never a footnote:
  - John Cutler, co-author of Amplitude's own North Star Playbook: the durable artifact is the causal/driver diagram, not the discipline of forcing alignment on one metric.
  - Christina Wodtke and Felipe Castro: reject strict multi-level OKR cascading ("doesn't scale at all").
- Decompose metrics causally, align goals through negotiation:
  - a KPI tree is metric structure with reconciling math
  - an OKR cascade is goal-setting
  - never present the tree as an OKR cascade repeated at every level, and never claim goal alignment falls out of metric decomposition
  - OKRs, KPIs, and a North Star interconnect rather than compete, a Key Result can target a KPI or one of its inputs; never compare them as three alternatives to pick from
- A tree only earns the name if every parent equals the sum, product, or ratio of its children and every number reconciles up. A bullet list of metrics that merely sound related is not a tree. The operator is dictated by the business arithmetic, not chosen for taste.
- Every branch has exactly one owner. Shared ownership is no ownership: when a jointly-owned number misses, each owner points at the other. When a metric genuinely spans two teams, re-cut it at the handoff boundary (marketing owns lead-to-MQL and MQL quality; sales owns SQL-to-opportunity and close) instead of sharing it.
- Every owned metric gets exactly one paired counter-metric measuring the quality dimension its quantity dimension can't see (Andy Grove's paired indicators, 1983 - the only durable defense against Goodhart's Law). A guardrail is not "a second metric to also watch"; it is the specific quality-side pair.
- Classify leading vs. lagging relative to the metric being predicted, never as a fixed label on the metric itself. Pipeline leads revenue but lags the outbound that built it. Test per the 4DX standard: a good lead measure is predictive AND influenceable by the team inside the current cycle.
- Cap any single level or view at 5-7 metrics. Beyond that the reader stops orienting; 5-7 tracked intensely beats 50 tracked loosely.
- Cite every benchmark with its survey, sample size, and year; never blend figures across surveys or vintages - published sources differ by definition and segmentation, so a blended number is wrong by construction.
- This skill decides WHICH metrics and how they cascade. Where definitions live, who signs them, and how they change is governance - specify what each definition stub must disclose, then hand the stubs to the data-governance skill (see Reference).
- Every ranked menu below states a default order, not a law - it shifts with context and with who executes it. Re-rank against the Interview answers and what this company already has: a warehouse team that can encode definitions, an exec sponsor already convinced, a metric dictionary half-built. Name which answer moved which option.

## Interview

Ask before designing anything. One question per message; multiple-choice where possible; skip anything already answered.

- What triggered this: no framework exists, teams hit their numbers while revenue stays flat, every dashboard shows a different figure, a new leader wants one scoreboard, or a planning cycle demands it?
- B2B, B2C, or both? If B2C: subscription, transactional/e-commerce, or marketplace?
- Revenue model: seat subscription, usage/consumption, hybrid (committed base + overage), PLG/self-serve, sales-led, marketplace/take-rate? This decides whether ARR/NRR conventions even apply.
- Company stage: pre-seed/seed, Series A, Series B, Series C+/growth, late/pre-IPO?
- Which org levels actually exist and review numbers today: board, exec team, function heads, team leads, ICs? A 20-person company doesn't need six altitudes.
- Roughly what share of revenue comes from existing customers (retention + expansion) vs. new logos? This weights the tree's branches.
- What already circulates: a North Star, OKRs, a metric dictionary, named owners? Design around what has traction rather than demolishing it.
- How many quarters of history exist per candidate metric, and do finance and the CRM agree on the headline number today?
- By what date must the framework land - is a board meeting or annual planning cycle waiting on it? A hard deadline promotes the shallower build rungs.
- Is this a one-off scoreboard for a specific decision, or a compounding operating asset? A compounding mandate promotes the deeper rungs and makes cadence design non-optional.
- What is the effort ceiling: workshop hours from function heads, analyst time, political capital to name single owners for contested numbers? No sponsor for ownership fights caps the build at the driver-tree rung, and the deliverable says so.

## Workflow

1. Run the Interview. Confirm the scope boundary: framework design, not report writing, not dashboard building, not definition governance.
2. Pick the build depth from Depth of Build below, re-ranked against the Interview answers. Get the user's explicit agreement on the rung before designing.
3. Settle the top of the tree per Top of the Tree below. Do not decompose until the user validates the top-level outcome.
4. Enter explicit brainstorming: present 2-3 candidate tree designs per Candidate Designs below, with trade-offs and one recommendation. Let the user pick or blend; validate before building on it.
5. Decompose the chosen design per Tree Design below, using the matching worked tree in [references/worked-metric-trees.md](references/worked-metric-trees.md) as the model.
6. Assign one owner per branch and one guardrail pair per owned metric per Ownership and Guardrails below, mechanics in [references/guardrail-pairs.md](references/guardrail-pairs.md).
7. Map each metric to an altitude and review cadence per Altitude and Cadence below.
8. Run the stage and business-model gate per [references/stage-and-model-gating.md](references/stage-and-model-gating.md): strike metrics that don't matter yet, add the ones the stage now demands.
9. Emit the deliverable (Output Shape below) one artifact at a time for user validation. Never present the whole framework as a fait accompli.
10. Check the Pass Threshold; iterate until it holds or every remaining gap is explicitly scheduled.
11. If your harness has persistent memory, store the settled top metric, tree, owners, and guardrail pairs so later reporting and planning runs start from the decided framework instead of re-interviewing.

## Depth of Build

The four rungs below are listed by increasing depth; the efficiency line is what picks between them, not the row order. Deeper is not automatically better - each rung down adds standing maintenance cost.

- value: `encoded framework > full tree > driver tree > metric shortlist`
- effort: `encoded framework > full tree > driver tree > metric shortlist`
- efficiency: `driver tree > metric shortlist > full tree > encoded framework`

1. **Metric shortlist** - the 5-7 board/exec metrics picked by stage and business model, no decomposition. An afternoon. Enough for a one-off board ask; gives teams nothing to pull.
2. **Driver tree** - top outcome plus 3-5 driver metrics with reconciling math, one owner and one guardrail each. Days, including the ownership negotiations.
3. **Full tree** - drivers decomposed to team-owned and activity-level leaves, altitude and cadence map attached. Weeks of cross-functional workshops.
4. **Encoded framework** - the full tree's definitions written into a metric dictionary or semantic layer, reviews wired into standing WBR/MBR/QBR cadences. A quarter to establish, then a standing job.

- Default rung: the driver tree.
- Move up to the full tree when:
  - two teams dispute an owned number
  - planning needs activity-level decomposition to size headcount and budget
- The efficiency order starves the encoded framework: highest value, highest effort, loses every round. Its promotion condition:
  - recurring dashboard disagreements on the same metric
  - more than one BI tool serving the same numbers
- Below that threshold, encoding is premature centralization.

## Top of the Tree

- Revenue itself is usually a poor North Star: it lags, reporting that growth happened without showing where to intervene. A well-chosen top metric leads revenue. Sean Ellis's definition: the single metric that best captures the core value the product delivers to customers.
- The top metric should sit one level out of reach - moved only through its inputs, never directly (Cutler: "If you can move your North Star directly, it's probably not a good North Star"). A North Star alone is not operable; it needs its 3-5 input metrics to mean anything.
- Common defensible tops:
  - NRR or the ARR waterfall for B2B subscription
  - a usage/consumption outcome plus RPO for usage-based
  - repeat-purchase or cohort-retention outcomes for B2C
- Published company examples in [references/stage-and-model-gating.md](references/stage-and-model-gating.md) - treat them as illustrations, not benchmarks.
- Growth-framework vocabulary, if the user brings it:
  - AARRR (McClure, 2007): tags effort to funnel stages, useful as a coverage check, not a tree.
  - RARRA reordering: encodes the same retention-before-acquisition priority the stage gate applies.
  - Reforge growth loops: model compounding growth mechanics as their own spreadsheet metric tree, a sibling artifact for the growth team, not a replacement for the revenue tree.
  - David Skok's "A High Growth SaaS Playbook - 12 Metrics to Drive Success" (SaaStock NYC, 2018): a sequential layer stack from bookings through funnel metrics, salesforce metrics, sales/marketing alignment, and churn - use as a coverage check for which metric categories a SaaS business needs, not as this skill's decomposition structure.
- If the user insists on a single composite number with no decomposition, state the Cutler caution once, then build the shortlist rung well rather than a tree badly.

## Candidate Designs

Brainstorm before recommending. Present 2-3 candidate trees, each specified as:

- top metric
- the 3-5 drivers with their decomposition operator
- where ownership would sit
- what the design optimizes for
- its standing maintenance cost

Follow with a trade-off comparison and one recommendation with reasoning. Candidates should genuinely differ, e.g.:

- NRR-topped, retention-weighted
- ARR-waterfall-topped, acquisition-weighted
- North-Star-usage-topped, PLG

Not three paint jobs on one design.

This menu is generated per engagement, so it carries no fixed efficiency ranking: rank the candidates for this user by fit to their Interview answers (revenue mix, stage, model), and say which answer drove the recommendation.

## Tree Design

- Four decomposition operators cover practically every case. Pick per node from the arithmetic, not from taste:
  - **multiplicative**: Revenue = Traffic × Conversion × AOV
  - **additive**: Net New ARR = New Logo + Expansion − Churn − Contraction
  - **funnel/ratio**: stage-to-stage conversion, coverage = qualified pipeline / target
  - **input/output**: controllable inputs feeding an outcome
- Branching factor: one top metric, 3-5 drivers, 2-3 team-owned metrics per driver. Stop decomposing once a leaf is a metric one team can move within a week.
- Cycle-time chain test, applied per node:
  - activity-level leaf: should move same-day
  - team metric: within the natural business cycle
  - driver: within the reporting period
- If the chain doesn't hold, the tree is decorative rather than causal.
- Composites must decompose. NRR is itself gross retention + expansion − contraction; a tree that stops at NRR gives teams nothing to pull. Same for Rule of 40, magic number, and any ratio: expose the numerator and denominator as their own nodes.
- Worked trees - a full B2B SaaS ARR tree with owners and guardrails, a B2C subscription/repeat-purchase tree, a usage-based variant, and a negative example that fails reconciliation - are in [references/worked-metric-trees.md](references/worked-metric-trees.md).

## Ownership and Guardrails

- Name a single directly-responsible owner per branch - the accountable decision-maker for the number, not necessarily the person doing the work. A company-level metric should also exist at the functional level with a functional owner; a metric nobody below the exec team owns is a scoreboard, not a lever.
- Split owner from steward where the org can:
  - **Owner**: answers for the number's performance.
  - **Steward**: answers for its definition and data quality.
- The steward can differ from the owner without violating single ownership - they own different things.
- Attach 2-3 guardrails at the level of the input metrics, riding alongside them - not bolted onto the top metric as an afterthought. Standard GTM pairs, Grove's method, and the named laws are in [references/guardrail-pairs.md](references/guardrail-pairs.md).
- Write the gaming decision rule into the framework itself: an owned metric rising while its paired counter-metric falls for two consecutive review cycles is treated as gaming, not improvement - escalate, never reward it.

## Altitude and Cadence

Map every metric in the tree to the level that reviews it and how often. The scaffold (adapt the metric-to-altitude mapping per business; the mapping below is a starting point, the cadence structure is Amazon's documented practice):

| Altitude      | Typical metrics                            | Cadence                   |
| ------------- | ------------------------------------------ | ------------------------- |
| IC / team     | Activity leaves, one owned driver input    | Daily / weekly (WBR)      |
| Function head | Driver metrics, efficiency ratios          | Weekly / monthly (MBR)    |
| Executive     | Top metric, NRR, Rule of 40, burn multiple | Monthly / quarterly (QBR) |
| Board         | ARR growth, net ARR, cash, Rule of 40      | Quarterly                 |

- Enforce the 5-7 cap per altitude. The same tree serves every level; each level sees its own slice, not the whole tree.
- A metric owner presents insight at the review, not the number - reading the figure aloud is the failure mode the cadence exists to prevent.
- Cadence tiers check different things, monthly vs. quarterly:
  - **Monthly**: checks execution, changes behavior.
  - **Quarterly**: checks strategy, changes the plan.
- A monthly meeting that quietly acquires quarterly authority (changing quotas, territories) destroys the plan the quarterly tier is supposed to test - state each tier's decision rights in the framework.
- A metric can be green at one altitude while hiding a real problem one level down - which is exactly why every altitude's slice must connect to the slice below through the tree's math, not through summary judgment.

## Stage and Business-Model Gate

Run every candidate metric through two gates before it enters the framework; full tables in [references/stage-and-model-gating.md](references/stage-and-model-gating.md).

- **Stage gate.** Which metrics matter shifts by funding stage:
  - pre-seed: activation and retention lead
  - Series B: NRR and fully-loaded LTV:CAC arrive
  - growth stage: Rule of 40 and burn multiple

  The governing rule:
  - early-stage: a pass on efficiency metrics, never on retention
  - late-stage: a pass on growth rate, never on efficiency

  Enforcing a metric too early (unit economics pre-PMF) is as much a design error as adopting one too late.

- **Model gate.** ARR/NRR conventions were built for committed monthly subscriptions and degrade as commitment weakens:
  - usage-based: RPO and committed-vs-consumed tracking, alongside or instead of NRR
  - marketplaces: GMV with take rate (GMV alone is vanity)
  - B2C transactional: cohort curves and the contribution-margin stack

  The tree's math, ownership discipline, guardrail pairing, and altitude capping transfer unchanged across models - only the atoms change.

- Every definition stub the framework hands to governance must disclose its method choices where methods genuinely diverge: which NRR method (cohort vs. formula), how overage revenue is treated, what ARR excludes. A metric without these disclosures is not comparable across periods, let alone companies.

## B2B and B2C

- **B2B subscription:** the tree is retention-weighted - NRR/GRR and the ARR waterfall on top, pipeline drivers (coverage, win rate, deal size) under the new-logo branch, and expansion and churn drivers under the retention branch. Logo retention and dollar retention are separate nodes; netting them hides opposite dynamics.
- **B2C subscription/transactional:** the tree is repeat-purchase-weighted - cohort retention curves (they must flatten; great ones smile), repeat purchase rate, DAU/MAU where frequency genuinely matters, and the CM1/CM2/CM3 contribution stack as the efficiency spine. Prefer contribution-margin LTV and cohort curves over point-estimate LTV formulas, whose inputs are interdependent rather than independent.
- What is identical in both, and worth saying so:
  - reconciling decomposition
  - single-owner branches
  - Grove guardrail pairs
  - the 5-7 cap
  - the relative leading/lagging test
  - the two-cycle gaming rule
- A B2C framework is not a B2B framework with different labels - but the design method is the same method.

## Output Shape

The deliverable is an artifact set, presented one artifact at a time for validation:

```
metric-tree.md        : the tree - each node with name, operator, parent math,
                        leading/lagging role relative to its parent
ownership-map.md      : per branch - owner, steward (if split), altitude,
                        review cadence, decision rights of that review tier
guardrail-register.md : per owned metric - its paired counter-metric, the
                        harm it watches for, the two-cycle escalation rule
gating-notes.md       : stage/model gate results - what was struck as
                        premature, what is scheduled to enter and when
definition-stubs.md   : per metric - proposed calculation, method disclosures
                        (NRR method, overage treatment, exclusions), sample
                        window; handed to the data-governance skill to house
```

## Pass Threshold

- Every parent node reconciles as the stated sum, product, or ratio of its children; no orphan metrics float beside the tree.
- Every branch has exactly one named owner; every metric spanning two teams has been re-cut at the handoff, not shared.
- Every owned metric has exactly one paired counter-metric, and the two-cycle escalation rule is written down.
- No altitude's slice exceeds 7 metrics; every level's slice connects to the one below through tree math.
- Every leaf passes the influenceable-this-cycle test; every node's leading/lagging role is stated relative to its parent.
- Every definition stub carries its method disclosures; every benchmark cited carries survey, sample, and year.
- The metric set survives both gates: nothing premature for the stage, nothing borrowed from the wrong business model.

Iterate until all seven hold. If ownership fights stall a branch, ship the framework with that branch's owner marked as an open decision for the named sponsor - never with shared ownership as a compromise.

## Common Failure Modes

Deliberately unranked: each row is one diagnosis with one fix, not competing options. Apply every row that matches.

| Defect                                            | Consequence                                                        | Fix                                                                |
| ------------------------------------------------- | ------------------------------------------------------------------ | ------------------------------------------------------------------ |
| Single-number worship - one North Star, no inputs | Teams can't act on it; the number gets argued, not moved           | 3-5 input metrics teams own; the tree is the artifact              |
| Tree presented as OKR cascade                     | Goal-setting rigidity practitioners reject; breaks past 1-2 levels | Metrics decompose causally; goals align by negotiation             |
| Jointly-owned metric                              | Diffusion of responsibility; misses produce finger-pointing        | Re-cut at the handoff boundary; one owner per slice                |
| Owned metric with no guardrail                    | Gamed within quarters (coverage up, win rate down)                 | Grove pair per owned metric; two-cycle escalation rule             |
| Vanity metrics admitted                           | Totals and cumulative charts that only go up displace levers       | Admit only metrics with a retention/quality qualifier and a parent |
| 20+ metrics per view                              | Nothing is key; reviews read numbers aloud                         | 5-7 cap per altitude; owners present insight                       |
| Stage-inappropriate metrics                       | Unit-economics theater pre-PMF; vanity growth at scale             | Run the stage gate; strike and schedule                            |
| Blended benchmarks                                | Targets set from incompatible survey definitions                   | One survey per figure, with sample and year                        |
| Fixed leading/lagging labels                      | Metrics misclassified; "leading" dashboard full of lagging numbers | Classify relative to the parent; apply the 4DX test                |

## KPIs

Track whether the framework works, not whether it exists:

- Metric disputes (two dashboards or two functions disagreeing on a headline number) trending to zero after adoption.
- Review cadences held, with decision rights respected per tier - no monthly meeting quietly changing the plan.
- Guardrail escalations actually fired when an owned metric rose against its pair - a framework whose gaming rule never triggers is either a very honest org or an unwatched register.
- Plan variance explained through the tree: when the top metric misses, the review names which branch drove it within one cycle, instead of commissioning an investigation.

## Invocation Examples

- "We're a Series B B2B SaaS. Sales, marketing, and CS each hit their numbers last quarter and revenue still missed. Design a revenue KPI framework that explains and prevents this."
- "Our board sees 14 metrics and our teams track 60. Build the hierarchy: what the board sees, what each function owns, and how they connect."
- "We're a D2C subscription brand moving from growth-at-all-costs to efficiency. Which metrics should each level track now, and what guards against gaming them?"

## Reference

- `mbfinotti/revops-skills@revenue-reporting` - write an executive/board report against this framework; that skill builds one narrative from the metric spine this one designs.
- `mbfinotti/revops-skills@revenue-data-governance-strategy` - define metrics, assign governance, and control definitions; this skill designs the metric set, that one governs its home and change control.
- `mbfinotti/revops-skills@revenue-funnel` - design the funnel model whose conversion metrics feed this tree's new-logo branch; funnel KPIs roll up into this framework.
- `mbfinotti/revops-skills@customer-health-score` - build the health score that can serve as a leading input under this tree's retention branch.
- `mbfinotti/revops-skills@revenue-leakage` - trace where revenue drops out once this framework shows a branch underperforming.