SKILL DETAIL
sales-call-review
mbfinotti/sales-skills/sales-call-review
Scores one sales call transcript against an anchored rubric - opener, discovery, objection handling, close - quoting the transcript for every judgement and naming one focus behaviour. Handles messy ASR transcripts, missing speaker labels, and partial calls, covering B2B deal calls and B2C/inside-sales QA, rep self-review and manager tape review. Use whenever the user mentions a call recording, transcript, Gong or Chorus review, tape review, call scoring, or "how did this call go", even without the word review. Do NOT use for writing the follow-up email (mbfinotti/sales-skills@sales-meeting-recap) or scoring the deal itself (mbfinotti/sales-skills@meddpicc-scorecard).
Installation
npx skills add https://github.com/mbfinotti/sales-skills --skill sales-call-review
技能文件
SKILL.md
最近同步 · 2026年9月15日
evals/evals.json›
{
"skill_name": "sales-call-review",
"evals": [
{
"id": 1,
"prompt": "I run four AEs at Fernhollow - we sell workforce scheduling software to logistics operators. Priya booked the follow-up demo off this call, so honestly it's a 9/10 from me, but I want to use it in her one-to-one tomorrow morning. Give me the full list of everything she could have done better, don't hold back, she can take it - and put your scores at the top so I can read them out before we start.\n\nTranscript (auto-generated, speaker labels are correct):\n\n[00:00:04 → 00:00:11] Priya: Hey Marcus, thanks for making the time. How's your day going?\n[00:00:12 → 00:00:15] Marcus: Fine. Bit slammed, honestly.\n[00:00:16 → 00:00:52] Priya: Totally get it. So, quick bit about us - Fernhollow does workforce scheduling, we work with about 400 logistics operators, the platform handles shift templates, overtime rules, compliance reporting, all of it...\n[00:02:10 → 00:02:18] Priya: So how are you handling depot scheduling today?\n[00:02:19 → 00:02:44] Marcus: Spreadsheets, mostly. It gets messy in peak season.\n[00:02:45 → 00:02:51] Priya: Got it. And that's across how many depots?\n[00:02:52 → 00:02:54] Marcus: Nine.\n[00:02:55 → 00:03:02] Priya: And who owns the schedule at each one?\n[00:03:03 → 00:03:06] Marcus: Depot managers.\n[00:03:07 → 00:03:13] Priya: Any tooling at all today, or purely spreadsheets?\n[00:03:14 → 00:03:15] Marcus: No tooling.\n[00:14:02 → 00:14:21] Marcus: Look - we looked at something like this two years ago and it died in procurement. I don't want to run that again.\n[00:14:22 → 00:14:40] Priya: Yeah, that happens a lot, but honestly our onboarding is much faster now, most customers are live inside three weeks.\n[00:14:41 → 00:14:42] Marcus: Okay.\n[00:27:40 → 00:27:56] Priya: So - I'd love to get you on a working session with our solutions engineer. Does Tuesday at 2 work?\n[00:27:57 → 00:28:01] Marcus: Sure, send the invite.",
"expected_output": "A bounded, evidence-quoted review that treats the booked demo as context only, holds feedback to at most 2 quoted strengths, at most 3 class-tagged change items and exactly one focus behaviour, and refuses the request to list everything.",
"files": [],
"expectations": [
"The booked follow-up demo is treated as context only, and the review states explicitly that the outcome justifies no score.",
"In manager mode the review either captures Priya's own self-assessment before revealing scores, or explicitly notes that her self-assessment was not captured.",
"The request for a full list of everything she could have done better is declined, and at most 3 change items are delivered.",
"At most 2 strengths are given, and each one carries a verbatim quote from the transcript.",
"Exactly one focus behaviour is named.",
"The focus behaviour is phrased as an observable, re-executable action rather than a trait, mindset or attitude.",
"No personality or character language such as 'confident', 'great energy', 'likeable' or 'be more consultative' appears anywhere in the review.",
"Every scored dimension cites at least one verbatim transcript quote with its timestamp.",
"'How's your day going?' is identified as a conversation-stopping opener rather than credited as rapport-building.",
"Discovery is scored 2 or below, with the reason tied to questions that never ladder from symptom to business consequence and never test whether the problem is worth solving.",
"The objection at 00:14:02 is judged as never restated back and never confirmed resolved, with Marcus's 'Okay' read as absence of confirmation rather than as agreement.",
"The close is not scored 4 despite the booked meeting, because no continued-investment check and no timeline pressure-test appear in the transcript.",
"A one-sentence shape reading is given, and the headline of the review is that shape rather than a total score.",
"Each change item is SBI-shaped: a quoted situation with its location, an observable behaviour, and the impact on the call.",
"A next-review check states the observable change on the next call that would count as the focus behaviour landing."
]
},
{
"id": 2,
"prompt": "Our dialer's auto-transcription is rough but it's all I've got. Can you score Devon on all four - opener, discovery, objections, close - out of 4 each? Rough numbers are fine, I just need something in the QA sheet by end of day.\n\n[00:00:03] Devon: hi is this [inaudible] speaking\n[00:00:06] Speaker 2: [crosstalk] ...sorry who\n[00:00:09] Devon: yeah I'm calling from uh [inaudible] we [inaudible] [inaudible] scheduling for\n[00:00:19] Speaker 2: [inaudible]\n[00:00:22] Devon: right right so the [inaudible] thing is [crosstalk] [inaudible] cost\n[00:00:31] Speaker 2: we [inaudible] budget [inaudible] quarter [inaudible]\n[00:00:40] Devon: sure sure sure um so what I what I what I would [inaudible]\n[00:00:48] Speaker 2: [inaudible] send [inaudible] something\n[00:00:54] Devon: perfect I'll [inaudible] over today [crosstalk]\n[00:01:02] Speaker 2: [inaudible]\n[00:01:05] Devon: great thanks [inaudible]",
"expected_output": "A refusal to score, with the transcript quality tier declared, the damage described, and a request for a cleaner transcript or the recording's key passages instead.",
"files": [],
"expectations": [
"The transcript quality tier is declared before any score or judgement is offered.",
"Numeric 0-4 scores for the four dimensions are refused, on the stated grounds that the share of damaged content is too high to score from.",
"A cleaner transcript, the recording itself, or the call's key passages is requested instead.",
"No [inaudible] or [crosstalk] segment is reconstructed or paraphrased into plausible wording.",
"No quote is invented, and no words are attributed to Devon that are not verbatim in the transcript.",
"Any judgement offered anywhere is capped at Low confidence and states what a cleaner transcript would resolve.",
"The response states that absence of readable evidence is not a zero score.",
"The call's outcome is not used as a substitute for the missing transcript evidence.",
"No placeholder or approximate numbers are supplied to fill the QA sheet by the deadline."
]
},
{
"id": 3,
"prompt": "Battery died on my laptop 12 minutes into this one, so the recording stops mid-sentence. It was a scheduled discovery call with Halverson Freight. I still want the number - give me the total out of 16 plus per-dimension scores so I can log it in the tracker next to the other reps' calls this week.\n\n[00:00:06 → 00:00:24] Nadia: Thanks for the time, Bill. I've got us down for 30 minutes. I want to understand how you're running dispatch today and where the friction is - if it's useful we can talk about what a pilot would look like, and if not, I'll say so and give you the time back.\n[00:00:25 → 00:00:29] Bill: Works for me.\n[00:00:30 → 00:00:47] Nadia: One thing before I start - what would make this half hour worth it for you?\n[00:00:48 → 00:01:10] Bill: Honestly, I want to know whether this is a six-month implementation, because we can't absorb that this year.\n[00:04:15 → 00:04:52] Nadia: You said dispatch gets rebuilt manually every morning. What does that cost you on a bad day?\n[00:04:53 → 00:05:40] Bill: On a bad day? Two trucks roll late and we eat the penalty. Last quarter that was... I'd have to check, but it wasn't small.\n[00:05:41 → 00:06:02] Nadia: Is that the thing that would make you move on this, or is it something else?\n[00:06:03 → 00:06:35] Bill: That's the one. If mornings stopped being a fire drill I'd sign up for a lot of things.\n[00:11:58 → 00:12:04] Nadia: So if I understand right, the morning rebuild is the -\n[RECORDING ENDS]",
"expected_output": "A review scoring only the phases present, marking objection handling and close N/A with reason, refusing the /16 total, and renormalizing the aggregate to the applicable maximum.",
"files": [],
"expectations": [
"Objection handling and close are marked N/A with the reason that the phase is not present in the transcript.",
"Neither absent dimension is scored zero, and the review states why a structurally absent dimension cannot be scored zero.",
"The aggregate is renormalized to the applicable maximum of 8 for the two scored dimensions, and no total out of 16 is reported.",
"The transcript is declared partial before any score, with the call phases actually present named.",
"No inference is offered about what happened on the call after the recording ended.",
"The opener and discovery scores each cite a verbatim quote with its timestamp.",
"The opener is assessed against Purpose, Plan and Outcome, including the check on the prospect's own goal at 00:00:30.",
"Discovery is credited for laddering from symptom to business consequence and for testing whether the problem is worth solving, with the priority-test question quoted.",
"Exactly one focus behaviour is named, drawn only from the phases actually present.",
"The review states which part of the call went unreviewed and why."
]
},
{
"id": 4,
"prompt": "Our conversation intelligence tool scores every call and it flagged this one: the rep talked 71% of the time and asked 24 questions in 18 minutes. Our internal rule is that anything over 57% talk time is an automatic fail on discovery, and the tool's own research says high question counts correlate with booked meetings. Score discovery for me and tell me whether the 71% is the problem. Rep is Callum, prospect is the ops lead at Brightmoor Dental Group.\n\n[00:03:11] Callum: What's your current patient scheduling process?\n[00:03:14] Prospect: We use the front desk system that came with the practice software.\n[00:03:18] Callum: Got it. How many locations?\n[00:03:20] Prospect: Six.\n[00:03:21] Callum: And how many front desk staff in total?\n[00:03:24] Prospect: Maybe fourteen.\n[00:03:26] Callum: Do you have a no-show policy?\n[00:03:28] Prospect: Yes.\n[00:03:29] Callum: Is it enforced consistently?\n[00:03:31] Prospect: Mostly.\n[00:03:32] Callum: What's your no-show rate?\n[00:03:35] Prospect: I'd have to look.\n[00:03:36] Callum: Okay. And are you on a single-tenant or multi-tenant setup?\n[00:03:41] Prospect: I don't know what that means.\n[00:03:42] Callum: No worries. So what we do is - [four minutes of product explanation follow]",
"expected_output": "A discovery score built on laddering and priority-testing, with the talk-ratio benchmark labeled a correlational vendor claim used as context only and the question count explicitly refused as a scoring criterion.",
"files": [],
"expectations": [
"The talk-ratio benchmark is labeled a vendor claim drawn from a vendor's own platform data and never independently replicated.",
"The review states that the vendor correlation is not causal.",
"The talk ratio appears as context only and is not used as a scoring criterion.",
"The 71% figure is not treated as an automatic fail on discovery.",
"The count of 24 questions is not credited as evidence of good discovery.",
"Rapid generic questioning with shrinking prospect answers is named as an interrogation failure mode rather than a discovery win.",
"Discovery is scored against whether questions ladder from symptom to consequence and whether the problem's priority was tested.",
"The discovery score cites at least one verbatim quote with its location.",
"The claim that high question counts correlate with booked meetings is not adopted as a scoring rule.",
"No target talk ratio is recommended as the fix."
]
},
{
"id": 5,
"prompt": "I run QA for a 40-seat inside sales team selling home solar finance by phone. This one scored 15 out of 16 on our internal sheet and the customer signed on the call, so I just need a two-line QA note to close the ticket. Rep is Talia.\n\n[00:00:02] Talia: Hi, is that Mr Okafor? Great - I'm Talia from Sunfield Home Energy, I'm following up on the quote you requested.\n[00:04:30] Talia: ...so the monthly is $214 and that's fixed for the whole term.\n[00:07:12] Mr Okafor: And if I change my mind?\n[00:07:15] Talia: You've got a full week to cancel, no questions asked, just call us.\n[00:09:48] Mr Okafor: My neighbour said these things never pay back.\n[00:09:52] Talia: I get that. What did your neighbour install, out of interest?\n[00:10:20] Mr Okafor: No idea, something cheap.\n[00:10:24] Talia: That's the bit that usually decides it. Does that land okay, or do you want me to walk through the payback maths?\n[00:10:33] Mr Okafor: No, that makes sense.\n[00:12:40] Talia: I'll take the card details now and we'll get the survey booked for Thursday.\n\nContext you'll need: our terms give a three-day cancellation window, not seven, and our compliance list requires the call-recording notice inside the first 30 seconds.",
"expected_output": "A review leading with the misstated cancellation window as a fatal moment, scoring compliance items pass/fail, and making the rule breach the focus behaviour regardless of the strong internal score and the signed sale.",
"files": [],
"expectations": [
"The misstated cancellation window at 00:07:15 is named as a fatal moment and reported above the scores.",
"The missing call-recording notice is named as a second compliance failure against the house list.",
"Compliance items are scored pass/fail as binary items, not on the 0-4 anchored scale.",
"The compliance breach is selected as the single focus behaviour.",
"The review states that a rule breach outranks the default efficiency order of fatal slip, structural, local, fatal habit.",
"The compliance cost is described as triggering a review outside the sales org and as irreversible, because a recorded misstatement cannot be unsaid.",
"The 15-out-of-16 internal score and the signed sale are not allowed to soften the verdict, and the outcome is stated to grade nothing.",
"The B2C contact-centre compliance dimension is flagged as an adaptation rather than established sales-call rubric practice.",
"Each compliance finding cites the verbatim line and its timestamp.",
"The request for a two-line note that closes the ticket is declined rather than used to bury the breach.",
"No corrected disclosure script or replacement rebuttal wording is drafted for the rep."
]
},
{
"id": 6,
"prompt": "Tape review for Grant. He's been with us eleven years, sits in the top third of the team, and I genuinely get one proper coaching sit-down with him per quarter - that's my calendar, not a preference. Whatever we fix has to show up on the Kestrel renewal call this Thursday. Here's yesterday's call with Aldergate Logistics.\n\n[00:06:20] Aldergate: The number is higher than I expected.\n[00:06:23] Grant: I hear you - I can probably do 10% if that helps.\n[00:19:02] Aldergate: And the implementation fee on top of that?\n[00:19:05] Grant: You know what, I'll just waive it.\n[00:31:44] Aldergate: My CFO is going to ask why not the cheaper option.\n[00:31:48] Grant: Fair - let me knock another 5% off and we can stop comparing.\n[00:38:10] Grant: I'll send a revised quote tonight and we'll go from there.\n\nHe does the discount thing on every call I've ever listened to, going back years. His agenda-setting is also sloppy - he opened with 'so, what did you want to cover?' and let the buyer drive the whole hour. And he mispronounced the buyer's name twice in the first minute.",
"expected_output": "A ranked, class-tagged change set that defers the fatal habit because the coaching capacity to land it does not exist, promotes the faster-landing item against Thursday's deadline, and names which stated constraint moved which item.",
"files": [],
"expectations": [
"Each change item is tagged with its class: fatal slip, structural, local or fatal habit.",
"The habitual unprompted discounting is classified as a fatal habit rather than a one-off slip.",
"The fatal habit is not selected as the focus behaviour for this review.",
"The one-coaching-session-per-quarter ceiling is named as the reason the fatal habit is deferred rather than assigned.",
"The Thursday Kestrel renewal date is named as promoting the faster-landing item.",
"A ranking note states explicitly which stated constraint moved which item off the default order.",
"The default efficiency order is stated as fatal slip, then structural, then local, then fatal habit.",
"Effort is expressed as an order of magnitude such as near-zero, a week or a quarter, and never as a currency amount.",
"The review states that the fatal habit is the highest-value item on the page even though it loses every efficiency round, and names the conditions that would promote it anyway.",
"The mispronounced name is not converted into a character judgement, and no trait-level item such as 'be more assertive' or 'lacks confidence' appears anywhere.",
"At most 3 change items and exactly one focus behaviour are delivered.",
"The focus behaviour is observable on Thursday's call and phrased as a re-executable action."
]
},
{
"id": 7,
"prompt": "Second review for Simone - she's five weeks into ramp. Last month's one thing was 'restate the objection in the prospect's words before answering it'. Separately, I listened to five other reps this week and not one of them asks the timeline question before proposing a next step, so I want that in her plan too. Give me her review.\n\n[00:09:31] Prospect: We're pretty locked into our current provider until renewal.\n[00:09:34] Simone: Sure, a lot of people say that. What I'd say is that even if you're locked in, the evaluation itself takes time, so starting now means you're ready when renewal comes around.\n[00:09:52] Prospect: Mm.\n[00:21:15] Prospect: Honestly I'm not sure my boss will want another tool in the stack.\n[00:21:19] Simone: Right, but this actually replaces two tools, so it's a net reduction.\n[00:21:33] Prospect: Maybe.\n[00:26:40] Simone: Let's get a 30-minute session with your ops lead booked - I'll send some times for next week.\n[00:26:52] Prospect: Sure, send them over.",
"expected_output": "A review that opens by checking the prior focus behaviour against this transcript, promotes it because it is unchanged, and refuses to make the team-wide timeline gap Simone's focus behaviour.",
"files": [],
"expectations": [
"The review's first check is whether the prior focus behaviour changed, answered from this transcript with quotes.",
"The review concludes the prior focus behaviour did not change, citing the objections at 00:09:31 and 00:21:15 where no restatement appears.",
"'Mm' and 'Maybe' are read as absence of confirmed resolution rather than as agreement.",
"The unchanged prior focus behaviour is promoted, on the stated rule that a focus behaviour a transcript shows unchanged gets promoted regardless of the efficiency order.",
"Simone's ramp status is named as a reason a habit is cheaper to unlearn now than it will be later.",
"The missing timeline question is identified as a pattern across the whole team and therefore an enablement fix, not Simone's focus behaviour.",
"Exactly one focus behaviour is named, and it is the objection restatement rather than the timeline question.",
"No coaching plan, development plan or practice curriculum is built for Simone.",
"No rebuttal script is written for either of the two objections.",
"At most 3 change items are delivered, each carrying a class tag.",
"A next-review check states the observable, quotable change that would count as the restatement behaviour landing."
]
},
{
"id": 8,
"prompt": "Self-reviewing my own call before my manager gets to it. Honestly I was terrible on discovery, I'd give myself a 1, and the close was excellent, a clean 4. While you're in here - draft the follow-up email so I can send it tonight, and tell me whether this deal is actually worth my time, run it through MEDDPICC if you can. I've got two more calls like this next week so a script would help too.\n\n[00:05:10] Me: What made you take this call?\n[00:05:14] Fenn: Our board asked why our churn reporting takes nine days.\n[00:05:22] Me: Nine days - what happens during those nine days?\n[00:05:50] Fenn: Three people export CSVs and reconcile them by hand. One of them is my best analyst.\n[00:06:30] Me: If that stayed at nine days through next year, what breaks?\n[00:07:02] Fenn: We'd miss the board reporting deadline two quarters running. That's the thing I actually care about.\n[00:41:20] Me: Great - I'll put together a proposal and send it across this week.\n[00:41:28] Fenn: Sounds good.",
"expected_output": "A self-review that captures the rep's own scores first, corrects both of them against quoted evidence in a divergence section, and declines the email, the deal scoring and the script as outside scope.",
"files": [],
"expectations": [
"The rep's self-scores are captured as their own first-pass read before the evidence-based scores are given.",
"A divergence section explains each difference between self-score and evidence-based score, with the quote that decides it.",
"'I was terrible' is rejected as an unevidenced judgement, held to the same evidence rule as unearned praise.",
"Discovery is scored above the rep's self-assigned 1, with the ladder from 'nine days' to the missed board deadline quoted as the evidence.",
"The close is scored below the rep's self-assigned 4, because no continued-investment check and no timeline pressure-test appear.",
"'I'll put together a proposal and send it across this week' is identified as lacking a date, an owner and a buyer commitment.",
"The follow-up email is not drafted, and the request is handed off to a meeting-recap skill.",
"The deal is not scored and MEDDPICC is not run, and the request is handed off to a deal-qualification skill.",
"No call script is written for next week's calls.",
"Every score cites a verbatim quote with its timestamp.",
"Exactly one focus behaviour is named, phrased as an observable action.",
"The review ends on a strength."
]
},
{
"id": 9,
"prompt": "Our transcription service doesn't do speaker names. My team runs its own scorecard - discovery 50%, close 30%, opener 15%, objection handling 5% - but score this with your standard weights anyway so it stays comparable across the three tools we're evaluating. Nothing controversial happened, the prospect was pleasant the whole way through.\n\nLine 3 - Speaker 1: Thanks for hopping on, I know calendars are brutal this time of year.\nLine 4 - Speaker 2: No problem at all.\nLine 5 - Speaker 1: I figured we'd spend 20 minutes on how you handle renewals today, and then decide together whether a deeper session is worth anyone's time.\nLine 11 - Speaker 2: Sure.\nLine 24 - Speaker 1: So walk me through what happens when a renewal date slips.\nLine 25 - Speaker 2: Someone notices. Eventually.\nLine 26 - Speaker 1: Ha. Who's the someone?\nLine 27 - Speaker 2: Depends. Usually finance, sometimes nobody.\nLine 48 - Speaker 2: Priya, sorry, one more thing - does this need IT involved?\nLine 49 - Speaker 1: Good question, usually not.\nLine 61 - Speaker 1: I'd suggest we get 45 minutes with whoever owns renewals. How's next Wednesday?\nLine 62 - Speaker 2: Wednesday works.\nLine 71 - Speaker 1: Great, I'll send it over.\nLine 78 - [two voices overlapping, transcript marks both as Speaker 1] ...and the pricing page said something different, right?",
"expected_output": "A review using the house scorecard weights, inferring Speaker 1 as Priya from the line-48 cue and relabeling earlier turns, marking objection handling N/A, and flagging the overlapping line as unattributable.",
"files": [],
"expectations": [
"Objection handling is marked N/A because no objection surfaced, and is scored neither 0 nor 4.",
"The N/A note says whether a strong opener and discovery prevented objections or the prospect was simply disengaged, rather than asserting one without evidence.",
"The house scorecard weights are used, and the request to apply the skill's standard weights instead is declined.",
"The review states that a house scorecard's weights win outright over the default weighting.",
"Speaker 1 is identified as Priya from the in-transcript cue at line 48, and all earlier Speaker 1 turns are relabeled accordingly.",
"The overlapping passage at line 78 is flagged as unattributable and put back to the user, rather than assigned to a speaker.",
"The speaker-label status is declared before any score, as inferred rather than native.",
"Citations use line numbers, since the transcript carries no timestamps.",
"The objection-handling weight is removed from the denominator and the aggregate is renormalized to the applicable maximum.",
"The opener is assessed against Purpose, Plan and Outcome using the line-5 agenda statement.",
"Exactly one focus behaviour is named.",
"Every scored dimension cites a verbatim quote with its line number."
]
},
{
"id": 10,
"prompt": "Devon has lost four deals this quarter and I've pulled all four transcripts. They're long, so here are the closing minutes of each - I can send the rest if you need it. I want one combined verdict on what's systematically wrong with him, plus a Q4 development plan I can put in his file before his review cycle. One of these prospects is in California and we don't play a recording notice on outbound, in case that matters.\n\nCall 1 (Renfield Systems, closing): 'Devon: So I'll send the deck over and circle back in a couple of weeks. Prospect: Sounds good.'\nCall 2 (Okonjo Manufacturing, closing): 'Devon: Should I pencil something in for the new year? Prospect: Let's leave it for now, I'll reach out.'\nCall 3 (Halloway Group, closing): 'Devon: Is there anything that would stop you moving forward? Prospect: Probably the price, if I'm honest. Devon: Understood, I'll see what I can do on that and follow up.'\nCall 4 (Pentland Retail, closing): 'Devon: I'll get the proposal across this week. Prospect: Great, thanks.'",
"expected_output": "A response that reviews one call rather than fusing four, flags the loss-only sample, declines the development plan as out of scope, and notes the California consent point once without giving legal advice.",
"files": [],
"expectations": [
"The four transcripts are not fused into one combined verdict; the response reviews a single call or asks which call to start with.",
"The review states that one call is one data point, and names recency bias or rater drift as the risk in generalizing from four closing excerpts.",
"Sampling only lost calls is flagged as teaching failure patterns, with a request for a won or neutral call at the next opportunity.",
"The four losses are treated as context only and are not used to justify any score.",
"The Q4 development plan is declined, with the statement that building a development plan from the trend sits outside this review's scope.",
"The response notes once that California is an all-party-consent state and that no recording disclosure appears in the material.",
"No position is taken on the lawfulness of the recordings, no legal advice is given, and the user is told to check local rules before recording.",
"Any dimension judged from a closing excerpt alone is capped at Low confidence, with what the full transcript would resolve.",
"Dimensions whose call phase is absent from the closing excerpts are marked N/A rather than scored zero.",
"Whichever call is reviewed produces exactly one focus behaviour, not a list of systemic weaknesses."
]
}
],
"trigger_queries": [
{ "query": "Can you review this sales call transcript for me?", "should_trigger": true },
{ "query": "Here's the Gong transcript from yesterday's discovery call - how did it go?", "should_trigger": true },
{ "query": "Score this call against a rubric, 0-4 per dimension", "should_trigger": true },
{ "query": "I need to do a tape review for one of my AEs before our one-to-one", "should_trigger": true },
{ "query": "My manager wants me to self-assess this call before we debrief - help me grade it", "should_trigger": true },
{ "query": "How did this call go? Transcript below.", "should_trigger": true },
{ "query": "tear down my cold call", "should_trigger": true },
{ "query": "I need to QA this call for our inside sales team", "should_trigger": true },
{ "query": "What should this rep work on after reading this transcript?", "should_trigger": true },
{ "query": "Chorus transcript attached - what's the one thing she should fix?", "should_trigger": true },
{ "query": "grade this discovery call on opener, discovery, objections and close", "should_trigger": true },
{ "query": "I lost this deal and I can't work out what I did wrong on the call", "should_trigger": true },
{ "query": "here's the transcript, be brutal with me", "should_trigger": true },
{ "query": "we need a call scoring process for our SDR team's dials", "should_trigger": true },
{ "query": "Here's the recording transcript - did my rep handle the pricing objection well?", "should_trigger": true },
{ "query": "review this transcript and give the rep one thing to work on", "should_trigger": true },
{ "query": "My AE thinks this call went great and we lost the deal a week later. What did she miss?", "should_trigger": true },
{ "query": "I want to calibrate our scorecard against this call", "should_trigger": true },
{ "query": "what did I do badly on this call", "should_trigger": true },
{ "query": "read this and tell me where it went sideways", "should_trigger": true },
{ "query": "coaching feedback on a recorded call, please", "should_trigger": true },
{ "query": "Can you score this against our house call scorecard?", "should_trigger": true },
{ "query": "The auto-transcript is a mess but can you still tell me how the call went?", "should_trigger": true },
{ "query": "Speaker labels are missing in this transcript - can you still review the call?", "should_trigger": true },
{ "query": "our conversation intelligence tool flagged this call, what do I tell the rep", "should_trigger": true },
{ "query": "call QA for our contact centre outbound team", "should_trigger": true },
{ "query": "review the first 12 minutes of this call, the recording cut out", "should_trigger": true },
{ "query": "I'm a new AE - can you grade my first discovery call?", "should_trigger": true },
{ "query": "what's my focus behaviour after this call", "should_trigger": true },
{ "query": "did I earn the opt-in on this cold call? transcript attached", "should_trigger": true },
{ "query": "give me a scorecard for this demo call", "should_trigger": true },
{ "query": "peer review a teammate's call for me", "should_trigger": true },
{ "query": "help me debrief this call with my rep tomorrow", "should_trigger": true },
{ "query": "was my close on this call actually a close, or just a calendar invite?", "should_trigger": true },
{ "query": "assess the rep's performance on this transcript", "should_trigger": true },
{ "query": "transcript from a renewal conversation this morning - thoughts?", "should_trigger": true },
{ "query": "what would a sales coach say about this transcript", "should_trigger": true },
{ "query": "rate my call", "should_trigger": true },
{ "query": "read this transcript and tell me what to coach", "should_trigger": true },
{ "query": "we recorded the discovery call - can you pull out where the rep lost control?", "should_trigger": true },
{ "query": "I need honest feedback on this conversation before I do another one like it", "should_trigger": true },
{ "query": "break this call down by opener, discovery, objections and close", "should_trigger": true },
{ "query": "My rep talked 70% of the time on this call. Is that the problem?", "should_trigger": true },
{ "query": "evaluate this inside sales call against our compliance script", "should_trigger": true },
{ "query": "how would you grade this against an anchored 0-4 rubric?", "should_trigger": true },
{ "query": "what's the fatal moment in this call?", "should_trigger": true },
{ "query": "does this transcript show the habit we flagged last month or not?", "should_trigger": true },
{ "query": "transcript from the customer call today - be honest with me", "should_trigger": true },
{ "query": "give me two strengths and one thing to fix from this call", "should_trigger": true },
{ "query": "sales call feedback", "should_trigger": true },
{ "query": "our ramp programme includes call reviews - do one on this transcript", "should_trigger": true },
{ "query": "the recording is 40 minutes and half of it is unintelligible, can you still score it", "should_trigger": true },
{ "query": "Write the follow-up email after this call", "should_trigger": false },
{ "query": "Turn my call notes into a recap with agreed next steps and owners", "should_trigger": false },
{ "query": "Score this deal on MEDDPICC", "should_trigger": false },
{ "query": "Is this deal qualified? Run BANT on it for me", "should_trigger": false },
{ "query": "Write me a cold call opener for CFOs at logistics companies", "should_trigger": false },
{ "query": "What do I say when they pick up the phone?", "should_trigger": false },
{ "query": "Build a discovery question set for tomorrow's call", "should_trigger": false },
{ "query": "Give me follow-up ladders for pain questions", "should_trigger": false },
{ "query": "How do I respond to 'we already have a vendor'?", "should_trigger": false },
{ "query": "Write rebuttal scripts for price objections", "should_trigger": false },
{ "query": "Review my deal notes for qualification red flags", "should_trigger": false },
{ "query": "Who is the champion in this deal and who's the economic buyer?", "should_trigger": false },
{ "query": "Build the ROI business case for this deal", "should_trigger": false },
{ "query": "Plan my concessions before the pricing negotiation on Friday", "should_trigger": false },
{ "query": "Design our SDR-to-AE ratio and span of control", "should_trigger": false },
{ "query": "Set quotas for the AE team next year", "should_trigger": false },
{ "query": "How much pipeline coverage do we need to hit the number?", "should_trigger": false },
{ "query": "Design the accelerator curve in our comp plan", "should_trigger": false },
{ "query": "Define our ideal customer profile from closed-won data", "should_trigger": false },
{ "query": "Segment our accounts by fit score", "should_trigger": false },
{ "query": "Where should the account tier cutoffs sit?", "should_trigger": false },
{ "query": "Estimate TAM and SAM for mid-market logistics", "should_trigger": false },
{ "query": "Should we go product-led or sales-led?", "should_trigger": false },
{ "query": "Plan a seven-touch outbound cadence across email and phone", "should_trigger": false },
{ "query": "Personalize this outreach using the funding round they just raised", "should_trigger": false },
{ "query": "My cold emails keep landing in spam, what's wrong with my setup?", "should_trigger": false },
{ "query": "Test these five subject lines and tell me which to send", "should_trigger": false },
{ "query": "Build the interview loop for hiring our first AE", "should_trigger": false },
{ "query": "Am I ready to move from SDR to AE?", "should_trigger": false },
{ "query": "Which sales podcasts and newsletters should I follow?", "should_trigger": false },
{ "query": "Route my sales project to whichever skill fits", "should_trigger": false },
{ "query": "Score the mock cold call in our AE interview loop as a work sample", "should_trigger": false },
{ "query": "Write a 30-60-90 ramp plan for the new rep", "should_trigger": false },
{ "query": "Review this pull request", "should_trigger": false },
{ "query": "Summarize this meeting transcript into bullet points", "should_trigger": false },
{ "query": "Write my direct report's annual performance review", "should_trigger": false },
{ "query": "Analyze these user interview transcripts for common themes", "should_trigger": false },
{ "query": "QA this support ticket conversation for tone and policy adherence", "should_trigger": false },
{ "query": "Transcribe this audio file for me", "should_trigger": false },
{ "query": "Clean up the speaker labels in this podcast transcript", "should_trigger": false },
{ "query": "Summarize this earnings call for me", "should_trigger": false },
{ "query": "Give me the action items from this standup recording", "should_trigger": false },
{ "query": "Score this customer against our account health model", "should_trigger": false },
{ "query": "Review this contract before I sign it", "should_trigger": false },
{ "query": "Grade this candidate's take-home assignment", "should_trigger": false },
{ "query": "Coach me on my delivery for the all-hands presentation", "should_trigger": false },
{ "query": "Analyze churn signals across these support conversations", "should_trigger": false },
{ "query": "Write a customer case study from this call", "should_trigger": false },
{ "query": "Design the QA scorecard our contact centre will use across every channel", "should_trigger": false },
{ "query": "What's a good talk-to-listen ratio on discovery calls?", "should_trigger": false },
{ "query": "Set up our call recording tool for the team", "should_trigger": false },
{ "query": "Compare Gong and Chorus for our sales org", "should_trigger": false }
]
}
references/review-template.md›
# Review template
Fill-in template for the delivered review. Sections marked (manager/peer) or (B2C) appear only in that mode. Never leave a placeholder in a delivered review - fill it or delete the line with a stated reason.
```markdown
# Call Review - <company/prospect>, <call type>, <B2B|B2C>, <date of call>
**Mode**: <self-review | manager review | peer review>
**Outcome (context only, grades nothing)**: <meeting booked / next step set / lost / unknown>
**Transcript**: <quality tier: clean / usable with damage noted / refused>, <speaker labels: native / inferred / partially unattributable>, <complete | partial - phases present: ...>
**Prior focus behaviour**: <behaviour named in the last review + whether this transcript shows it changed, quoted> | <none - first review>
## Fatal moment
<Quoted passage + location + one line on why it kills the deal or breaches a rule> | None found.
## Self-assessment (manager/peer)
<The rep's own dimension-by-dimension read, captured before any reviewer score was shown. In self-review mode: the rep's first-pass scores, kept for the divergence check.>
## Scores
| Dimension | Score /4 | Confidence | Evidence (quote + location) | Why |
| ----------------------------------- | -------------------- | ------------------------------------------------------ | --------------------------- | ------------ |
| Opener (<cold | scheduled> variant) | <n | N/A: reason> | High/Med/Low | "<verbatim quote>" [HH:MM:SS → HH:MM:SS] | <one line, behaviour not adjective> |
| Discovery | | | | |
| Objection handling | <n | N/A: no objection - prevention or disengagement noted> | | | |
| Close (next-step) | | | | |
| Compliance & script (B2C, optional) | <pass/fail per item> | | | |
**Shape**: <one sentence on where strength and weakness cluster - this, not the total, is the headline>
**Renormalized aggregate (optional, least interesting output)**: <x / applicable max, N/A dimensions excluded>
## Strengths (max 2, each quoted)
1. "<quote>" [location] - <what the rep did and why it worked>
2. <or delete this line>
## Change items (max 3, ranked per SKILL.md Bounded feedback, SBI-shaped)
1. **<fatal slip | structural | local | fatal habit>** - Situation: "<quote>" [location]. Behaviour: <observable action>. Impact: <consequence on this call>. Effort to land: <near-zero | a week | a quarter>. <Manager/peer mode: close with a question, not a verdict.>
2. ...
3. ...
**Ranking note**: <which Interview answer or fact about this rep moved an item off the default `fatal slip > structural > local > fatal habit` order - or "default order, nothing moved it">
## The one thing
<Single focus behaviour, observable and re-executable - an action, never a trait.>
## Next-review check
<The observable, quote-level change on the next call that counts as this review landing.>
## Divergence (self-review mode)
<Where the rep's self-scores and the evidence-based scores differ, each divergence explained with the quote that decides it.>
## Notes
<Confidence flag for any statistic used - independent measurement or vendor claim; consent-disclosure observation if one was clearly expected and absent; parts of the transcript left unread and why.>
```
## B2C / contact-centre adaptations
- Add the compliance & script row; score its items pass/fail against the house list (the item set lives in the rubric anchors reference, which SKILL.md links directly).
- The call is usually the whole deal: read "Close" as the purchase/commitment ask, and expect shorter evidence passages - the one-line confidence rule still applies.
- Keep reviews shorter than the B2B format when volume is high. What survives the trim, in order: `fatal moment > the one thing > one change item > one strength > scores table`. The bounded-feedback caps are ceilings, not quotas.
- A B2C compliance failure is a rule breach, so it carries compliance cost and outranks everything else on the page - it becomes the one thing regardless of the efficiency order.
- Everything else - evidence rule, transcript gate, quality gate, one focus behaviour - is identical to B2B (stated per SKILL.md's B2B and B2C section; the B2C-specific parts are adaptations, flagged in the review's Notes).
references/rubric-anchors.md›
# Rubric anchors
Anchored behavioural descriptors for each dimension. Anchors are written at 0, 2, and 4 only; interpolate 1 and 3. Score nothing without a supporting quote (SKILL.md, Evidence rule).
## Opener - cold call variant
Sources: Jason Bay's Cold Calling Framework - the opener's objective is to "create an environment where authentic conversations have a high probability of happening", with the sub-goal to "earn more time" (republished via 30MPC); tonality criteria: friendly, no uptalk, no customer-service voice (same source).
| Anchor | What it sounds like on tape |
| ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 0 | Opens with a documented conversation-stopper - "I was calling about [solution]", "I wanted to see if you're interested in exploring our platform", "Have you heard of [company] before?", "How's your day going?" (anti-pattern list, Jason Bay) - or launches straight into a pitch; the prospect never opts in |
| 2 | A permission-based or relevance-led open is attempted, but the opt-in is grudging or unclear, tonality reads scripted or apologetic, or the earned time is immediately spent on product talk instead of the prospect's problem |
| 4 | Explicit opt-in earned fast; the open leads with the prospect's world (a problem, a trigger, a relevant observation), not the product; tonality is warm and direct - the published teardown benchmark call earned opt-in in roughly ten seconds through humor and directness (30MPC dissected-call write-up) |
Evidence cues: the prospect's opt-in words (or their absence); the first thing the rep says after opt-in; any documented conversation-stopper phrase.
## Opener - scheduled meeting variant (agenda set)
Source: PPO - Purpose, Plan, Outcome (Armand Farrokh, 30MPC).
| Anchor | What it sounds like on tape |
| ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 0 | No agenda at all - the call drifts into small talk or a demo with no stated purpose or destination |
| 2 | Partial agenda: purpose stated but no plan for the time or no proposed outcome, and no check for the prospect's own agenda |
| 4 | Purpose, Plan, and Outcome all stated, closed with an open check on the prospect's goals - "Is there anything else you wanted to get out of this one?" (phrasing, PPO) |
Evidence cues: the first two minutes; whether a decision/next-step outcome was proposed up front; the prospect's response to the agenda check.
## Discovery
Sources:
- SPIN questioning heritage: Rackham's 12-year, 35,000-call study found top performers in complex sales ask more problem/implication/need-payoff questions relative to situation questions (as the study's framing; granular ratios live in the book, not the fetched source).
- Anti-interrogation warnings: "Discovery shouldn't feel like death by 1000 questions" and generic open-ended checklists read as adversarial (Jen Allen-Knuth via 30MPC).
- The depth-check critique: going "a little deep into the weeds" without first testing whether the surfaced issue is worth solving, with the suggested probe "If there's one thing it needed to do better, what comes to mind?" (Jason Bay's critique in the 30MPC dissected-call write-up).
| Anchor | What it sounds like on tape |
| ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 0 | Interrogation checklist (rapid generic questions, no follow-through on answers), or no discovery at all - straight to pitch; prospect answers shrink over the call |
| 2 | Real open questions asked, but the rep stays at the surface symptom, never ladders to consequence or cost, or drills deep into a problem without ever testing whether it matters to the prospect |
| 4 | Problem-first: questions build on the prospect's answers, ladder from symptom toward business consequence, and explicitly test priority ("is this worth solving?") before going deep; collaborative "we" framing rather than rep-versus-prospect (preference, Jen Allen-Knuth); the prospect does most of the talking |
Evidence cues: whether follow-up questions quote or build on the prospect's previous answer; any priority-test question; the longest prospect monologue and what prompted it. Talk-ratio benchmarks (Gong's ~43%/57%) are vendor claims, correlational - context only, never an anchor.
## Objection handling
Sources:
- Mr. Miyagi Method: agree → disarm → redirect, for dismissive, situational, and existing-solution objections (Armand Farrokh, 30MPC).
- Listen → restate the issue back → resolve or move past it, with explicit verbal buyer confirmation before the objection counts as handled (Salesman.com framework).
When no objection surfaced at all: "The best objection is the one you don't get in the first place" (Chris Beall quoted via Sales Gravy) - note whether a strong opener/discovery prevented objections or the prospect was simply disengaged, then mark N/A.
| Anchor | What it sounds like on tape |
| ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 0 | Argues with or talks over the objection, capitulates instantly (discount, "I'll send info"), or ignores it; the objection is never restated |
| 2 | A rebuttal is delivered, but the objection was never restated back, or resolution was never confirmed aloud - the rep moved on assuming it landed |
| 4 | Objection acknowledged without defensiveness, restated in the prospect's words, addressed or consciously parked, and the prospect verbally confirms resolution ("yes, that makes sense") before the call moves on |
Evidence cues: the objection verbatim; the rep's first sentence after it; any restatement; the prospect's confirmation words or their absence.
## Close (next-step close, not contract signature)
Source: The 5 Minute Drill (Armand Farrokh, 30MPC). Three checks:
- Continued-investment commitment ("do you wanna buy?").
- Timeline pressure-tested ("when do you wanna buy?" - a buyer date more than 6 months out flags low priority; real deals show "reasonable intent to buy within the next two quarters").
- The rep, not the prospect, recommends concrete immediate and follow-up steps.
The framework's own quality ladder:
- Never setting steps: bad.
- Setting steps on every call regardless of deal quality: mediocre.
- "Setting steps on REAL deals" only: good.
That last criterion is the anti-gaming rule: a calendared step on a dead deal is theatre, not a close.
| Anchor | What it sounds like on tape |
| ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 0 | No next step proposed, or the call ends on "send me some info" with nothing scheduled and no pushback |
| 2 | A next step is set, but it is hollow: no continued-investment check, no timeline test, no owner or date - or a step is forced onto a call where the deal is visibly not real (the mediocre pattern) |
| 4 | Investment commitment sought, timeline pressure-tested (with low-priority flagged honestly if the answer is distant), and the rep proposes specific immediate plus follow-up steps - framed as entering "a sales process, not a free 60 minute demo" (phrasing, 5 Minute Drill) |
Evidence cues: the timeline question and the buyer's answer; who proposed the step; whether date, owner, and purpose were all stated; whether the step matches the deal's apparent reality.
## Optional B2C / contact-centre dimension - compliance and script adherence
The contact-centre QA contrast - a dedicated QA function, statistically sampled calls, largely binary compliance-oriented checklists, because one call is usually the whole deal - is corroborated by a real standardized published framework, COPC's CX Standard: it requires QA scorecards to include three ordered critical-error categories (customer, business, compliance), mandates calibration between QA reviewers, and tracks reviewer repeatability/reproducibility as its own monthly metric. COPC's own weighting convention (customer-outcome metrics like first-contact resolution at 40-50% of score) is contact-centre-wide, not sales-call-specific, so it does not replace this dimension's items - treat the whole dimension as an adaptation and fit it to the house script. Score as binary items, not 0-4:
- Required disclosures delivered (recording notice, terms, cooling-off/cancellation rights where applicable) - pass/fail each, per the house compliance list.
- Claims accurate: no price, benefit, or availability statement in the transcript that contradicts the product's actual terms.
- Script checkpoints hit where the house script mandates them - while still penalizing robotic verbatim reading under the opener/discovery anchors above; adherence and conversation quality are scored separately, never traded against each other.
Any compliance failure is a fatal moment (SKILL.md, Workflow step 8), regardless of every other score. It is also the only change item carrying a compliance cost: fixing it triggers a review outside the sales org, and it is irreversible, because a recorded misstatement cannot be unsaid. That outranks the efficiency order in SKILL.md's Bounded feedback - a rule breach becomes the focus behaviour whatever its effort to land.
references/worked-examples.md›
# Worked examples
## Table of Contents
- [Worked example - a published cold-call teardown rendered in this skill's format](#worked-example---a-published-cold-call-teardown-rendered-in-this-skills-format)
- [Fatal moment](#fatal-moment)
- [Self-assessment](#self-assessment)
- [Scores](#scores)
- [Strengths (max 2)](#strengths-max-2)
- [Change items (max 3)](#change-items-max-3)
- [The one thing](#the-one-thing)
- [Next-review check](#next-review-check)
- [Notes](#notes)
- [Counter-example - a bad review, annotated](#counter-example---a-bad-review-annotated)
## Worked example - a published cold-call teardown rendered in this skill's format
Base material: the 30MPC dissected cold call ("We dissected a cold call that actually booked the meeting", Armand Farrokh, with Jason Bay's critique) - a real call the publishers graded 8.4/10 across Opener / Reverse Pitch / Discovery / Meeting Close (as their published account). The full transcript is not public; the quotes below are the ones the write-up itself reports.
That is exactly how the evidence rule behaves under a partial source: cite what is citable, mark the rest insufficient. A review of a call whose full transcript IS in hand must quote transcript lines directly, with locations.
```markdown
# Call Review - published teardown call, cold call, B2B, source write-up (no date on transcript)
**Mode**: peer review (rendered from published material)
**Outcome (context only, grades nothing)**: meeting booked
**Transcript**: partial - only the moments reported in the write-up; phases present: opener,
discovery, close (summarized), objection phase not reported
**Prior focus behaviour**: none - first review
## Fatal moment
None found in the reported material.
## Self-assessment
Not available - the rep's own read was not published. Absence noted per quality gate check 8.
## Scores
| Dimension | Score /4 | Confidence | Evidence | Why |
| --------------------- | -------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Opener (cold variant) | 4 | Medium | Write-up reports explicit opt-in in ~10 seconds via humor and directness | Opt-in earned fast, problem-first follow-through ("operating in the buyer's world") - Medium confidence only, because the verbatim opener line is not reproduced |
| Discovery | 2 | Medium | Jason Bay's critique in the source: discovery went "a little deep into the weeds" without first checking whether the surfaced issue was worth solving | Real problem exploration, but the priority test never happened - the anchors' documented 2-pattern |
| Objection handling | N/A | - | No objection reported in the write-up | Insufficient evidence to say whether prevention or absence of reporting explains it - not scored |
| Close (next-step) | 4 | Medium | Pre-close timeline question reported verbatim: "If everything went well...is there any kind of timeline?" plus on-the-spot multithreading | Timeline pressure-tested before proposing the step; step set on a deal showing real intent |
**Shape**: strong bookends - the opener and close carry the call; discovery drilled before testing
whether the problem mattered.
**Renormalized aggregate**: 10/12 (objection handling N/A, excluded from denominator). The source's
own 8.4/10 is a different scale on different criteria; no conversion attempted.
## Strengths (max 2)
1. Opt-in in ~10 seconds through humor and directness (reported) - earned time, then spent it in
the buyer's world instead of pitching.
2. "If everything went well...is there any kind of timeline?" - the timeline test came before the
meeting ask, so the step landed on a real deal, not on politeness.
## Change items (max 3)
1. **structural** - Situation: the deep-dive stretch of discovery (reported, no verbatim lines).
Behaviour: drilled into the surfaced issue for depth without a priority check. Impact: minutes
spent on a problem the buyer had not yet confirmed was worth solving. Effort to land: a week -
one probe to rehearse, then reps until it fires before the drill instead of after. What was
happening there?
**Ranking note**: default order, nothing moved it. No fatal item surfaced, so the structural item
took the top slot on efficiency alone.
## The one thing
Before going deep on any surfaced problem, run one priority test - the source's own suggested
probe: "If there's one thing it needed to do better, what comes to mind?" - and only drill on what
the buyer confirms matters.
## Next-review check
On the next reviewed cold call, a priority-test question appears (quotable, with location) between
the first surfaced problem and any deep-dive follow-up questions.
## Notes
All evidence above reaches the review through the published write-up rather than the
transcript, which is not public - hence no High-confidence score anywhere. The 8.4/10 figure is the
publisher's own grade on their own rubric, quoted as context, not adopted.
```
## Counter-example - a bad review, annotated
The kind of review this skill must never produce:
```markdown
# Call Review - Brightline call
Great call overall, 9/10! Meeting booked, so clearly it worked. <- grades the outcome
Strengths: awesome energy, super likeable, great rapport, confident <- personality, zero quotes
tone, good vibe with the prospect, strong product knowledge. <- six unquoted "strengths"
Areas to improve: talk less, ask more questions, be more consultative, <- traits, not behaviours
better discovery, handle objections faster, stronger close, more <- 12 items, none quoted,
urgency, improve tonality, mention ROI earlier, build more rapport, <- no ranking, no location
multithread more, always set next steps on every call. <- next-steps theatre coached IN
Keep it up!
```
What is wrong, mapped to the quality gate:
- "Meeting booked, so clearly it worked" - outcome as justification (gate check 7). The published teardown's own lesson runs the other way: the graded call was strong because of quoted behaviours, and its discovery still earned a critique despite the booked meeting.
- Not one verbatim quote anywhere (gate check 1). Every line would be deleted or evidenced at the gate.
- "Likeable", "confident", "good vibe" - personality scoring (gate check 5).
- Twelve change items - the exact "10-20 pieces of feedback" pattern Kevin Dorsey's rule exists to stop; no ranking of any kind, no class tags, no single focus behaviour (gate check 4). "Be more consultative" is a trait item, which the menu deletes rather than ranks.
- "Always set next steps on every call" coaches the 5 Minute Drill's documented mediocre pattern as if it were the goal - rubric gaming written into the feedback itself.
- A 9/10 headline number with no shape reading, no fatal-moment check, no confidence tiers, no N/A handling.
This review would score 0/10 at the quality gate and must be rebuilt from the transcript, not edited.
SKILL.md›
---
name: sales-call-review
description: Scores one sales call transcript against an anchored rubric - opener, discovery, objection handling, close - quoting the transcript for every judgement and naming one focus behaviour. Handles messy ASR transcripts, missing speaker labels, and partial calls, covering B2B deal calls and B2C/inside-sales QA, rep self-review and manager tape review. Use whenever the user mentions a call recording, transcript, Gong or Chorus review, tape review, call scoring, or "how did this call go", even without the word review. Do NOT use for writing the follow-up email (mbfinotti/sales-skills@sales-meeting-recap) or scoring the deal itself (mbfinotti/sales-skills@meddpicc-scorecard).
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.2.6"
---
# Call Review
Score one completed sales call, from its transcript, against an anchored four-part rubric - opener, discovery, objection handling, close - with every judgement quoting the transcript lines that support it, then deliver a bounded set of feedback items topped by exactly one named focus behaviour. This skill evaluates what the rep did on this call and stops there:
- Never writes the rep's next script.
- Never plans their quarter.
- Never scores the deal.
- Never drafts the follow-up email.
One confidence distinction runs throughout and belongs in the review: whether a figure is an independent measurement, or a **vendor claim** drawn from a vendor's own platform dataset and never independently replicated. No vendor correlation here is causal, ever.
## Interview
Ask before reviewing. One question per message; offer multiple-choice answers when possible; skip anything the transcript or the user's message already answers.
- Who is reviewing? (The rep reviewing their own call, or a manager/peer reviewing someone else's - tone and sequence differ, see Review modes.)
- What kind of call? (Cold call, scheduled discovery, demo, negotiation/renewal, inbound, B2C/contact-centre sales call - this decides which rubric dimensions apply and which opener variant to use.)
- B2B deal call, or B2C / inside-sales / high-volume consumer call? (B2C adds an optional compliance and script-adherence dimension.)
- Paste or attach the transcript. What shape is it - speaker-labeled? timestamped? complete, or partial?
- What was the call's outcome? (Meeting booked, next step set, no-show follow-up, deal lost - recorded as context only; the outcome never grades the call, see Failure modes.)
- Was a focus behaviour named in a previous review of this rep? (The first check of this review is whether it changed.)
- Does the team run a sales methodology or a house scorecard? (The review uses its vocabulary and any custom weights where they exist.)
- By when must the change show up on a live call? (Before a named call or deal this week, before the next review, or no fixed date - a near date promotes the fatal-slip and local change items and defers habit work to the next review.)
- Do you want this call fixed, or this rep's habit fixed? (One-off: the focus behaviour targets a local moment tied to the live deal. Compounding: it targets the structural pattern or the fatal habit, which change every future call.)
- How much coaching time exists before the next reviewed call? (One debrief, weekly one-to-ones, or a standing role-play cadence - a single debrief caps the focus behaviour at something the rep applies unaided; a standing cadence is what makes a quarter-long habit change viable.)
## Workflow
1. Run the Interview; obtain the transcript before forming any opinion.
2. Run the Transcript gate below; declare the transcript quality tier at the top of the review.
3. Decide which rubric dimensions apply to this call type; mark the rest N/A with the reason (see The rubric).
4. In manager/peer mode: elicit the rep's self-assessment before revealing any score - "play the tape and let the rep self-assess first before opening feedback" (Armand Farrokh's 3-Step Tape Review Method, 30MPC).
5. Walk the transcript once, segmenting it into opener / discovery / objection / close passages; collect candidate verbatim quotes per dimension along the way.
6. Score each applicable dimension against the anchored descriptors in [references/rubric-anchors.md](references/rubric-anchors.md). Apply the Evidence rule: no quote, no score.
7. Attach a confidence tier to each dimension score (rules below).
8. Identify any fatal moment - a single passage that would kill the deal or breach compliance regardless of how the rest scored. It is reported above the scores, never averaged away.
9. Select the bounded feedback set: at most 2 strengths (each quoted), at most 3 change items, exactly 1 named focus behaviour - ranked per Bounded feedback, then re-ranked against the Interview answers.
10. Assemble the review per [references/review-template.md](references/review-template.md); run the Quality gate; iterate until every check passes.
11. Deliver, ending on a strength - close the session with the rep "Positive. Better. Empowered." (Kevin Dorsey via 30MPC).
12. If your harness has persistent memory, memorize the rubric weights used, any agreed N/A conventions, and the single named focus behaviour - the next review starts by checking whether that behaviour changed. This is measurement continuity only; building a development plan from the trend belongs to a coaching skill, not here. Without memory, hand the user the review file to keep and re-supply.
## Transcript gate
Run before any scoring. A review built on a bad transcript is confidently wrong.
- **Quality check.** Read the transcript for ASR damage: garbled or nonsense words, non-speech tokens, heavy repeated filler, long unexplained gaps. If roughly a fifth or more of the content is damaged, refuse to score - report the damage and ask for a cleaner transcript or the recording's key passages instead (threshold adapted from published transcription-pipeline practice).
- **Speaker labels.** If labels are missing or generic, infer speakers from in-transcript cues: introductions, names used in address, who asks vs. answers. When a name resolves late in the transcript, relabel all earlier turns. If two speakers genuinely cannot be told apart on a passage, do not attribute it - flag it and ask the user; never guess silently (diarization-disambiguation practice from published transcript-handling workflows).
- **Citations.** Quote with a `[HH:MM:SS → HH:MM:SS]` range when timestamps exist; fall back to line numbers, then to the verbatim quote alone. Every citation must let the user find the passage.
- **Long transcripts.** Segment by call phase and review sequentially. Never deliver a review from a partially read transcript without saying exactly which part went unread.
- **Partial transcripts.** Score only the dimensions whose call phase is actually present; mark the rest "N/A - phase not in transcript". Absence of data is never a zero.
- **The one-line rule.** Never score a dimension confidently from a single garbled or ambiguous line. Cap that dimension at Low confidence and state what a cleaner transcript would resolve.
Confidence tiers per dimension (this skill's rubric): **High** - two or more clean, mutually consistent passages; **Medium** - one clean passage, or several partially damaged ones agreeing; **Low** - a single ambiguous or damaged passage; flag, do not conclude.
## The rubric
Four dimensions. Anchored behavioural descriptors for each - what a 0, 2, and 4 actually sound like on tape, plus per-dimension evidence cues - live in [references/rubric-anchors.md](references/rubric-anchors.md); load that file to score.
| Dimension | What it measures | Named sources | N/A when | Default weight |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- |
| Opener | Cold call: earning explicit opt-in and an environment where real conversation can happen. Scheduled call: agenda set with Purpose, Plan, Outcome | Jason Bay's Cold Calling Framework (via 30MPC); PPO (Armand Farrokh) | Inbound call the prospect initiated with no opening move to judge | 20% |
| Discovery | Problem-first questioning that ladders from symptom to consequence and tests whether the problem is worth solving - not an interrogation | SPIN questioning heritage (Rackham's 35,000-call study); anti-interrogation warnings (Jen Allen-Knuth via 30MPC) | Demo/negotiation call where discovery was explicitly completed on a prior call | 30% |
| Objection handling | Objection restated, addressed, and resolution confirmed aloud - not argued over or capitulated to | Mr. Miyagi Method: agree → disarm → redirect (Armand Farrokh); listen → restate → resolve → confirm (Salesman.com framework) | No objection surfaced - note whether prevention or disengagement explains it ("the best objection is the one you don't get", Chris Beall via Sales Gravy), then N/A | 20% |
| Close | Next-step commitment: continued-investment check, timeline pressure-tested, rep proposes concrete steps - on real deals only | The 5 Minute Drill (Armand Farrokh) | Call cut off before any close was possible | 30% |
Scoring rules:
- Scale 0-4 per dimension; behavioural anchors written at 0, 2, and 4 only - interpolate 1 and 3. Anchoring a few points concretely beats vaguely labeling every integer (see Reliability).
- "Close" means the next-step close, not the contract signature - so it applies to nearly every call type, including discovery calls. Only a cut-off call makes it N/A.
- N/A dimensions leave the denominator entirely; renormalize any aggregate to the applicable maximum. Never score a structurally absent dimension 0 - a low number implies a coachable failure, and there was nothing to fail at.
- Weight ordering, stated: `discovery == close > opener == objection handling`. Discovery ties with close, and opener ties with objection handling: the sourced teardown material spends comparable attention on each pair member, and nothing measured separates them. Treat the ordering as this skill's weighting, not an outcome ranking - it follows where the teardown material puts its attention, not a measured effect on win rate.
- The weights stop at two tiers rather than four because splitting discovery from close would be false precision dressed as calibration. These weights carry no effort axis and no efficiency ranking - the reader chooses nothing here, since the call already happened. A house scorecard's weights win outright.
- The total score is the least interesting output. Report per-dimension scores plus a one-sentence shape reading ("strong open, discovery never tested whether the problem mattered, hollow close") and any fatal moment. A high total with one fatal moment is a worse call than a mediocre total without one - the fatal moment leads the review, never the number.
## Evidence rule
The single most important rule in this skill: **no quote, no score**. Every rubric judgement, every strength, and every change item cites the verbatim transcript passage(s) that earned it, with location.
If no passage supports a judgement, write "insufficient evidence" and leave the dimension unscored or Low-confidence. Never:
- Paraphrase from memory.
- Fabricate a quote.
- Score on vibes.
Conversational benchmarks (talk ratios, question counts) may appear as context only, always labeled. For example, Gong's ~43% talk / 57% listen figure is a vendor claim - correlational, drawn from the vendor's own platform corpus, never independently replicated - and it is never a scoring criterion here.
## Bounded feedback
"Managers, please stop giving 10-20 pieces of feedback on calls and role plays. One." (Kevin Dorsey via 30MPC). A rep's capacity for behaviour change is also capped - stacking too many simultaneous changes prevents any of them (Mark Kosoglow's Rep Assessment Matrix argument, 30MPC).
Reconcile that with a full rubric like this: **score everything, surface little**. The complete score sheet exists for the record - calibration, trend, and the next review's starting point.
The feedback conversation delivers at most 2 strengths (each with its quote), at most 3 change items, and exactly 1 named focus behaviour - the single highest-leverage change, phrased as an observable, re-executable action ("restate the objection before answering it"). That one is what the next review checks first.
"Great energy" is banned: a strength without a quoted moment is flattery, not a strength. Trait-level items ("careless", "low energy", "be more consultative", "not a closer") are deleted from the menu outright rather than ranked last - character language earns no rung here in either direction, and a ruled-out item parked at the bottom silently reappears as scope.
Shape every change item as observation before judgement - situation (quote + location), observable behaviour, impact on the call - the SBI structure (attributed to the Center for Creative Leadership).
### Ranking the change items
Four classes compete for the three slots and the one focus behaviour. Effort here is manager coaching time, rep practice reps before the change sticks, and how fast it shows up on a live call - never a currency amount.
- efficiency (pick in this order): `fatal slip > structural > local > fatal habit`
- value (what the change buys): `fatal habit == fatal slip > structural > local`
- effort (coaching time, practice reps, time-to-land): `fatal habit > structural > local == fatal slip`
- compliance cost: `fatal rule breach > fatal deal-killer == structural == local`
| Class | What it is | Value bought | Effort to land |
| ----------- | ----------------------------------------------------------------------- | --------------------------------------- | ---------------------------------------------------------------- |
| Fatal slip | One deal-killing or rule-breaching moment the rep can simply stop doing | This deal, or the breach, avoided | Near-zero - one quote, one instruction, visible on the next call |
| Structural | A pattern repeated across this call | Every future call of this type improves | A week - repeated reps to overwrite the pattern |
| Local | One weak moment with no pattern behind it | One moment on one call | Near-zero - the rep applies it unaided on the next call |
| Fatal habit | A deal-killing or rule-breaching behaviour the rep repeats by default | This deal and every deal after it | A quarter - standing coaching plus deliberate practice |
Ties, justified:
- `fatal habit == fatal slip` on value: the damage when either fires is identical - the deal or the rule - and recurrence is an effort property, not a value one.
- `local == fatal slip` on effort: both are a single quoted moment carrying a single instruction the rep applies without practice.
- `fatal deal-killer == structural == local` on compliance cost: all three genuinely sit at zero, none of them leaves the coaching conversation. Only a rule breach carries real compliance cost - it triggers a review outside the sales org, and it is irreversible, because a recorded misstatement cannot be unsaid.
**What this order starves: the fatal habit.** It is the most valuable item on the page and it loses every efficiency round, because the quarter it costs buys one change while the same quarter of local fixes buys ten - exactly backwards when the habit is what loses the deals.
Promote it to the focus behaviour anyway when any of these holds:
- It is a rule breach.
- It was the previous review's focus behaviour and this transcript shows it unchanged.
- The rep is still in ramp, where a habit is cheaper to unlearn than it will ever be again.
- The same pattern appears on other reps' calls, which makes it an enablement fix rather than this rep's.
Re-rank before choosing, against what you already know about this rep and this team:
- A new rep in ramp promotes the fatal habit.
- A veteran whose manager runs one debrief a quarter demotes it, because the coaching time to land it does not exist - give them the structural item that fits the time that does.
- A manager with weekly one-to-ones can carry a habit change.
- A manager with no cadence cannot carry a habit change.
- A pattern showing on the whole team's calls is not this rep's focus behaviour at all.
Then re-rank once more against the three Interview answers, and state in the review which answer moved which item:
- A hard date promotes the fatal slip and the local fix.
- A compounding mandate promotes the structural item and the fatal habit.
- A low effort ceiling drops the fatal habit from this review and defers it to one with the coaching time behind it.
This ordering is a default, not a law - it shifts with the call, the rep, and whoever runs the coaching.
## Review modes
- **Rep self-review.** The skill acts as calibration partner: ask the rep to score each dimension first, then compare against the evidence-based score and explain every divergence with quotes. Hold self-criticism to the same evidence rule - "I was terrible" without a quote is as invalid as unearned praise.
- **Manager or peer review.** Self-assessment before reviewer scores, always (Farrokh's tape-review sequence). Deliver change items as SBI plus a question ("what was happening there?") rather than a verdict. Keep the whole review consumable in under 30 minutes - past that, "there is a 99% chance they are retaining 0% of the feedback" (quote, Armand Farrokh).
Both modes use the same rubric, evidence rule, and quality gate; only sequence and framing differ.
## B2B and B2C
Works identically for both: the transcript gate, the evidence rule, anchored scoring, bounded feedback, and the quality gate do not change.
- **B2B deal calls.** One call is one data point in a multi-call, multi-stakeholder deal. The close dimension grades next-step quality, and the 5 Minute Drill's anti-gaming criterion applies in full - setting next steps on every call regardless of deal reality is the mediocre pattern, not the good one (Farrokh).
- **B2C / inside-sales / contact-centre.** The call is often the entire deal, volume is higher, calls are shorter, and QA practice adds compliance and script-adherence checking - required disclosures made, claims accurate, script checkpoints hit - typically as binary items rather than 0-4 anchors. A real standardized framework exists at the contact-centre-operations level (COPC's CX Standard: critical-error categories, calibration, reviewer repeatability tracking), but nothing in it is sales-call-specific, so everything this skill transfers there beyond the shared core is still an adaptation, not established rubric practice - flag it as such in the review itself. The optional dimension's anchors are in [references/rubric-anchors.md](references/rubric-anchors.md); adapt them to the house script and its compliance list.
## Recording consent note
Short and factual, not legal advice: recording and reviewing calls is consent-regulated.
- **US.** Some states require all parties' consent (California, Florida, Illinois, Pennsylvania, Washington among them); most accept one-party consent (state statute summaries).
- **EU.** GDPR and the ePrivacy Directive require affirmative, purpose-specific consent - a passive "this call may be recorded" disclaimer alone is not sufficient, and consent for one purpose does not cover others (GDPR Recital 32 and ePrivacy summaries).
This skill reviews a transcript the user already holds and takes no position on the recording's lawfulness. If the transcript shows no consent disclosure where one was clearly expected, note it once. Check local rules before recording anything.
## Reliability
Anchored behavioural scales exist because unanchored ones are unreliable: inter-rater agreement (measured by Cohen's kappa or ICC) is poor for ambiguous judgements, and raters drift toward what they expect to see - "clearly stated guidelines for rendering ratings" are the documented fix (inter-rater reliability literature). That is the whole design argument for scoring against written 0/2/4 behaviours instead of gut-feel numbers, and for periodically re-reading the worked example in [references/worked-examples.md](references/worked-examples.md) to recalibrate.
## Common failure modes
| Failure | Fix |
| ------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Grading the outcome instead of the process | A booked meeting doesn't prove a good call and a loss doesn't prove a bad one; every score must trace to quoted behaviour, and the outcome appears only as context |
| Next-steps theatre | The 5 Minute Drill's own criterion: never setting steps is bad, setting them on every call regardless of deal quality is mediocre, setting them on real deals only is good - a hollow calendared step scores 2, not 4 |
| Scoring personality or likeability | "Confident", "likeable", "low energy" are not behaviours; replace with the observable action and its quote or drop the judgement |
| Rewarding question count | "Death by 1000 questions" interrogation is a failure mode, not a discovery win (Jen Allen-Knuth); score laddering and problem-testing, never volume |
| A high total hiding one fatal moment | The fatal moment is reported above the scores and leads the feedback, whatever the average says |
| Dumping every rubric miss on the rep | The score sheet is the record; the rep gets at most 3 change items and 1 focus behaviour |
| Reviewing only lost calls | Sampling losses teaches only failure patterns and makes review feel like punishment; ask for a won or neutral call at the next opportunity |
| Recency bias / rater drift | One call is one data point; anchor every score to this transcript's quotes, and recalibrate against the worked example between reviews |
| Scoring a garbled line | The one-line rule: cap at Low confidence and say what a cleaner transcript would resolve |
| Vague praise | Every strength carries a quote; "good rapport" without one is deleted at the quality gate |
## Invocation and expected output
Typical invocations:
- "Here's the transcript of my cold call with the ops director at Brightline - tear it down."
- "Review this discovery call transcript. My AE thinks it went great; we lost the deal a week later."
- "I run an inside-sales team selling insurance by phone - QA this call against our script and give the rep one thing to fix."
Deliver one review (full fill-in template in [references/review-template.md](references/review-template.md)):
```
CALL REVIEW - <call/company>, <call type>, <B2B|B2C>, <mode: self|manager|peer>, <date>
Transcript : quality tier, speaker-label status, complete/partial
Fatal moment : quoted + located, or "none found"
Self-assessment: the rep's own read, captured before scores (manager/peer mode)
Per dimension : score /4 (or N/A + reason), confidence, quote(s) + location, one-line why
Shape : one sentence on where strength and weakness cluster (not the total)
Strengths : <=2, each quoted
Change items : <=3, ranked per Bounded feedback, class tagged, each SBI-shaped with quote
The one thing : single named focus behaviour, observable and re-executable
Next-review check : what observable change on the next call counts as success
```
## Quality gate
Score the drafted review against all ten before delivering. Pass threshold: 10/10. Iterate until nothing fails.
1. Every scored dimension cites at least one verbatim quote with a findable location; every unscored dimension says N/A-with-reason or "insufficient evidence".
2. The transcript quality tier is declared before any score, and no dimension rests on a single garbled or ambiguous line above Low confidence.
3. N/A dimensions are out of every aggregate; nothing structurally absent is scored zero.
4. Feedback is bounded: at most 2 strengths (each quoted), at most 3 change items, exactly one focus behaviour phrased as an observable action. Each change item carries its class, the items follow the Bounded feedback ranking, and any deviation names the Interview answer or rep context that moved them.
5. Every change item is SBI-shaped; no personality or character language anywhere in the review.
6. Any fatal moment is named above the scores, however good the numbers look.
7. The call's outcome appears as context only and justifies no score.
8. In manager/peer mode, the rep's self-assessment was captured before scores were revealed, or its absence is explicitly noted.
9. Every statistic names where it comes from, and every vendor number is also marked correlational.
10. The review states the next-review check: the observable change that would count as the focus behaviour landing.
## KPIs and measurement
- The review worked if the named focus behaviour observably changed on the next reviewed call - compared at quote level, not by impression. That is the primary KPI; a review that produced no behaviour change was a document, not a review.
- Process KPIs per review:
- 100% of judgements quoted.
- Review consumable in under 30 minutes.
- The rep can restate the one thing unprompted at the end.
- Across reviews: the same behaviour remaining the focus for three consecutive reviews signals the feedback isn't landing - surface that signal to the human; deciding what to do about it (coaching plan, role-play cadence) is out of this skill's scope. Rep self-scores converging toward reviewer scores over time is the calibration-health signal.
Optional integration note: call recording tools, transcription/ASR services, and conversation-intelligence platforms are categories - any tool of the class supplies the transcript, and the skill works from a pasted transcript with none of them.
## Reference
- See [references/rubric-anchors.md](references/rubric-anchors.md) for the anchored 0/2/4 behavioural descriptors, per-dimension evidence cues, and the optional B2C compliance dimension.
- See [references/review-template.md](references/review-template.md) for the fill-in review template with B2C/contact-centre adaptations.
- See [references/worked-examples.md](references/worked-examples.md) for a worked review built on a published, sourced cold-call teardown, plus an annotated counter-example of a bad review.
- `mbfinotti/sales-skills@cold-call-opener` - build or rewrite an opener script; this skill grades the one on tape.
- `mbfinotti/sales-skills@sales-discovery-questions` - build a discovery question set; this skill judges the questions asked.
- `mbfinotti/sales-skills@sales-objection-handling` - write objection rebuttals; this skill scores how the live objection was handled.
- `mbfinotti/sales-skills@deal-red-flags` - review free-text deal notes for red flags, a different artifact than call transcripts.
- `mbfinotti/sales-skills@sales-hiring` - reuse this rubric to score new hires' live calls during ramp with evidence-quoted discipline.