SKILL DETAIL
sales-hiring
mbfinotti/sales-skills/sales-hiring
Employer-side hiring workflow for SDR/BDR and AE roles, producing an outcome-based scorecard, a structured interview loop with question bank, a scored mock-call work sample, and a 30-60-90 ramp plan with certification gates. Recommends SDR vs AE vs full-cycle from ACV, cycle length and inbound volume, and rests on selection-validity evidence - structured interviews, independent scoring, no brainteasers. Use whenever the user mentions hiring a rep, sales interview questions, a hiring scorecard, a mock call interview, rep onboarding, or a ramp plan, even without the word hiring. Do NOT use for candidate-side prep (mbfinotti/sales-skills@sales-career) or comp plan design (mbfinotti/sales-skills@sales-comp-design).
Installation
npx skills add https://github.com/mbfinotti/sales-skills --skill sales-hiring
Skill-Dateien
SKILL.md
Zuletzt synchronisiert · 15.09.2026
evals/evals.json›
{
"skill_name": "sales-hiring",
"evals": [
{
"id": 1,
"prompt": "I run sales at Kettleform, a workflow tool at about $9K ACV. Our sales cycle is roughly five weeks and we get around 180 inbound demo requests a month - my three AEs are booked solid and pushing first calls out two weeks. The board wants us to look like a real sales org, so I've got budget for four SDRs starting in January. Can you put together what I need to hire them - scorecard, interview questions, ramp plan?",
"expected_output": "A challenge to the four-SDR plan before any artifact is drafted, grounded in the $9K ACV, five-week cycle and inbound volume, with the SDR + AE split removed from the options while AE calendars stay full, a named condition that would restore it, and the remaining intake questions asked one at a time.",
"files": [],
"expectations": [
"Challenges the four-SDR request before producing any artifact, rather than drafting SDR scorecards as asked",
"Cites the $9K ACV, the five-week cycle and the inbound volume explicitly as the basis of the role recommendation",
"States that the SDR + AE split is removed from the options entirely, not merely ranked lowest, while inbound keeps closers' calendars full",
"Names the condition that would put the SDR + AE split back on the table: AEs showing empty-calendar symptoms, or a company commitment to outbound as a strategy",
"Recommends adding closing capacity (full-cycle or AE) rather than SDRs, given the booked-solid AE calendars",
"States that below roughly $25K ACV on cycles under three months with inbound flow, the full-cycle rep leads on efficiency",
"Warns that adding headcount requires scaling each demand-generation source in proportion or the new hires starve",
"Labels the evidence tier behind the role recommendation rather than presenting it as unsourced fact",
"Asks the remaining intake questions, including hiring jurisdiction and comp philosophy, before producing artifacts",
"Asks intake questions one at a time rather than delivering a full questionnaire in a single message",
"States the role reasoning in a short paragraph the user is invited to veto rather than presenting it as settled",
"Does not deliver the scorecard, question bank, work sample and ramp plan all at once in the first reply"
]
},
{
"id": 2,
"prompt": "We're hiring a mid-market AE at Brantlow Systems - about $40K ACV, four-month cycles, outbound-led, US only, we're at eleven reps already, no hard deadline, and this is the first of maybe six AE hires over the next 18 months. Here's the loop I drafted: (1) recruiter screen 30 min, (2) hiring manager 45 min, (3) VP Sales 45 min, (4) panel with two AEs and a CS lead, 60 min, (5) founder chat 30 min, (6) mock call with the whole panel watching, final round before offer. Does this look right?",
"expected_output": "A revised loop with the mock call moved to stage 2 or 3, the late placement removed rather than ranked, the value tie and the hours difference stated explicitly, panel size corrected, competencies assigned per interviewer, and the six-hire horizon treated as compounding.",
"files": [],
"expectations": [
"Moves the mock call from the final round to stage 2 or 3 of the loop",
"States that the last placement is deleted from the options rather than ranked below stage 2-3",
"States that the two placements are equal on value because it is the same instrument, the same rubric and the same score wherever it sits",
"Names interviewer hours as the only axis separating the placements, quantifying roughly 10 or more hours burned per candidate who cannot sell",
"Names the extra days of time-to-fill, and candidates lost to competing offers, as a second cost of the late placement",
"States that a panel adds no measurement validity over a single trained interviewer, and that a panel's value is governance rather than measurement",
"Assigns each interviewer one named scorecard competency instead of having every interviewer assess the candidate overall",
"Requires the identical question set, in the same order, for every candidate for this role",
"Keeps the loop within 4-6 stages for a mid-market AE",
"Treats the six-hire horizon as compounding, which justifies building the question bank and the frozen rubric properly rather than thinly",
"Requires interviewers to submit scores independently before any debrief opens",
"Does not recommend adding interviewers or widening the panel as a way to raise selection accuracy"
]
},
{
"id": 3,
"prompt": "Sanity-check my SDR interview question list for Velmarch Analytics before I send it to the panel. 1) Sell me this pen. 2) What's your greatest weakness? 3) Are you coachable? 4) How many ping-pong balls fit in a school bus? 5) Where do you see yourself in five years? 6) Tell me about yourself. Then whoever's free that week runs the interview and we all chat afterwards.",
"expected_output": "All four banned questions removed with the evidence behind the removal, the coachability question replaced by a behavioral test, the bank rebuilt per competency with behavioral and situational questions and probes, and the ad hoc interviewer assignment replaced with fixed interviewers and a fixed question set.",
"files": [],
"expectations": [
"Removes \"sell me this pen\" from the list",
"Removes the \"greatest weakness\" question",
"Removes \"are you coachable\" as pure self-report with no validity",
"Removes the ping-pong-ball brainteaser",
"States that a large employer's own analysis of tens of thousands of its interviews found no relationship between brainteaser performance and job performance",
"Replaces the coachability question with a behavioral test: re-running a segment of the mock call after feedback and scoring the delta",
"Rejects \"whoever's free that week\" and fixes the same interviewers and the same questions across every candidate for the role",
"Rebuilds the bank as 2-3 past-behavior plus 1-2 situational questions per scorecard competency, each carrying drill-down probes",
"Maps every replacement question to a named scorecard competency",
"States roughly how little performance variance unstructured interviews explain (about 14%, and less after modern corrections)",
"Does not defend any of the four banned questions as still useful for rapport, culture fit or pressure-testing"
]
},
{
"id": 4,
"prompt": "Our COO at Harrowgate Labs wants every AE candidate to take the Objective Management Group sales assessment - their rep told him it has 95% predictive validity - and he wants it as a hard gate before anyone reaches the hiring manager. We also had Culture Index pitched to us last month, they quoted reliability figures in the .86 to .96 range. We hire in Austin and Munich. What should I tell him?",
"expected_output": "A refusal to let any assessment gate or veto, the 95% and reliability claims characterised as vendor marketing rather than psychometrics, Buros MMY named as the independence test, Hogan and a validated cognitive test named as the evidenced alternatives, and the Munich compliance duties flagged.",
"files": [],
"expectations": [
"Refuses to let the assessment gate or veto any candidate, at any tier",
"States that a vendor validity claim above roughly .55 is a marketing claim, not a psychometric one",
"States that genuine criterion validities rarely exceed about .50 even for the best predictors",
"Notes that the 95% figure is not a standard psychometric metric and has no independent replication outside the vendor's own materials",
"States that the vendor tool does not appear in the Buros Mental Measurements Yearbook and has no published peer-reviewed criterion validity",
"Applies the same unvalidated verdict to Culture Index, and states that reliability measures consistency rather than predictive validity",
"Names Hogan and a validated cognitive test such as Wonderlic as the options here carrying independent peer-reviewed evidence",
"Flags that deploying an assessment on the Munich role typically requires works-council approval first",
"Flags the high-risk AI obligations attaching to selection tools used for EU-based roles",
"States that a timed assessment must accommodate disability",
"Permits unvalidated vendor output at most as a conversation prompt, never as a score inside the mechanical combination"
]
},
{
"id": 5,
"prompt": "I read that general mental ability is the single best predictor of job performance, so at Pellamy Group I want a 12-minute cognitive test at the very top of the AE funnel to cut our 400-applicant pile down fast, and I want it weighted heaviest in the final decision. Walk me through setting that up.",
"expected_output": "The sales-specific split between supervisor-rating validity and objective-sales validity stated with figures, achievement/conscientiousness named as the better-weighted signal, the heaviest weighting refused, the validation commitment and compliance load spelled out, and the structured interview placed as the highest-value instrument.",
"files": [],
"expectations": [
"Reports that cognitive ability correlates about .40 with supervisor ratings but only about .04 with objective sales results",
"States that the test predicts what managers think of reps rather than what reps sell",
"Names achievement/conscientiousness as the inverse case, around .41 against objective sales, and the more trustworthy signal to weight",
"Refuses to make the cognitive test the heaviest-weighted input in the final decision",
"States that adding a cognitive test requires a commitment to validate it against objective sales outcomes, roughly a quarter of work on data most teams do not have",
"Flags the cognitive test as carrying the heaviest compliance load of the selection instruments under discussion",
"Uses the corrected operational validity figures and notes that earlier published figures were overstated",
"Directs the user to validate any gate against objective sales outcomes rather than against performance reviews",
"Places the structured interview, not the cognitive test, as the highest-value instrument in the loop"
]
},
{
"id": 6,
"prompt": "Our AE debrief at Sondermere works like this: everyone who met the candidate gets in a room on Friday, we go around the table, talk it through, and land on a thumbs up or thumbs down. My VP usually goes first since she's the most experienced person in the room. The last two hires interviewed brilliantly and then missed quota badly. What would you change?",
"expected_output": "Independent evidence-quoted scores submitted before the debrief, per-competency scores combined mechanically against pre-set weights with the mechanical-versus-holistic figures, the VP-first order named as the anchoring failure, a multi-tier recommendation replacing the binary, and the interview-well/miss-quota pattern read as the trigger to build the work sample first.",
"files": [],
"expectations": [
"Requires every interviewer to submit written scores before the debrief opens",
"Requires one verbatim evidence quote from the interview behind each score",
"Replaces the blended verdict with per-competency scores combined mechanically using pre-set weights",
"Gives the comparison figures for mechanical versus holistic combination (about .44 versus .28)",
"States that holistic judgment burns up to roughly half the validity",
"Names anchoring by the most senior or loudest voice as the specific failure the VP-goes-first order produces",
"Replaces the binary thumbs up/thumbs down with a multi-tier recommendation",
"States that the debrief calibrates and interrogates scores but never originates them",
"Sets decision thresholds before the loop opens, with the hiring manager deciding within them",
"Reads \"interviewed brilliantly, missed quota\" as the condition that promotes the work sample to first build",
"Requires debrief comments in Situation-Behavior-Impact form rather than trait labels such as \"not a closer\""
]
},
{
"id": 7,
"prompt": "New AE starts at Corvantis on the 6th - we're at $50K ACV with five-month cycles. The plan: day one she gets her laptop, the full product deck, CRM login, her territory list and the same $900K annual quota everyone else carries, then we review at 90 days and see if she's working out. Her manager will double as her onboarding buddy. Write me the onboarding plan.",
"expected_output": "A 30-60-90 plan with a day-30 certification gate blocking pipeline ownership, day-60 coached activity and day-90 self-sourced pipeline milestones, a stepped quota replacing the flat $900K, a non-recoverable draw, a peer buddy separate from the manager, week-one spread orientation, SBI check-ins, the missing-all-three action rule, and the long-cycle inspection extension.",
"files": [],
"expectations": [
"Structures the plan as 30-60-90 phases rather than a day-one handoff",
"Places a product/ICP certification exam with a pass bar at day 30 and blocks live pipeline ownership before it is passed",
"Sets a day-60 milestone of live coached activity with calls recorded and reviewed",
"Sets a day-90 milestone of self-sourced pipeline plus first deals or forecast accuracy",
"Replaces the flat $900K from month one with a stepped quota ramp of 25/50/75/100% across successive quarters",
"Adds a non-recoverable draw for the AE across months 1-3",
"Assigns a peer buddy distinct from the manager, correcting the manager-as-buddy plan",
"Spreads orientation across week 1 with deep work starting week 2, rather than loading day 1",
"Specifies check-ins in Situation-Behavior-Impact form",
"States the action rule: missing all three 90-day milestones triggers a decision - a dated remediation plan or exit - not continued hope",
"Quotes an AE ramp benchmark near 6.2 months with its evidence tier labeled rather than stated as bare fact",
"Flags a 60-90-180 or 180-day inspection extension because a five-month cycle puts the first closeable deal after day 90"
]
},
{
"id": 8,
"prompt": "Vermilya Health is opening AE reqs in New York City, Denver and Berlin this quarter. To handle volume we bought an AI video-interview tool that scores candidates on sales aptitude and auto-rejects the bottom 60% before a human sees them. Our application form asks for current base and commission so we can put together a fair offer. The postings say \"competitive OTE, DOE.\" Anything I should tighten up?",
"expected_output": "Salary-history question deleted and real ranges published, LL144 bias-audit and notice duties flagged for NYC, Colorado and EU high-risk AI obligations flagged, automated rejection flagged against the human-involvement requirement, German works-council approval flagged, and dropping the tool named as the single move that removes four duties at once.",
"files": [],
"expectations": [
"Removes the current base and commission question from the application form",
"Replaces \"competitive OTE, DOE\" with a published base and OTE range",
"Flags the NYC bias-audit and candidate-notice obligations for automated employment decision tools",
"Names the NYC penalty exposure of up to $1,500 per day, or the impact-ratio threshold below 0.80 that turns a failed audit into evidence against the employer",
"Flags Colorado's high-risk AI employment obligations for the Denver reqs",
"Flags the EU classification of recruitment and selection AI as high-risk for the Berlin reqs",
"States that the EU obligations attach to AI used for EU-based roles regardless of where the employer is located",
"Flags that auto-rejecting candidates without meaningful human involvement conflicts with the right not to be subject to solely automated decisions",
"Flags German works-council approval as typically required before the tool is deployed",
"Names dropping the AI screening tool as the single decision that removes the bias audit, the human-review duty, the high-risk classification and the works-council question together",
"States that publishing ranges and deleting the salary-history question is the near-zero-effort move satisfying the largest number of jurisdictions at once",
"States that structured interviews and scored work samples are the legally defensible instruments under adverse-impact challenge, via content validity"
]
},
{
"id": 9,
"prompt": "First sales hire ever at Thistlebourne. I'm the founder and I still do all the selling myself - $30K ACV, four-month cycles, mostly outbound. I've got about three weeks of my own time before I want candidates in a loop, no recruiter, nobody else here to interview anyone. What do I build first, and is there anything cheap I can bolt on to make the decision more objective?",
"expected_output": "The work sample promoted to first build because a first hire has no incumbent benchmark, with that promotion named as an override of the default efficiency order, biodata deleted rather than ranked, the job-knowledge exam pushed into the ramp, reference checks scoped to finalists, the ordering axes kept distinct, and the founder-hour re-rank applied.",
"files": [],
"expectations": [
"Promotes the work sample to first build because a first sales hire has no incumbent benchmark to calibrate an interview against",
"States explicitly that this promotion overrides the default efficiency order, which otherwise buries the work sample",
"Names why the efficiency order starves the work sample: priciest to build, the only one needing continuous upkeep, the only one spending the candidate's hours",
"Keeps the validity ordering and the efficiency ordering distinct and does not quote one as if it were the other",
"Rules out biodata / empirical keying because it needs outcome data on hundreds of past hires this team does not have",
"States that biodata is deleted from the options rather than ranked last",
"Places the job-knowledge exam in the ramp as the day-30 certification gate rather than as a pre-hire instrument",
"Scopes reference checks to finalists only and weights them low as a fraud and claim check rather than a stage that outvotes the work sample",
"Re-ranks against the no-recruiter answer, noting every hour in the instrument table becomes the founder's own hour",
"Applies the first-hire profile - coachability, curiosity, intelligence, work ethic - and treats heavy big-company experience as a risk to probe rather than a credential"
]
},
{
"id": 10,
"prompt": "Draft the job scorecard for our second AE at Ovanth Retail Systems - $60K ACV, enterprise-ish buyers, five-month cycles, US. What I want is someone who can do a bit of everything: prospect, close, handle renewals, help marketing with content, mentor the SDRs. Ideally out of Salesforce or Oracle with 10+ years. Must-haves: MEDDIC, Salesforce, Gong, Clari, Outreach, Sales Navigator, SaaS background, enterprise background, strong closer mentality. Responsibilities: own the territory, build pipeline, hit the number, be a team player.",
"expected_output": "A scorecard with a one-sentence mission, 3-8 ranked and quantified outcomes replacing the vague responsibilities, 5-7 weighted competencies on enterprise weighting, the all-around-athlete profile rejected, must-haves cut to 5-7, acceptance defined as bands, red flags listed with their failure modes, and the dual reuse of the scorecard stated.",
"files": [],
"expectations": [
"Produces a one-sentence mission for the role",
"Replaces the vague responsibilities list with 3-8 measurable outcomes, each quantified with a number and a deadline",
"Ranks the outcomes by importance",
"Defines 5-7 competencies, each carrying an explicit weight",
"Rejects the do-a-bit-of-everything profile and scopes the role to narrow, deep competence against the outcomes",
"Cuts the must-have list to at most 5-7 items and states that more shrinks the funnel without adding signal",
"Weights the hire toward the competency the current team lacks rather than cloning the strongest incumbent",
"Applies enterprise weighting: discovery depth, business acumen and stakeholder navigation highest, raw activity volume lower",
"States that the scorecard doubles as the interview rubric and as the ramp milestone map",
"Defines acceptance as bands - must have, strongly prefer, acceptable, hard no - rather than a single pass mark",
"Lists role-specific red flags with the failure mode each predicts, and states a red flag triggers debrief discussion rather than automatic rejection"
]
},
{
"id": 11,
"prompt": "Designing the work sample for SDR candidates at Kirnwell Data. My plan: the candidate does a 10-minute cold call pitching whatever they sold at their last job, so they're comfortable and we see their best. Panel of five watches, then we score them 1-10 on overall impression over coffee afterwards. I also want to use the rule that top reps talk 43% of the time, so anyone over that loses points.",
"expected_output": "The candidate switched to selling the hiring company's product with a prep pack and a hard scenario, an unscripted objection injected, a scored coachability re-run added, the 1-10 impression score replaced by a frozen behaviorally-anchored 0-3 rubric, the panel cut to two roles, and the 43% talk-share figure refused as a criterion and labeled as conflicting vendor data.",
"files": [],
"expectations": [
"Changes the work sample so the candidate sells the hiring company's product rather than the one from their last job",
"States the reason: it tests adaptability to an unfamiliar pitch instead of a rehearsed one",
"Sends a prep pack ahead of the call - example prospect, company one-pager, deck or demo video",
"Makes the live scenario genuinely hard rather than comfortable, on the grounds that a soft scenario measures nothing",
"Injects one unscripted objection mid-call",
"Adds a re-run of one segment after feedback and scores the delta as the coachability test",
"Replaces the 1-10 overall impression score with a behaviorally-anchored 0-3 rubric across 5-8 categories, writing what a 0 and a 3 sound like",
"Freezes the rubric for a quarter and requires frame-of-reference calibration for scorers before live use",
"Cuts the five-person panel to two roles: one playing the prospect, one scoring silently",
"Refuses to make the 43% talk-share figure a scoring criterion, labels it vendor data and correlational, and notes it conflicts with the same vendor's later figure near 57%",
"Scores improvisation and adaptability rather than polish, with both scorers submitting independent evidence-quoted scores before comparing notes"
]
},
{
"id": 12,
"prompt": "Hiring 15 inside-sales reps at Lumaflex Home Solar to sell residential systems over the phone - commission-heavy, high turnover expected. I'm copying the loop and ramp we used at my last B2B SaaS company: five stages over three weeks, six-month ramp, 3.5x pipeline coverage targets, base-heavy package. What needs to change?",
"expected_output": "A 2-4 stage loop on days, the SaaS benchmarks relabeled as directional rather than baseline, a commission-heavy package with a draw and budgeted backfill, B2C predictor weighting with compliance discipline, an end-to-end consumer call closing for the sale, and the shared rigor foundation held unchanged.",
"files": [],
"expectations": [
"Shortens the loop to 2-4 stages over days, on the grounds that speed wins candidates in this segment",
"States that the published ramp, tenure and attainment benchmarks skew North American B2B SaaS and are directional here, not this user's baseline",
"Replaces the base-heavy package with a commission-heavy structure plus a draw",
"Budgets for backfill against the higher expected washout",
"Weights rejection tolerance, drive and activity resilience highest, and multi-threading and long-cycle forecasting lowest",
"Adds compliance discipline to the weighting and required-disclosure checkpoints to the mock call for the regulated product",
"Designs the work sample as an end-to-end consumer call closing for the sale rather than for a meeting",
"Keeps the written scorecard, structured interviews, scored work sample, independent scoring and mechanical combination unchanged despite the shorter loop",
"States explicitly that a shorter loop earns no reduction in rigor"
]
},
{
"id": 13,
"prompt": "I've got a final-round AE interview at a Series B fintech on Thursday and they said there'd be a mock discovery call. Can you build me a scorecard for myself and run me through the questions they're likely to ask, so I can prep answers?",
"expected_output": "Recognition that the requester is the candidate rather than the hiring manager, a stop before any prep material is produced, and a pointer to the candidate-side sibling skill.",
"files": [],
"expectations": [
"Identifies the requester as a candidate rather than a hiring manager",
"Stops rather than producing the requested interview prep, likely-question list or candidate-side scorecard",
"Points the user to the candidate-side sibling skill by name",
"Does not produce a hiring-side scorecard, question bank or ramp plan for this request either",
"Does not coach the candidate on how to perform in the mock discovery call"
]
}
],
"trigger_queries": [
{ "query": "help me hire our first sales rep", "should_trigger": true },
{ "query": "I need an interview scorecard for an SDR role", "should_trigger": true },
{ "query": "design the interview loop for a mid-market AE", "should_trigger": true },
{ "query": "what questions should I ask a BDR candidate", "should_trigger": true },
{ "query": "build a 30-60-90 plan for a new account executive", "should_trigger": true },
{ "query": "our new rep starts Monday and I have nothing written down", "should_trigger": true },
{ "query": "should I hire SDRs or just more closers", "should_trigger": true },
{ "query": "how do I run a mock call as part of an interview", "should_trigger": true },
{ "query": "write a ramp plan for an inside sales hire", "should_trigger": true },
{ "query": "we keep hiring reps who interview great and then miss quota", "should_trigger": true },
{ "query": "set up a structured interview process for sales candidates", "should_trigger": true },
{ "query": "what should the day 30 milestone be for a new AE", "should_trigger": true },
{ "query": "I want to add a sales personality test to our hiring process", "should_trigger": true },
{ "query": "is the OMG assessment worth using on sales candidates", "should_trigger": true },
{ "query": "how many interview stages for an enterprise AE", "should_trigger": true },
{ "query": "our debrief is just everyone chatting, how do we score candidates properly", "should_trigger": true },
{ "query": "help me write the outcomes section of a rep job scorecard", "should_trigger": true },
{ "query": "what does a good sales work sample look like", "should_trigger": true },
{ "query": "onboarding plan for a new SDR cohort", "should_trigger": true },
{ "query": "we're opening six AE reqs next quarter, what do I build first", "should_trigger": true },
{ "query": "do I need a bias audit for the AI tool screening our sales candidates", "should_trigger": true },
{ "query": "can I ask sales candidates what they currently earn", "should_trigger": true },
{ "query": "what should go in a rep hiring scorecard", "should_trigger": true },
{ "query": "designing a certification exam for new reps before they touch pipeline", "should_trigger": true },
{ "query": "how long should a new AE ramp before carrying a full number", "should_trigger": true },
{ "query": "first sales hire after founder-led selling - what profile do I look for", "should_trigger": true },
{ "query": "should the mock call be the final round", "should_trigger": true },
{ "query": "our SDR attrition is brutal, is that the hiring loop or the onboarding", "should_trigger": true },
{ "query": "help me pick between a full-cycle rep and an SDR plus AE split", "should_trigger": true },
{ "query": "interview questions that actually predict sales performance", "should_trigger": true },
{ "query": "we're scaling from 3 reps to 10 next year, what process do I need", "should_trigger": true },
{ "query": "how do I test whether a candidate is actually coachable", "should_trigger": true },
{ "query": "grading rubric for a candidate's cold call", "should_trigger": true },
{ "query": "who should be in the room for a sales interview", "should_trigger": true },
{ "query": "my VP dominates every debrief and we end up hiring whoever she liked", "should_trigger": true },
{ "query": "what do I do when a new rep misses every 90 day milestone", "should_trigger": true },
{ "query": "we need a repeatable way to evaluate AE candidates", "should_trigger": true },
{ "query": "planning the first 90 days for a new enterprise seller", "should_trigger": true },
{ "query": "should we use a cognitive test to screen 400 sales applicants", "should_trigger": true },
{ "query": "how do I stop hiring the all-around athlete", "should_trigger": true },
{ "query": "build the competency weights for an outbound SDR role", "should_trigger": true },
{ "query": "every time we double the team the new people starve for leads", "should_trigger": true },
{ "query": "what red flags should I watch for in an AE interview", "should_trigger": true },
{ "query": "put together the evaluation pack for our next revenue hire", "should_trigger": true },
{ "query": "we need someone selling by Q2, what's the fastest defensible process", "should_trigger": true },
{ "query": "how do we make sure every candidate gets the same interview", "should_trigger": true },
{ "query": "the candidate couldn't name a single deal they lost - how much should that matter", "should_trigger": true },
{ "query": "how much weight should reference checks carry for a seller", "should_trigger": true },
{ "query": "designing a work sample where the candidate has to sell our product", "should_trigger": true },
{ "query": "our job post says competitive OTE, is that a problem", "should_trigger": true },
{ "query": "hiring reps in New York and Berlin, what compliance do I need to worry about", "should_trigger": true },
{ "query": "is a 100 question product exam before live accounts overkill", "should_trigger": true },
{ "query": "help me structure the first three months for a commission-only phone rep", "should_trigger": true },
{ "query": "what's a reasonable time to fill for an enterprise AE req", "should_trigger": true },
{ "query": "should each interviewer cover a different competency", "should_trigger": true },
{ "query": "we want to hire more people like our top rep, how do we profile that", "should_trigger": true },
{ "query": "draft the must-have versus nice-to-have list for a sales req", "should_trigger": true },
{ "query": "can I use Hogan on sales candidates", "should_trigger": true },
{ "query": "our mock call is too easy and everybody passes it", "should_trigger": true },
{ "query": "what does a stepped quota ramp look like for a new hire", "should_trigger": true },
{ "query": "I'm the founder and I've never interviewed a salesperson before", "should_trigger": true },
{ "query": "how do I choose between two AE finalists without just going on gut feel", "should_trigger": true },
{ "query": "prep me for my SDR interview next week", "should_trigger": false },
{ "query": "how do I get promoted from SDR to AE", "should_trigger": false },
{ "query": "review my sales resume", "should_trigger": false },
{ "query": "evaluate this AE offer I just received", "should_trigger": false },
{ "query": "design the commission plan for our new AEs", "should_trigger": false },
{ "query": "what OTE should we offer a mid-market AE", "should_trigger": false },
{ "query": "set quotas for next year's AE team", "should_trigger": false },
{ "query": "how much ramp relief should a new rep get on their quota", "should_trigger": false },
{ "query": "should we run pods or an assembly line sales org", "should_trigger": false },
{ "query": "what's the right SDR to AE ratio for our team", "should_trigger": false },
{ "query": "how many reps per sales manager", "should_trigger": false },
{ "query": "score this call transcript from my rep yesterday", "should_trigger": false },
{ "query": "how did this discovery call go", "should_trigger": false },
{ "query": "write a cold call opener for a CTO persona", "should_trigger": false },
{ "query": "our cold emails are all landing in spam", "should_trigger": false },
{ "query": "test these three subject lines for me", "should_trigger": false },
{ "query": "build an 11 touch outbound cadence", "should_trigger": false },
{ "query": "handle the we already use a competitor objection", "should_trigger": false },
{ "query": "what discovery questions should I ask on Thursday's call", "should_trigger": false },
{ "query": "score this deal against MEDDPICC", "should_trigger": false },
{ "query": "who's the economic buyer on this deal", "should_trigger": false },
{ "query": "is this deal real or is it going to slip again", "should_trigger": false },
{ "query": "build the ROI case for this opportunity", "should_trigger": false },
{ "query": "the buyer wants 25% off, what do I trade for it", "should_trigger": false },
{ "query": "write the recap email from this morning's meeting", "should_trigger": false },
{ "query": "define our ideal customer profile", "should_trigger": false },
{ "query": "what's the TAM for our market", "should_trigger": false },
{ "query": "build the account fit score from our closed-won data", "should_trigger": false },
{ "query": "where should the tier cutoffs sit for our key accounts", "should_trigger": false },
{ "query": "how much pipeline coverage do we need to hit the number", "should_trigger": false },
{ "query": "should we go product-led or sales-led", "should_trigger": false },
{ "query": "which sales podcasts and newsletters should I follow", "should_trigger": false },
{ "query": "which sales skill do I need for this project", "should_trigger": false },
{ "query": "find me a personalization angle for this prospect", "should_trigger": false },
{ "query": "clean up our pipeline hygiene, half the stages are wrong", "should_trigger": false },
{ "query": "our forecast keeps missing, help me diagnose it", "should_trigger": false },
{ "query": "design the handoff from sales to customer success", "should_trigger": false },
{ "query": "improve our user onboarding flow so more signups activate", "should_trigger": false },
{ "query": "write the customer onboarding checklist for a new enterprise account", "should_trigger": false },
{ "query": "help me hire a senior backend engineer", "should_trigger": false },
{ "query": "design the interview loop for a product manager role", "should_trigger": false },
{ "query": "write a job description for a marketing manager", "should_trigger": false },
{ "query": "source candidates on LinkedIn for our open AE role", "should_trigger": false },
{ "query": "set up an applicant tracking system for our team", "should_trigger": false },
{ "query": "optimize my resume for ATS keyword scanning", "should_trigger": false },
{ "query": "run a user research interview with our customers", "should_trigger": false },
{ "query": "build a marketing scorecard for our campaigns", "should_trigger": false },
{ "query": "what should I ask in an exit interview", "should_trigger": false },
{ "query": "design a performance review process for the sales team", "should_trigger": false },
{ "query": "write a PIP for an underperforming rep", "should_trigger": false },
{ "query": "negotiate my salary for the AE offer I just got", "should_trigger": false },
{ "query": "write my 30-60-90 day plan to present in my final interview", "should_trigger": false },
{ "query": "practice answers for sell me this pen", "should_trigger": false },
{ "query": "build a recruiting pipeline dashboard for our talent team", "should_trigger": false },
{ "query": "write the sales playbook for our existing reps", "should_trigger": false },
{ "query": "onboard a new customer onto our platform", "should_trigger": false },
{ "query": "plan the agenda for our sales kickoff offsite", "should_trigger": false },
{ "query": "how do I coach a tenured rep out of a slump", "should_trigger": false },
{ "query": "write LinkedIn outreach to passive candidates for our open AE role", "should_trigger": false },
{ "query": "post our AE job ad to the usual boards", "should_trigger": false },
{ "query": "design the SDR to AE promotion criteria", "should_trigger": false },
{ "query": "run a compensation benchmarking study across our go-to-market org", "should_trigger": false }
]
}
references/assessment-validity-audit.md›
# Commercial sales-assessment validity audit
The substance behind the SKILL.md assessment rule: treat any commercial sales-personality assessment as unvalidated unless it appears in the Buros Mental Measurements Yearbook (MMY) - the standard independent test-review index - with published peer-reviewed criterion validity. Never let any assessment gate or veto a candidate.
| Assessment | Headline claim | Evidence status | In Buros MMY? | Peer-reviewed criterion validity? |
| -------------------------------- | ------------------------------ | ---------------------------------------------------------------------------------------------- | --------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| Objective Management Group (OMG) | "95% predictive validity" | Vendor self-reported only; unreplicated | No | None found |
| Culture Index | Reliability α=.86-.96 | Vendor self-reported - and reliability is consistency, not predictive validity | No | None found |
| Caliper Profile | Validity r=.29-.39 | Vendor-reported; plausible magnitude, not independently verified | No | Not independently verified |
| Predictive Index | 400+ studies; EFPA-certified | Mostly vendor; some external audit (EFPA/DNV-GL); the 400+ studies are internal client studies | No | Not in peer-reviewed journals |
| Hogan (HPI) | True validities .25-.43 | Independent and peer-reviewed | Yes (13th MMY onward) | Yes - Hogan & Holland 2003, _JAP_ 88(1); sales-specific validation in Hogan, Hogan & Gregory 1992, _J. Business and Psychology_ |
| Wonderlic | General cognitive ability test | Independent and peer-reviewed | Yes (14th/17th MMY) | Yes (GMA r≈.31 corrected) |
## Why OMG's "95%" is a red flag, not a coefficient
It is not a standard psychometric metric: genuine criterion validities rarely exceed ~.50 even for the best predictors in the peer-reviewed literature. The figure appears only in OMG's own materials and reseller pages and has zero independent replication. Generalize the lesson: any vendor quoting a validity above ~.55 is making a marketing claim.
## If the user insists on an assessment
Ranked, with gating deleted from the menu rather than ranked last - the SKILL.md assessment rule forbids it outright, at every tier. Read `>` as "more of this axis".
- efficiency: Hogan > validated cognitive test > unvalidated vendor tools
- effort: Hogan == validated cognitive test == unvalidated vendor tools (a genuine three-way tie: all are bought off the shelf, all take under an hour of candidate time, none needs an instrument built - which is exactly the trap, since near-zero effort makes the worthless one feel free)
- compliance cost: unvalidated vendor tools > validated cognitive test > Hogan
Effort separates none of them, so validity and compliance cost decide alone:
- Use one with real evidence (Hogan; a validated cognitive test such as Wonderlic) - and remember the sales caveat on cognitive tests: GMA predicts supervisor ratings (.40) but objective sales at only .04 (peer-reviewed), so validate any cognitive gate against actual sales outcomes before trusting it.
- Treat OMG / Culture Index / Caliper / Predictive Index output as unvalidated color: a conversation prompt at most, never a score in the mechanical combination, never a gate, never a veto. It is also the tier that triggers the compliance duties in the last bullet below without any validity to justify them.
- Traits are not skills: no style inventory substitutes for the scored mock call.
- **Compliance:**
- assessments used on EU-based roles may fall under high-risk AI obligations
- Germany/France typically require works-council approval before deployment
- timed assessments must accommodate disability
See legal-landscape.md via SKILL.md.
references/interview-question-bank.md›
# Interview question bank and debrief scoring form
Evidence labels as defined in SKILL.md.
## Banned questions - delete on sight
- **Brainteasers** ("how many golf balls fit in a 747"). Google's analysis of tens of thousands of its own interviews found zero relationship between brainteaser scores and job performance; Laszlo Bock: "a complete waste of time... They serve primarily to make the interviewer feel smart" [measured internal analysis at scale, reported via NYT 2013].
- **"Sell me this pen."** Decontextualized improv under artificial pressure; no established predictive validity as a selection instrument, however well the candidate answers.
- **"What's your greatest weakness"** and unstructured rapport chat. Unstructured interviews explain roughly 14% of performance variance and less after modern corrections (peer-reviewed).
- **"Are you coachable?"** Pure self-report, no validity. Test coachability behaviorally: re-run part of the mock call after feedback (see mock-call-design.md).
## Question construction rules
- 2-3 past-behavior questions plus 1-2 situational questions per scorecard competency, with drill-down probes. Ask the identical set, in the same order, of every candidate for the role - structure is the validity lever (peer-reviewed).
- Drill for specifics: a strong candidate survives three or four "and then what happened?" layers without contradiction. Depth of drill-down discriminates more than the prompt itself.
- A strong story answer runs about 90 seconds to 2 minutes with a concrete result; rambling or a missing result is itself a signal.
## SDR/BDR bank (map to the SDR competencies)
- Activity discipline: "Walk me through a normal Tuesday in your last prospecting role - hour by hour. What did you actually complete?" Probe: how do you know those numbers are right?
- Resilience: "Tell me about the worst stretch of rejection you have had. What changed in your approach by the end of it?"
- Curiosity: "Pick a company you prospected recently. What did you learn about them before the first touch, and where did you find it?"
- Written communication (live sample): have the candidate write a cold email or a 3-touch sequence during the interview, to a prospect brief supplied with the scenario. Score it on the rubric, not on taste.
- Process hygiene: "A meeting you booked no-showed twice. Show me exactly what you would log and do next."
## AE bank (map to the AE competencies)
- Discovery depth: "Take a deal you won. What business problem did the buyer put money against - not the feature list - and how did you find it?" Probe until you hit either a real problem statement or a feature pitch.
- Loss analysis / self-awareness: "Tell me about a deal you lost that you should have won. What was your mistake, specifically?"
- Multi-threading: "Describe the deal with the most stakeholders you have managed. Who could have killed it, and what did you do about them?"
- Forecasting honesty: "Tell me about a time you pulled a deal out of your commit. What did it cost you and why did you do it?"
- Closing under complication: "Give me a specific example of a late-stage deal that hit a complication - legal, budget freeze, champion change. Walk me through each decision to the close (or the loss)."
- Recovery: "Give me a specific example of a deal everyone else had written off that you closed - step by step."
## Debrief scoring form
Each interviewer completes this independently and submits it before the debrief opens. The debrief calibrates and interrogates: it never originates scores.
- Per competency: score on the anchored scale, one verbatim evidence quote from the interview supporting the score, and a 3-way fit key - strong match / acceptable / mismatch. The 3-way key surfaces the sticking-point competency that a blended average hides.
- Write comments in Situation-Behavior-Impact form (specific context → observable action → effect), never trait language ("not a closer", "low energy").
- A manager-fit note: predicted friction or alignment between the candidate's working style and the would-be manager's coaching style (e.g. autonomy-seeking rep under a highly directive manager).
- Combination is mechanical: multiply competency scores by the scorecard weights, sum, and compare to the pre-committed threshold. The hiring manager decides within the thresholds set before the loop opened.
Final recommendation is 5-tier, not binary:
- **Proceed** - strong fit, high confidence
- **Proceed with note** - strong fit, flag items to verify in ramp
- **Proceed with awareness** - moderate fit, plan onboarding adjustments
- **Discuss** - weak fit or unresolved concerns
- **Pause** - red flag hit, weigh severity in debrief, never auto-reject
Carry a caveat block in every debrief output:
- interview signals are predictions, not facts
- low-confidence signals may be wrong
- no candidate matches a perfect ideal
## Reference checks
Weight them low: validity ~.26 or lower (peer-reviewed) - a tiebreaker and fraud check, never a stage that outvotes the work sample or structured interview. Verify claimed numbers and titles; do not expect performance insight.
references/legal-landscape.md›
# Legal landscape for sales hiring (2026 status)
Informational only, not legal advice. Status is in flux - EU transposition is incomplete and country-specific, and AI-act guidance was still being negotiated in 2026. Confirm current status with counsel before acting.
| Jurisdiction / law | What it covers | 2026 status |
| ------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| US Uniform Guidelines / four-fifths rule | Adverse impact; validation of selection procedures | In force; structured interviews and work samples are defensible via content validity |
| US salary-history bans | ~20+ states/localities ban asking prior pay | In force, state-by-state patchwork |
| US pay transparency (CO, CA, NY, WA, IL, others) | Salary ranges required in postings | In force |
| NYC Local Law 144 | Bias audits for automated employment decision tools; annual independent audit; candidate notice | In force since 2023; enforcement intensifying after a critical Dec 2025 Comptroller audit; penalties up to $1,500/day; a bad audit (impact ratio <0.80) is evidence under Title VII/NYCHRL |
| Illinois AI Video Interview Act + HB 3773 | Consent/transparency for AI video analysis; broader AI-in-employment duties | In force since 2020; HB 3773 effective Jan 2026 |
| Colorado AI Act (SB 205) | High-risk AI including employment; algorithmic-discrimination duties | Effective Feb 2026 |
| EU AI Act | Recruitment/selection AI classified high-risk (Annex III); conformity assessment, human oversight, documentation | High-risk obligations apply from 2 Aug 2026; applies to AI used for EU-based roles regardless of employer location |
| EU Pay Transparency Directive (2023/970) | Salary ranges pre-interview; ban on salary-history questions; gender pay-gap reporting | Transposition deadline 7 June 2026 passed with only Italy, Slovakia, Lithuania, Malta fully in force; Germany, France, Spain and others missed it; some provisions may have direct effect; pay-gap reporting for 250+ employers from June 2027 on 2026 data |
| UK/EU GDPR Art. 22 | Right not to be subject to solely automated decisions | In force; requires meaningful human involvement in AI screening |
| UK Equality Act 2010 | Discrimination/adverse impact | In force |
| Germany/France works councils | Co-determination over assessment tools | Approval often required before deploying any assessment |
| ADA / accommodation (US) | Timed assessments must accommodate disability | In force; LL144 requires telling candidates they may request an alternative process or accommodation - the ADA may require providing one |
## Compliance moves, ranked by exposure removed per hour spent
Read `>` as "more of this axis"; the efficiency line is the order to work in.
- efficiency: publish ranges and drop salary-history > structured interview and scored work sample > human review of automated rejections > independent bias audit > works-council submission
- effort: works-council submission (a quarter, and the council can refuse) > independent bias audit (a quarter, external auditor, annual thereafter) > human review of rejections (a standing job) > structured instruments (a week, and Artifacts 1-3 already build them) > publish ranges and delete salary-history questions (near-zero)
**The move that deletes three rows:** declining to adopt an AI screening tool at all removes, in one decision:
- the bias audit
- the human-review duty
- the EU high-risk classification
- the works-council question
Rank that first whenever the tool's contribution to validity is unproven - which assessment-validity-audit.md finds it usually is.
The moves themselves, in efficiency order:
- Remove salary-history questions everywhere and publish real ranges - near-zero effort, satisfies the largest number of jurisdictions at once (20+ US states/localities plus the EU directive), and the same move the funnel evidence recommends anyway.
- Use structured interviews and scored work samples: not just higher-validity, but the legally defensible option under adverse-impact challenge, via content validity. Effort already spent building Artifacts 1-3, so the compliance benefit is free.
- Ensure a human meaningfully reviews every automated rejection (GDPR Art. 22) - low effort per batch, but a standing obligation for as long as the tool runs.
- Commission an independent bias audit before any AI screening tool is used in NYC, and check high-risk-AI obligations before use on EU-based roles. LL144 penalties reach $1,500/day and a failed audit (impact ratio <0.80) becomes evidence against the employer.
- Involve works councils in Germany/France before deploying any assessment, regardless of the tool's validity evidence - highest effort, and it can block deployment outright.
What this order starves: the bias audit and the works-council submission, both bottom on ratio and non-negotiable where they apply. Promote either to first the moment the loop touches its trigger - an AI tool on a NYC candidate, any assessment on a German or French role.
2026 is an inflection year (EU AI Act high-risk obligations and the Pay Transparency deadline both land in it); a loop designed earlier needs a compliance re-review now.
references/mock-call-design.md›
# Mock-call work sample: design and rubric
Evidence is labeled inline with a short parenthetical naming its source, as in SKILL.md. Work samples carry .33 operational validity (peer-reviewed) - second tier behind structured interviews - and show smaller group differences than cognitive tests, though that adverse-impact advantage is overstated in incumbent samples (peer-reviewed).
## Scenario design (4-step method from 30 Minutes to President's Club, a practitioner source)
1. The candidate sells **the hiring company's** product, not theirs - tests adaptability to an unfamiliar pitch instead of a rehearsed one.
2. Send a prep pack: example prospect, company one-pager, deck, demo video. Serious preparation is part of the test.
3. Once live, make it genuinely hard. A soft scenario measures nothing.
4. Two people in the room: one plays the prospect ("champion"), one scores silently ("below-the-line" analyst). Two is enough - panels add no measurement validity (peer-reviewed).
Timing: ~30 minutes of mock call, ~30 minutes reserved for feedback, the re-run, and buffer.
**The coachability re-run:** stop, give one specific piece of feedback, and re-run that segment. Score the delta, not the first attempt. This is the behavioral replacement for "are you coachable?".
## Content by role and segment
- **SDR / high-velocity / B2C:** mock cold call - opener, permission, two objections, close for a meeting (B2C: close for the sale; add required-disclosure checkpoints where the product is regulated).
- **Mid-market AE:** mock discovery call - problem-first questioning, one budget objection, next-step close.
- **Enterprise AE:** mock discovery plus a short deal-strategy exercise (a multi-stakeholder scenario: who can kill this deal, what do you do this week) or a territory/30-60-90 presentation.
## Rubric
Behaviorally-anchored rating scale, 0-3, 5-8 categories. Write what a 0 and a 3 sound like for each category; never bare numbers. Freeze the rubric for a quarter so cross-candidate comparison holds; retrain scorers with frame-of-reference calibration (walk them through anchor examples together) before live use.
Example category anchors (adapt per product):
| Category | 0 sounds like | 3 sounds like |
| --------------------- | --------------------------------- | --------------------------------------------------------------------------------- |
| Opening/agenda | Launches into pitch | Earns explicit permission, sets purpose and plan |
| Questioning | Feature-checklist interrogation | Ladders from symptom to business consequence; builds on answers |
| Listening/adaptation | Talks over answers, rigid script | Uses the prospect's words; adapts when the scenario shifts |
| Objection handling | Argues or folds | Restates, probes (when does budget reset? who else decides?), confirms resolution |
| Close/next step | No ask, or vague "I'll follow up" | Specific next step with owner, date, and stated purpose |
| Coachability (re-run) | Same behavior repeated | Feedback visibly integrated in the re-run |
Conversation-metric anchors are all vendor data, correlational, from conversation-intelligence platforms' own customer bases, never independently replicated - context for scorers, never scoring criteria:
- a widely cited analysis put top-performer talk share near 43% (talking over ~65% of the call correlated with losses); the same vendor's 2025 refresh across 326,000 calls found closed-won deals near 57% rep talk - the two figures conflict, which is itself the caveat
- discovery success peaked around 11-14 targeted questions spread evenly across the call, per a 519,000-call analysis by the same vendor
## Anti-gaming countermeasures
Candidates game mock calls with:
- memorized qualification frameworks recited as checklists
- over-rehearsed scripts
- paid interview coaching for common formats
- AI-assisted live prep
Counter all four the same way:
- Inject one unscripted objection mid-call ("I'm not sure we have budget for this right now").
- Shift the scenario mid-call off any rehearsed path.
- Score improvisation and adaptability, never polish. A strong candidate probes the injected objection and closes for a specific next step; a weak one folds or replays a script that no longer fits.
## Scoring discipline
Both scorers submit independent, evidence-quoted scores before comparing notes - the same mechanical-combination rule as the rest of the loop (peer-reviewed). The work sample is a "can do" capability check; never substitute a personality or style signal for it - traits are not skills.
references/ramp-plan-template.md›
# 30-60-90 ramp plan template and benchmarks
Evidence is labeled inline with a short parenthetical naming its source, as in SKILL.md. Benchmarks skew North American B2B SaaS; label them directional for other markets.
## Benchmarks to size the plan against
- SDR ramp ~3.2 months; median tenure ~1.5-1.9 years; productive window ~15 months (survey, 2025 SDR edition, n=351).
- AE ramp 6.2 months - the highest in that survey's history; experience required at hire 3.7 years, up from 2.7 in 2022; 48% of reps hit annual quota, down from 51% (2024) and 66% (2022); median OTE $200K; median quota $960K; quota-to-OTE 4.6x (survey, 2026 AE edition, n=158 - same source lineage as the 2024 figures, newer edition; cite 2026 as current).
- Cohorts ramping in ≤3 months posted 29% higher pipeline scores, and formal onboarding correlates with faster ramp (survey) - the strongest single argument for a structured plan over ad hoc onboarding.
- SDR total attrition averages 39%, nearly two-thirds involuntary, materially worse below $20M revenue (survey). Crowdsourced rep-level data reads attainment lower (~43%) than the company-leader surveys (crowdsourced) - individual and leader reports differ; keep the tiers apart.
- Cost of a bad hire: the circulated "30% of first-year earnings" floor is loosely sourced; replacement estimates run 50-200% of salary (survey and vendor-derived). The real cost is ramp carry plus foregone pipeline, which dwarfs interview spend - rigor upstream is cheap by comparison.
## Template
**Pre-start:**
- accounts and tooling provisioned
- buddy assigned (a peer, never the manager)
- calendar seeded with shadowing sessions
- prep pack sent
**Day 1 / week 1:**
- Orientation spread across the week; do not overload day 1, deep work starts week 2.
- Team 1:1s, first shadowing, and ICP/product self-study begin.
**Day 0-30 - certify:**
- Product, ICP, tooling and methodology training.
- Shadowing live calls; sandbox mock calls scored on the same rubric used in hiring.
- **Certification gate: a real product/ICP exam with a pass bar** (a job-knowledge test - .40 validity class, peer-reviewed; one large SaaS org used a 100-question exam, per a practitioner account). No live pipeline ownership before passing.
**Day 31-60 - coached activity:**
- Live activity at reduced targets; every call recorded and reviewed in the first 30 live days.
- Weekly coaching against the hiring rubric; first opportunities (AE) or first booked meetings (SDR).
- Check-ins in Situation-Behavior-Impact form: specific context → observable behavior → impact; document agreed actions and the next check-in date.
**Day 61-90 - own the number:**
- Self-sourced pipeline; first closed deals (AE, cycle permitting) or full meeting quota trajectory (SDR).
- Forecast accuracy review (AE); handoff-brief quality review (SDR).
- Day-90 review against the scorecard outcomes, prorated.
**Quota and pay during ramp:**
- Stepped quota 25/50/75/100% over successive quarters (compress for SDR/high-velocity).
- SDRs typically carry ~50% quota during ramp on full base.
- AEs get a non-recoverable draw in months 1-3 (up to 6 for enterprise).
Publishing the ramp draw in the offer is part of what makes an aggressive-variable plan hireable.
**Check-in cadence on milestones:**
- nudge at 50% of a milestone deadline
- urgent at 80%
- escalate to manager at 100%
**Optional 180-day extension:** inspect expectations on a 60-90-180 cadence for long-cycle enterprise roles, where day 90 predates the first closeable deal (practitioner estimate).
## Early-warning signals (act, don't archive)
- Low activity discipline in weeks 2-4.
- No improvement between coached call reps - the ramp-stage coachability failure, mirroring the mock-call re-run test.
- Cannot articulate product/ICP after certification.
- No self-generated pipeline by day 75.
- Poor CRM hygiene from the start.
**Action rule:** a hire missing all three milestones - certification, coached improvement, self-sourced pipeline - by day 90 needs a decision (structured remediation plan with a date, or exit), not hope. Involuntary exits cluster early and the productive window (tenure minus ramp) is ~15 months for SDRs (survey).
## Retro metrics
Track per cohort:
- time-to-productivity vs the benchmarks above
- ramp-step attainment
- 90-day milestone completion
- attrition split regrettable vs involuntary
Regrettable exits indict the ramp or manager; early involuntary exits indict the hiring loop.
references/scorecard-template.md›
# Role scorecard template and worked examples
Evidence is labeled inline with a short parenthetical naming its source, as in SKILL.md - peer-reviewed, survey, crowdsourced, vendor data, or a named practitioner.
## Structure
The outcome-based scorecard method (Smart & Street, _Who: The A Method for Hiring_, 2008) [practitioner, book-length framework]:
1. **Mission** - one sentence stating the job's core purpose.
2. **Outcomes** - 3-8 measurable results, ranked by importance, each quantified with a number and a deadline. These are the KPIs the hire will be judged on.
3. **Competencies** - 5-7 behaviors describing how the person must operate, each with a weight.
The scorecard is reused twice: as the interview rubric (each competency gets questions and a score) and as the ramp milestone map (outcomes become the 30-60-90 targets, prorated).
## Worked example - mid-market AE
**Mission:** Convert qualified pipeline into $1.0M net-new ARR in year one by running disciplined discovery-to-close cycles on 3-6 month deals.
**Outcomes (ranked):**
1. Close $1.0M net-new ARR in the first 12 months ($250K by month 6 on the stepped ramp).
2. Maintain 3.5x pipeline coverage from month 4 onward.
3. Hold stage-to-stage conversion at or above team median by month 6.
4. Forecast within ±10% of committed number from month 5.
5. Multi-thread every deal over $30K to 3+ contacts.
6. Pass product/ICP certification by day 30.
**Competencies (weights for mid-market; reweight per segment below):** discovery depth (25%), closing/next-step discipline (20%), objection handling (15%), achievement drive/conscientiousness (15%), coachability (15%), written communication and CRM hygiene (10%).
## Worked example - outbound SDR
**Mission:** Generate 12 qualified meetings per month for the AE team from cold outbound within 90 days of start.
**Outcomes (ranked):**
1. 12 sales-qualified meetings/month from month 4 (ramp: 3/6/9 in months 1-3).
2. Sustain target daily activity (calls + emails + social touches) at team standard from month 2.
3. Meeting-to-opportunity acceptance rate at or above team median by month 4.
4. Pass product/ICP certification by day 30.
5. Handoff briefs complete on 100% of booked meetings (context, need, next step).
**Competencies:** activity discipline (25%), resilience to rejection (20%), coachability (20%), curiosity (15%), written communication (15%), process/CRM hygiene (5%).
## Segment reweighting
| Competency emphasis | Enterprise B2B | High-velocity | B2C/consumer |
| ------------------- | -------------------------------------------------------- | ---------------------------------------- | ---------------------------------------------------------------------- |
| Highest weights | Discovery depth, business acumen, stakeholder navigation | Activity resilience, coachability, drive | Rejection tolerance, drive, script adaptability, compliance discipline |
| Lower weights | Raw activity volume | Multi-threading | Multi-threading, long-cycle forecasting |
The validity evidence does not differ by segment - only which competencies the org weights against it. Achievement/conscientiousness evidence deserves weight everywhere: it predicts objective sales at .41, better than it predicts manager ratings [peer-reviewed, sales-specific meta-analysis].
## Team-stage adjustments
- **First hire:** weight coachability, curiosity, intelligence, work ethic; treat long big-company tenure as a flag to probe, not a credential. Hire the builder/entrepreneur profile over the out-of-industry sales manager [practitioner, Roberge].
- **Hires 5-30:** build the competency weights from the org's own top performers' observed traits, then hire against that profile every time. Do not import another company's profile once internal data exists.
## Tolerance bands and red flags
Define acceptance as a band, not a point:
- "must have" - hard requirement
- "strongly prefer"
- "acceptable"
- "hard no"
Cap must-haves at 5-7 - more shrinks the funnel without adding signal [practitioner, corroborated across established HR hiring practice].
List role-specific red flags with the failure mode each predicts (e.g. "cannot name a lost deal and what they learned → low self-awareness, coaching resistance"; "every story is a solo win → will not multi-thread"). A red flag triggers a debrief discussion, never an automatic rejection on its own.
Before finalizing weights, check team composition: weight the hire toward the competency the current team lacks rather than cloning the strongest incumbent.
SKILL.md›
---
name: sales-hiring
description: Employer-side hiring workflow for SDR/BDR and AE roles, producing an outcome-based scorecard, a structured interview loop with question bank, a scored mock-call work sample, and a 30-60-90 ramp plan with certification gates. Recommends SDR vs AE vs full-cycle from ACV, cycle length and inbound volume, and rests on selection-validity evidence - structured interviews, independent scoring, no brainteasers. Use whenever the user mentions hiring a rep, sales interview questions, a hiring scorecard, a mock call interview, rep onboarding, or a ramp plan, even without the word hiring. Do NOT use for candidate-side prep (mbfinotti/sales-skills@sales-career) or comp plan design (mbfinotti/sales-skills@sales-comp-design).
license: MIT
metadata:
author: Maya-Beth Finotti
version: "1.1.2"
---
# Sales Hiring
Build the four artifacts a hiring manager needs to recruit an SDR/BDR or AE with evidence-backed rigor: a role scorecard, a structured interview loop with question bank, a scored mock-call work sample, and a 30-60-90 ramp plan. Structure beats intuition at every stage - the whole skill exists to enforce that.
Label every benchmark output with a short parenthetical naming its source - peer-reviewed, survey, crowdsourced, vendor data, or a named practitioner - and never present a lower tier as a higher one, or a vendor claim as fact.
If the user turns out to be a candidate preparing for a sales interview rather than a hiring manager running one, stop and point them to `mbfinotti/sales-skills@sales-career` - this skill never coaches the candidate.
Out of scope, do not drift into:
- sourcing candidates
- writing or distributing job ads
- running an applicant tracking system
- designing the commission plan as a strategy exercise
- SDR-to-AE promotion mechanics
Compensation appears here only as the published range and the base/variable split and ramp draw that shape the candidate funnel and the ramp.
## Interview
Ask before producing anything. Stop at these ten questions.
- One question per message.
- Offer the multiple-choice options.
- Skip any question the user's message already answers.
Questions 8-10 exist to re-rank the orderings in this skill before the user commits to a path - ask them here, never later.
1. Team stage: (a) first sales hire, founder still selling; (b) hires 2-4, playbook forming; (c) hires 5-30, repeatable playbook exists; (d) 30+, scaling a known model.
2. Role in mind: (a) SDR/BDR; (b) AE; (c) full-cycle rep; (d) not sure - recommend one for me.
3. Pipeline source mix: (a) mostly inbound; (b) mostly outbound; (c) balanced mix.
4. ACV and typical sales-cycle length: (a) under ~$5K, under a month; (b) ~$5-25K, 1-3 months; (c) ~$25-100K, 3-6 months; (d) over ~$100K, 6+ months.
5. Segment: (a) enterprise B2B; (b) high-velocity inside sales (SMB/mid-market B2B); (c) B2C/consumer.
6. Comp philosophy: base-heavy (60:40 or more) or aggressive variable (50:50, uncapped accelerators)? And will you publish the range in the posting?
7. Hiring jurisdiction(s): US (which states, notably NYC/CO/CA/NY/WA/IL), EU (which countries), UK, other - this drives the compliance checks.
8. Deadline: by what date must the offer be signed? (a) this month; (b) this quarter; (c) no hard date.
9. One-off or compounding: a single hire, or the first of several where the scorecard, question bank and rubric get reused? (a) one-off; (b) compounding.
10. Effort ceiling: how many interviewer-hours per candidate can you spend, and do you have a recruiter and an existing trained panel? (a) under ~5 hours, no recruiter, no panel; (b) ~5-15 hours; (c) more, with recruiting support.
## Reading the answers
Every answer above drives an output; use them, don't just record them.
- **Q3 + Q4 decide the role recommendation (Q2), and they flip the ordering.** Default for ACV under ~$25K on cycles under three months with inbound flow (read `>` as "more of this axis"; the efficiency line is the recommendation):
- efficiency: full-cycle rep > AE only > SDR + AE split
- value (pipeline generated and revenue closed per hire): SDR + AE split > AE only > full-cycle rep
- effort (loops to run, ramps to write, handoffs to police, time-to-fill): SDR + AE split > full-cycle rep > AE only
Above ~$25K ACV on 3-6+ month cycles with an outbound motion, the value line wins and the split leads - one rep cannot both prospect a mapped account list and run six-month deals. **Go/no-go, stated as a deletion:** while inbound volume keeps a closer's calendar full, delete the SDR + AE split from the menu instead of ranking it last, since a ruled-out option parked at the bottom comes back as headcount. Put it back only once AEs show empty-calendar syndrome or the company commits to outbound as a strategy **[survey + practitioner]**. If the user picked a role in Q2 that contradicts this, challenge the pick before proceeding and say why.
- **Q1 decides the candidate profile.**
- First hire: a founder-led-sales handoff problem. Hire a builder who sells without infrastructure, and treat heavy big-company experience as a risk (the "coin-operated" rep who imports a formula instead of building the company's own). Mark Roberge's regression on his own team found aggression and objection-handling nearly uncorrelated with success; coachability, curiosity, intelligence, and work ethic predicted it (his own single-company regression, not a broader study). For a founder or CEO running this interview personally, SaaStr's repeated framing for the first 2-10 reps is a single gut-check test layered on top of the scorecard, not a replacement for it: would you, the founder, actually buy the product from this person (practitioner opinion).
- Hires 5-30: a repeatability problem. Profile from the org's own top performers and hire the same rep every time.
Also warn: doubling the team requires doubling each demand-generation source in proportion, or new hires starve.
- **Q5 sets loop length, work-sample content, predictor weighting, comp split and sourcing** - see B2B and B2C below.
- **Q6 shapes the funnel and the ramp.**
- Base-heavy splits attract risk-averse candidates and suit reps who start with zero pipeline.
- Aggressive 50:50 plans with uncapped accelerators self-select confident closers.
- An unpublished "competitive OTE" repels strong candidates, who read it as a signal the quota is set to fail (practitioner opinion).
Publishing a range is also legally required in a growing set of jurisdictions (Q7).
- **Q7 triggers the compliance pass** in workflow step 4.
- **Q8 promotes the fast-acting options.** A this-month date:
- deletes the 5-6 stage enterprise loop
- promotes the single-loop roles over the SDR + AE split
- promotes reusing an existing question bank over writing one
No hard date leaves every default order below in place.
- **Q9 decides whether the work sample stays starved.**
- One-off: the build amortizes over a single candidate pool, so run the mock call but build it thin - no quarterly rubric refresh, no scorer-training programme.
- Compounding: the frozen rubric and question bank serve every future hire, which promotes building both properly and first - this is the condition in Selection evidence that overrides the efficiency order.
- **Q10 deletes rungs rather than reordering them.**
- Under ~5 interviewer-hours per candidate: cut to 3-4 stages and keep only the structured interview and the work sample.
- No recruiter and no trained panel: every hour in the Selection evidence table becomes the hiring manager's own - re-rank against that, not against a staffed team.
## Workflow
Produce the four artifacts in order. Deliver each one, validate it with the user section by section, then move to the next - never dump all four unreviewed.
1. Run the Interview. Confirm or challenge the role choice using Reading the answers, and state the reasoning in one short paragraph the user can veto.
2. **Artifact 1 - role scorecard.** Build from [references/scorecard-template.md](references/scorecard-template.md). Three parts:
- a one-sentence mission
- 3-8 measurable outcomes ranked by importance and quantified (e.g. "$1.2M net-new ACV in year one", "40 SQLs/quarter"), never vague responsibilities
- 5-7 competencies weighted per segment and team stage
The scorecard doubles as the interview rubric and the ramp milestone map, so quantify outcomes even when the user resists. Avoid the all-around-athlete profile: narrow, deep competence against these outcomes.
3. **Artifact 2 - interview loop + question bank.**
- Design the stage sequence for the role and segment: SDR 3-4 stages over days-weeks; AE 4-6 stages; enterprise AE 5-6 stages over weeks, with an inverted funnel (hiring manager in early) when the candidate pool is small and hot.
- Place the scored work sample at stage 2 or 3, never last. The two placements are identical on value: `value: stage 2-3 == last`, a genuine tie because it is the same instrument, the same rubric and the same score wherever it sits. They separate only on hours: `interviewer hours burned: last > stage 2-3`, at 10+ hours per candidate who cannot sell, plus the extra days of time-to-fill that lose candidates to competing offers. Last placement is therefore deleted from the menu, not ranked below stage 2-3; quality gate 5 enforces it.
- Write 2-3 behavioral plus 1-2 situational questions per scorecard competency with drill-down probes, from [references/interview-question-bank.md](references/interview-question-bank.md). Delete every brainteaser, "sell me this pen", "greatest weakness" and "are you coachable" - they have no demonstrated predictive value **[peer-reviewed + measured internal analysis at scale]**.
- Assign each interviewer a competency, and keep the same interviewers and the same questions across all candidates for the role. Structure, not headcount, is what raises validity; panels add no validity over a single trained interviewer (peer-reviewed) - a panel's value is governance, not measurement.
- Define the scoring mechanics now: every score needs an evidence quote, scores are submitted independently before any debrief, and the final number is combined mechanically with pre-set weights. The hiring manager decides within pre-committed thresholds.
- If the user wants a personality assessment in the loop, apply the assessment rule below and load [references/assessment-validity-audit.md](references/assessment-validity-audit.md).
- Run the compliance pass for the Q7 jurisdictions with [references/legal-landscape.md](references/legal-landscape.md): remove salary-history questions everywhere, publish the range, and flag any AI screening tool for bias-audit and high-risk-AI obligations.
4. **Artifact 3 - work-sample design + rubric.** Build from [references/mock-call-design.md](references/mock-call-design.md). Design a mock call where the candidate sells the user's product, not their own:
- a prep pack sent ahead
- a hard live scenario
- one unscripted objection injected mid-call
- one segment re-run after feedback (the coachability test, replacing the useless self-report question)
Score on a behaviorally-anchored 0-3 rubric with 5-8 categories, frozen for a quarter, with frame-of-reference training for scorers. Score improvisation and adaptability, never polish. Match content to segment: cold call + objection handling for SDR/high-velocity/B2C, discovery depth + multi-threaded deal strategy + territory plan for enterprise.
5. **Artifact 4 - 30-60-90 ramp plan.** Build from [references/ramp-plan-template.md](references/ramp-plan-template.md).
- Day 0-30: product/ICP/tooling training ending in a certification gate (a real exam, pass required); shadowing and sandbox mock calls.
- Day 31-60: live activity with coaching, every call reviewed, first opportunities.
- Day 61-90: self-sourced pipeline, first deals, forecast accuracy.
Add:
- a stepped quota ramp (25/50/75/100% over successive quarters)
- a non-recoverable draw for AEs in months 1-3
- a peer buddy distinct from the manager
- check-ins in Situation-Behavior-Impact form
Include the early-warning list and the action rule: a hire missing all three 90-day milestones (certification, coached improvement, self-sourced pipeline) needs a decision, not hope - involuntary exits cluster early and the productive window is short (survey).
6. Run the Quality gate below on all four artifacts. Iterate until it passes.
7. If your harness has persistent memory, store the scorecard, the frozen rubric, the loop design and the ramp milestones - later coaching and review sessions start from them. Without memory, tell the user to keep the artifacts as the canonical pack and re-supply them.
## Selection evidence
Two orderings govern the loop and they disagree, so never quote one as if it were the other. Validity says what an instrument buys; efficiency says what to build first with the hours available. Read `>` as "more of this axis"; the efficiency line is the build order.
Validities are corrected operational values (Sackett et al. 2022 correction, peer-reviewed; earlier figures were overstated).
- efficiency: structured interview > job-knowledge exam > reference check > cognitive test > work sample
- value: structured interview > job-knowledge exam > work sample > cognitive test > reference check
- effort: work sample > structured interview > job-knowledge exam > cognitive test == reference check (tied: nothing to build for either, under an hour to run)
- compliance cost: cognitive test > work sample > reference check > job-knowledge exam == structured interview (tied: content-valid instruments whose only exposure is the questions inside them - neither triggers an audit, a consent duty, or third-party data handling)
Rows sit in efficiency order, not validity order:
| Instrument | Validity | Build | Run per candidate | What the row costs |
| ----------------------- | ------------- | ----------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Structured interview | .42 | a week - bank, anchors, interviewer training | an hour per interviewer | Highest validity, and the week is reused by every later hire. Validity climbs with structure (.20 unstructured to .57 highly structured), so the build is what buys the coefficient |
| Job-knowledge exam | .40 | a week to write | near-zero, self-marking | Your own product cannot be tested pre-hire, so this lands as the day-30 certification gate (Artifact 4) - it buys a ramp decision, not a hire decision |
| Reference check | ~.26 or lower | near-zero | an hour per finalist | Cheap and honest about what it is: a fraud and claim check. Never weight it like an interview or work sample |
| Cognitive test | .31 | near-zero, off the shelf | near-zero | The cheapest row, and the reason cheapness must not lead: near-useless against objective sales (caveat below) and the heaviest compliance load of the five |
| Work sample (mock call) | .33 | a week to design, then a standing job - quarterly rubric refresh, frame-of-reference recalibration per new scorer | an hour x two scorers, plus candidate prep hours that cost drop-off and days of time-to-fill | Lowest ratio here, mandatory anyway. Lower adverse impact than cognitive tests, though that gap is overstated in incumbent samples |
**Deleted from the menu, not ranked last: biodata (empirically keyed), .38 (peer-reviewed).** Empirical keying needs outcome data on hundreds of past hires; no team at Q1 (a)-(c) has it, and a demoted row would silently return as scope.
**What the efficiency order starves: the work sample.** It carries real validity, and it is the only "can do" check in the loop. It is simultaneously:
- the priciest instrument to build
- the only one needing continuous upkeep
- the only one spending the candidate's hours
So a ratio buries it every round. Promote it to first build anyway when any of these hold:
- this is the first sales hire, so no incumbent benchmark exists to calibrate an interview against
- the loop keeps passing candidates who interview well and miss quota
- Q9 answered compounding, which amortizes the build across every future hire
**Default:**
- build the structured interview, then the work sample
- run reference checks on finalists only
- leave the job-knowledge exam in the ramp
Add a cognitive test only with a commitment to validate it against objective sales outcomes, which is a quarter's work on data most teams do not have.
Sales-specific caveat (peer-reviewed): cognitive ability correlates .40 with supervisor ratings but only .04 with objective sales results - it predicts what managers think of reps, not what reps sell. Achievement/conscientiousness shows the inverse pattern (.41 against objective sales), making it the more trustworthy signal to weight. Validate any gate against objective sales outcomes, not performance reviews.
Combine scores mechanically: formula-based combination predicts at .44 versus .28 for experts blending impressions in their heads - holistic judgment burns up to half the validity (peer-reviewed). This is the one choice in the skill needing no ratio: mechanical combination is both the higher-value option and the cheaper one - a weighted sum against an hour of argument. Implement it through independent scoring before the debrief: the loudest voice in the room cannot anchor scores that are already submitted.
**Re-rank before using any ordering above.** Every one is a default, not a law, and each shifts with context and with who executes it:
- a trained interview panel already in place drops the structured interview's build toward zero and moves the work sample to first
- no recruiter turns every hour in the table into the hiring manager's own hour, which shortens the loop and promotes the cheap rows
- a first hire has no top-performer data, which is what promotes the work sample; a thirtieth hire has it, which is the only condition that puts empirical keying back on the menu
**Assessment rule:**
- Treat any commercial sales-personality assessment as unvalidated unless it appears in the Buros Mental Measurements Yearbook with published peer-reviewed criterion validity.
- Never let any assessment gate or veto a candidate.
- Vendor "predictive validity" claims above ~.55 are marketing, not psychometrics - genuine criterion validities rarely exceed .50 even for the best predictors.
The named-vendor audit lives in [references/assessment-validity-audit.md](references/assessment-validity-audit.md).
## B2B and B2C
The shared foundation is identical across enterprise B2B, high-velocity inside sales, and B2C/consumer - no segment gets a pass on rigor because its loop is shorter:
- the written scorecard
- structured interviews
- the scored work sample
- independent scoring before the debrief
- mechanical score combination
- the banned-question list
- the legal constraints
Where the segments genuinely diverge - not ranked, deliberately: Q5 selects the column, so these are a lookup rather than a menu, and ordering options the user cannot choose between would be false precision.
| Dimension | Enterprise B2B | High-velocity inside sales | B2C/consumer |
| ------------------- | ----------------------------------------------------------------------------- | ----------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
| Loop length | 5-6 stages, several weeks; invert the funnel for scarce talent | 3-4 stages, days to ~2 weeks | 2-4 stages, days; speed wins candidates |
| Work-sample content | Mock discovery, multi-threaded deal strategy, territory/30-60-90 presentation | Mock cold call, objection handling | Mock consumer call end-to-end (the call is often the whole deal), objection handling, script adaptability |
| Predictor weighting | Discovery depth, business acumen, stakeholder navigation | Activity resilience, coachability, drive | Activity resilience, rejection tolerance, drive; compliance discipline where the product is regulated |
| Comp split | ~50:50 with large absolute variable | Base-heavy at entry (new reps have no pipeline) | Often commission-heavy with a draw; expect higher washout - budget backfill |
| Sourcing | Proactive headhunting of a small mapped list | Inbound, high-volume funnels, fast screening | High-volume funnels, fast screening |
Flag where the benchmarks come from when working B2C: the published ramp, tenure and attainment numbers skew North American B2B SaaS (survey). Present them to a B2C user as directional, not as their baseline.
## Quality gate
Score the four artifacts against all twelve checks before final delivery. Pass threshold: 12/12 - each check traces to a specific finding, so a miss is an evidence violation, not a style choice. Iterate until every check passes; report the checklist with the artifacts.
1. The scorecard has 3-8 outcomes, every one quantified and ranked (outcome-based scorecard method).
2. The role recommendation (SDR vs AE vs full-cycle) explicitly cites the user's ACV, cycle-length and inbound answers, including the empty-calendar rule where SDRs were requested.
3. Every interview question maps to a named scorecard competency and is behavioral or situational; the identical question set applies to every candidate (structured > unstructured, peer-reviewed).
4. Zero brainteasers, "sell me this pen", "greatest weakness", or "are you coachable" anywhere in the bank.
5. The scored work sample sits at stage 2 or 3 of the loop, never last.
6. The work sample sells the hiring company's product, injects one unscripted objection, and re-runs one segment after feedback as the behavioral coachability test.
7. The work-sample rubric is behaviorally anchored (0-3, 5-8 categories) and declared frozen for a quarter; any conversation-metric anchors are labeled as vendor data, correlational.
8. Interviewers submit independent, evidence-quoted scores before any debrief, and the final combination is mechanical with pre-set weights (mechanical > holistic, peer-reviewed).
9. No assessment gates or vetoes any candidate; any assessment in the loop is labeled per the assessment rule.
10. The posting content states a real base + OTE range; no salary-history question appears anywhere in the loop; jurisdiction flags from Q7 are addressed.
11. The ramp plan contains a day-30 certification gate, a day-60 coached-activity milestone, a day-90 self-sourced-pipeline milestone, a stepped quota ramp, and the missing-all-three action rule.
12. Every benchmark number in all four artifacts carries its evidence label, and no vendor claim is presented as fact.
## Common failure modes
Not ranked, deliberately: every row is a defect with one mandatory fix, and ordering defects by ratio would imply the bottom ones are optional.
| Failure | Fix |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| Hiring the all-around athlete | Score narrow, deep competence against the 3-8 outcomes; an impressive-but-mismatched background is a mismatch |
| Coin-operated first hire | For hire #1, weight builder traits over big-logo experience; the imported formula rarely survives contact with the company's motion |
| Mock call run last | Move it to stage 2-3; apply the placement ordering in workflow step 3 |
| Debrief-first scoring | Collect written independent scores before anyone speaks; the debrief calibrates, it never originates scores |
| One blended gut score | Per-competency scores with evidence quotes and mechanical combination; a single average hides the sticking point |
| Assessment as gate | Apply the assessment rule; unvalidated tools are color at most, and no tool vetoes alone |
| "Competitive OTE", no range | Publish base + OTE; vagueness filters out exactly the target candidates and violates pay-transparency rules in several jurisdictions |
| Adding SDRs to look grown-up | Apply the empty-calendar rule; SDRs before AEs are at capacity just adds cost and handoff loss |
| Day-1 overload in the ramp | Spread orientation across week 1; deep work starts week 2 |
| Hoping past day 90 | Missing all three 90-day milestones triggers a decision; late rescue attempts rarely beat the base rates |
| Cloning the team or the top performer | Weight the hire toward what the team is missing; different profiles add coverage, not risk |
| Quota set from hope | Sanity-check quota against ramp benchmarks in the ramp reference; unattainable year-one numbers drive early regrettable attrition |
## KPIs
Judge the hire - and this skill's output - after the fact with:
- **Ramp attainment vs the stepped schedule**: percent of the 25/50/75/100 quota steps hit on time, per cohort.
- **Time-to-productivity** vs role benchmarks (SDR ~3.2 months **[survey, 2025 ed.]**; AE 6.2 months, the highest recorded **[survey, 2026 ed.]** - an earlier edition of the same survey read 5.7; use the most recent edition available and treat older ones as trend context, not the baseline).
- **90-day milestone completion**: certification pass, coached-call improvement, self-sourced pipeline - tracked per hire.
- **Attrition, split regrettable vs involuntary**: SDR total attrition averages 39%, nearly two-thirds involuntary, worse below $20M revenue (survey); involuntary exits clustering early usually indict the loop, regrettable exits the ramp or the manager.
- **Loop process KPIs**: time-to-hire vs role norms (SDR 21-35 days; enterprise AE 45-70+ days, practitioner estimate), stage-to-stage drop-off, offer-accept rate.
If you can browse the web or query current market data, refresh comp ranges and ramp benchmarks before quoting them; otherwise label every figure with its edition year and tier, and tell the user to verify locally.
Compensation and legal content in this skill is informational, not legal advice - have counsel and HR review postings, assessment use, and any offer language, especially commission/OTE and guaranteed-vs-variable pay terms.
## Reference
- See [references/scorecard-template.md](references/scorecard-template.md) for the scorecard structure, worked SDR and AE examples, and segment weighting.
- See [references/interview-question-bank.md](references/interview-question-bank.md) for the banned list, per-competency question banks, and the debrief scoring form.
- See [references/mock-call-design.md](references/mock-call-design.md) for the scenario design, the 0-3 rubric with labeled anchors, and anti-gaming countermeasures.
- See [references/ramp-plan-template.md](references/ramp-plan-template.md) for the full 30-60-90 template, certification gates, and ramp benchmarks.
- See [references/assessment-validity-audit.md](references/assessment-validity-audit.md) for the named-vendor validity audit behind the assessment rule.
- See [references/legal-landscape.md](references/legal-landscape.md) for the 2026 jurisdiction-by-jurisdiction compliance table.
- See `mbfinotti/sales-skills@sales-career` for the candidate's side of this table - interview prep, positioning, progression; this skill never coaches candidates.
- See `mbfinotti/sales-skills@sales-comp-design` for the OTE and pay mix an offer positions, and the draw schedule that backs the ramp period.
- See `mbfinotti/sales-skills@sales-org-structure` for the org design that decides which seats exist before recruiting into them - role mix, SDR-to-AE ratio, and management spans.
- See `mbfinotti/sales-skills@sales-quota-setting` for the ramp-relief schedule the 30-60-90 plan's quota expectations key off - this skill sets the certification gates, that one sets what the rep carries at each stage.
- See `mbfinotti/sales-skills@sales-call-review` to score the new hire's live calls during ramp with the same evidence-quoted discipline.
- See `mbfinotti/sales-skills@sales-objection-handling` for objection-handling craft useful when writing mock-call scenarios.