Consistent
Every candidate for a role answers the same core questions in the same stages.
Consistency means every candidate is measured against what actually drives success in the role, instead of against the impression they happened to leave.
Sales hiring · The full process
Most sales advice is gut-feel or a generic questions list. This is the repeatable system I used to take a sales org from under $1M to over $100M ARR — the 4 Pillars of Talent Selection and the five jobs a hiring loop has to do.
By Lucas Price, founder of Yardstick. He scaled the sales organization at Zipwhip from under $1M to over $100M ARR, which was acquired by Twilio.
Why it is worth the work
Most sales teams are not hitting their number. In Ebsta’s 2024 B2B Sales Benchmarks report — built on 4.2 million opportunities across 530 companies — 69% of reps missed quota, even after quota targets were cut 19% year over year. Just 15% of sales teams had more than half their reps reach 80% of quota. When most of the field misses, who you hire matters more than almost anything else you do as a leader.
Bad sales hires are also expensive. Leadership IQ’s study of more than 20,000 hires found 46% fail within 18 months — fired, pushed out, or a hire the manager wouldn’t make again — while only 19% become unequivocal successes, and 89% of those failures trace to attitude and fit, not skill. Sales runs worse than average: HBR pegs annual sales turnover at 25–30%. The direct cost of replacing a rep is real money before you count a dollar of lost pipeline: DePaul University research puts it at roughly $115k per rep, and SBI’s fuller accounting — unmet quota, ramp, management time, relationship damage — lands between $500k and $1.35M for a single bad hire.
Earlier in my career, roughly 40% of the salespeople I hired didn’t work out. Building a structured process around the four pillars below is what moved that to about 10% for me. That’s my experience as an operator, not a guarantee — but the difference between a 40% and a 10% miss rate on a sales team compounds into a completely different company.
A rep who is 10% better doesn’t stay 10% ahead. They ramp faster, learn faster, close more, and raise the bar for everyone around them, so the gap widens every quarter. Structured hiring is how you tilt the odds toward that person on purpose instead of by luck.
The framework
Everything in a good sales hiring process serves four pillars. They are not specific to sales. The same four hold for any role you hire, and what changes for sales is what you put inside them. If an interview loop isn’t doing all four, it’s leaking predictive power. Every stage, question, and scorecard below is one of these pillars made practical.
Every candidate for a role answers the same core questions in the same stages.
Consistency means every candidate is measured against what actually drives success in the role, instead of against the impression they happened to leave.
Questions are tuned to what the job requires, and ask what a candidate did — not what they would hypothetically do.
Real past behavior predicts better than a hypothetical, and gets harder to fake as your follow-ups sharpen.
Every interviewer rates every candidate on the same scorecard, on the same scale, against the same competencies, skills, and role outcomes.
Scoring puts every candidate on one scale. It adds objectivity to the evidence; it does not make the decision for you.
Somewhere in the loop, the candidate does the job — a role play, a mock discovery call, a call-coaching exercise.
A work sample is the most honest signal you get: behavior you watched instead of behavior they described.
The architecture
A sales loop has five distinct jobs to do. That is not the same as five interviews. Adding conversations is not the goal, and past a certain point it stops helping at all (there are numbers on exactly where that point sits further down). The goal is that every job here gets done deliberately by someone, rather than three of them happening by accident in the same call.
How you package them is a judgment call about the role and the level. High-volume AE and SDR hiring can cover screening, competencies, and a work sample in three well-designed conversations. For a frontline manager, you spend more of the loop on the career history, because by then there is a real track record to examine and less reason to infer from traits alone. Either way the loop is built for that specific role. Every interview should be. What changes with seniority is where you spend the depth, not whether the questions are tailored. Whatever the shape, design each stage by asking the same three questions: what is this stage for, how do we prepare the interviewer, and what tools does it need? For the questions themselves, we keep competency-grouped banks for account executives, SDRs, and sales managers.
The hardest interview to design well — decide who is worth the loop's time by confirming required traits and skills, the right experience, and genuine interest and fit.
Deep behavioral interviews on the handful of traits that separate strong reps: grit, adaptability, emotional intelligence, resourcefulness, planning, and coachability.
A structured walk through the candidate's career — the “Who” interview from Geoff Smart and Randy Street, with roots in Topgrading. Its superpower is that the follow-ups are boring — and therefore unspinnable.
The Assessed pillar in action: a role play or mock discovery call where you watch the candidate sell, coach them once, and see whether they incorporate it.
Done right, not as a rubber stamp: wait until you are offer-ready, ask structured fact-based questions, and go beyond the candidate's hand-picked list.
Behavioral questions
The most common way a sales interview loop leaks signal is bad questions. Three fixes matter most.
Hypothetical “what would you do if a deal stalled” questions lose predictive power fast as the role gets more complex — for complex roles their validity drops to around r = .30, while past-behavior questions hold near r = .51 (Huffcutt et al., 2004). Lead with a real situation and ask how they actually handled it.
In a 45–60 minute interview, three well-followed-up questions beat a checklist of ten. Depth is where the signal is — a great follow-up chain is what separates a rehearsed story from a real one. Our framework for follow-up questions is the engine that makes this work.
“Tell me about a time you showed grit” tells the candidate exactly what to perform. Ask for the situation first, then find the grit in their follow-ups — and skip the tired “tell me about a time” opener so they can’t pattern-match to a canned answer.
The research on behavioral interviewing looks mixed until you notice what is actually varying, which is the interview itself. In the most recent comprehensive meta-analysis, structured interviews are the single best predictor of job performance among all selection methods (Sackett et al., 2022), and the word doing the work in that sentence is “structured.” The same conversation, run without that structure, predicts far less.
So a bad interview is usually a process failure rather than a people problem. Most managers have never been handed a set of questions, a reason behind each one, or a chance to practice before it counted. Give them those three things and they run good interviews. Write the questions before the loop opens, write the follow-ups you want asked when an answer comes back thin, and let people run the same interview enough times to get good at it. The quality you get out is mostly a property of what you put in.
Scoring
A scorecard does its most valuable work before anyone interviews. Writing one forces the hiring team to say out loud what this role actually needs, and that is usually where you discover how much you disagreed. Four people can walk into a kickoff with four different jobs in mind and never find out, because nobody made them write it down.
Arguing your way to one list is most of the value. So is cutting that list down, which is the harder half: every line you cut is someone admitting a thing they cared about is not actually what separates a strong rep from an average one. What you are left with is a shared definition of the job, and a loop that is aimed at it. Only then does the scoring do its part, which is to keep every candidate on one scale instead of ranking whoever left the strongest impression.
The shape of the card matters less than that agreement. What follows is the shape we have found works, but a card your team argued its way to and actually fills out beats a perfectly structured one nobody bought into.
One card scores three things. Competencies are behavioral: how a rep operates, like resourcefulness, coachability, grit, and curiosity. Skills are the selling motions a candidate has to arrive with, because you can’t teach them quickly: orchestrating a complex, multi-stakeholder deal, selling to executives, winning against an entrenched incumbent. Your product, your CRM, and your sales process don’t belong here, since a strong hire absorbs those in weeks; score the ability to learn fast instead, which is a competency. Role outcomes are the results that would make you call the hire a success, written as forward-looking predictions with real numbers, like “will source 40% of first-year quota from self-generated pipeline.”
Cap each section at three or four items, and treat that as a ceiling, not a target. Brainstorm wide, then cut hard: a card with a dozen lines on it is a card no interviewer scores honestly, and it buries the few things that separate strong reps. Some sections should come out nearly empty. An entry-level SDR role may need no hard skills at all, or only the basics, like holding a conversation on the phone and writing an email someone replies to. Writing three items when the honest answer is none is how a card fills up with things you never actually screen for.
Naming a competency is not the same as defining it. Put “resilience” on a card and it means whatever each interviewer privately assumes it means. Say what a rep on your team actually has to bounce back from: a six-month enterprise cycle that dies at procurement, forty dials a day into a market that has never heard of you, a champion who leaves three weeks before close. Once it is that specific, the interview question mostly writes itself, and two interviewers scoring the same answer land in roughly the same place. Do this for every competency you keep. It is the step most teams skip, and it is why their scorecards produce numbers nobody trusts.
What goes on the card comes from the role, not from a generic list of sales virtues. Break down the sales motion and work backward from what winning actually requires, then decide whether this hire should mirror a strength your best reps already have or cover a gap nobody covers. Building a sales interview scorecard works through all three sections with examples.
If you don’t yet have a scorecard for the role, and most teams don’t because there’s no performance data on day one, that’s the normal starting point rather than a failure. Yardstick’s AI scorecard generator is built for exactly this cold start: it drafts a first-pass sales scorecard (competencies, skills, outcomes, and the anchored scale) that you react to instead of composing from nothing. The draft won’t be perfect. It is meant to be wrong in places and easy to fix, and reacting to a real draft is far faster than staring at a blank page. Generate a first-draft sales scorecard and start editing instead of starting from zero.
The decision
Here’s the counterintuitive part: more interviews stop helping surprisingly fast. Google’s internal analysis, published on its re:Work blog in 2017, found four interviews were enough to predict a hiring decision with 86% confidence — after that, each additional interviewer improved accuracy by less than 1%. Four good, independent interviews beat six rushed ones.
The word doing the work there is independent. Interviews only add signal if opinions form on their own — so no comparing notes before the debrief. In the debrief itself, the leader speaks last, so the room isn’t anchored to the most senior opinion.
And the most important rule: never turn the decision into a score cutoff.“Anything over a 3.5 is a hire” feels objective, but it hands the decision to a formula. The scorecard’s job is to add objectivity to parts of the process — to make sure every candidate was evaluated on the same evidence. The decision itself is a human one, made by a team looking at that evidence together. Structure raises the quality of the inputs; it never replaces the judgment.
Put it into practice
The hardest part of adopting a structured process isn’t believing in it — it’s the blank page. Writing a competency question bank, a scorecard, and an interview plan for every role is real work, and it’s why most teams stay unstructured.
That’s the specific problem Yardstick is built to solve. The four pillars aren’t a philosophy you have to operationalize on your own — they’re the product. Yardstick is a structured-interview ATS where your team creates job-specific interview plans, runs consistent interviews, and collects scorecards on one scale, with AI assistance to get you off the blank page. If you work with a coding agent like Claude Code or Codex, it can operate Yardstick through the yardstick CLI — drafting interview plans, questions, and scorecards for you to review — while every sensitive action waits for your approval. Agents prepare the work; you decide.
The AI generators are designed around the cold start: a first-draft scorecard or question set that’s useful precisely because it’s easy to correct. You’re not asking the model to be right; you’re asking it to give you something to react to. That’s the fastest path from “we should interview more consistently” to actually doing it.
FAQ
Run a structured process instead of a series of conversations: the same questions for every candidate (Consistent), tuned to the role and focused on real past behavior (Behavioral), scored by every interviewer on one scorecard (Scored), and including a work sample where the candidate actually sells (Assessed). Then decide as a team on the evidence — without turning the score into a cutoff.
Around four independent interviews is the sweet spot. Google's re:Work analysis found four interviews predicted the hiring decision with 86% confidence, and each additional interviewer after that added less than 1%. What matters more than the count is that opinions form independently before the debrief.
Ask fewer, deeper behavioral questions — about three in a 45–60 minute interview — focused on what the candidate actually did, with strong follow-ups. Tie them to the specific competencies that predict success in the role, such as grit, adaptability, resourcefulness, and coachability.
Consistent (the same questions for every candidate), Behavioral (questions tuned to the job, asking what they did), Scored (one scorecard and one scale for every interviewer), and Assessed (a work sample or role play, not just conversation). They are the role-agnostic foundation of a structured hiring process — here applied to a sales hire.
Use one scorecard for every candidate, scoring three things: competencies (behavioral traits like resourcefulness and coachability), skills (the selling motions a candidate has to arrive with because you can’t teach them quickly, such as orchestrating a complex enterprise deal), and role outcomes (the results that would make you call the hire a success, written as predictions with real numbers). Keep each section to three or four items, rate every item 0–4 where 0 means “not enough information,” score independently, and resist weighting the criteria. The scores exist to confirm everyone evaluated the same evidence, not to auto-decide the hire.
Making the decision a score cutoff. A scorecard adds objectivity to the evidence, but the best hiring decisions are human decisions made by a team looking at that evidence together. Handing the call to “anything over a 3.5 is a hire” throws away the judgment the whole process was built to inform.
Generate a first-draft sales interview question set and a matching scorecard, then edit them to fit your role — or see how Yardstick runs the whole process in one system.