Hiring process · Scored
How to build an interview scorecard.
A scorecard gets every candidate rated on the same scale, whatever the role. Building a good one is mostly upstream work: break the role down into what success actually requires, then translate that into the three things a scorecard scores. Those are competencies, skills, and role outcomes.
These are the scorecard design rules Yardstick’s founder, Lucas Price, used building the sales organization that scaled from under $1M to over $100M ARR at Zipwhip, later acquired by Twilio, and has since taught to other leaders. The examples on this page apply them to engineering, operations, and marketing roles; the method is the same for any role you hire.
Start here
Start with the role, not the card.
What belongs on a scorecard comes from the role, not from a template. Break the job down and work backward from what winning actually requires. For a backend engineer on a small team, winning might mean shipping reliably in a messy codebase with little spec. For an operations manager, it might mean building processes that survive a doubling of volume. For a content marketer, it might mean turning search demand into pipeline without a team behind them. The card falls out of that question, and the question is different for every role.
Then look at your team, not just the role. Two questions are worth asking before you finalize the card. One: do you want someone who mirrors what your best people already do well, with the same competencies aimed at the same outcomes? That is a proven, lower-risk profile. Two: is there something your team is missing, a strength no one has yet, that would make the whole team better if someone brought it? Sometimes the highest-leverage hire covers a gap instead of cloning a strength, and deciding which you are optimizing for changes what goes on the card. That is the heart of context-aware hiring: cloning a strength or covering a gap on purpose.
Once you know what the role needs and what your team is missing, translate it into the three things a scorecard scores: competencies, skills, and role outcomes. The discipline is the same for all three: brainstorm wide, then cut hard to the 3–4 items that actually matter. This is the “Scored” pillar of the 4 Pillars of Talent Selection, applied to a single role. And the structure pays off: structured interviews are the single best predictor of job performance among all selection methods (Sackett et al., 2022), when the structure is disciplined rather than sprawling.
Part one · Competencies
The behavioral traits. Pick three to four.
Competencies are the durable, behavioral traits that make someone good at the job: how they operate, not what they have done. Choose the ones this role actually demands and keep it to 3–4. If picking feels like cutting something essential, you are doing it right. A card with twelve competencies is a card no interviewer scores honestly, and it buries the traits that separate strong candidates.
Example — competencies for a Senior Backend Engineer
- Ownership — takes a production issue from first alert to shipped fix without waiting to be assigned, and follows through on the boring parts.
- Learning speed — gets productive in an unfamiliar codebase within days, and asks questions that show they read the code first.
- Communication — explains a technical decision to a non-engineer without losing the substance, in writing and out loud.
- Pragmatism — ships the simple version first, and can say out loud when good enough is good enough.
A different role picks different traits. An operations manager card might score process thinking and calm under load; a content marketer card might score audience empathy and editorial judgment. The test is the same everywhere: would the person who wins in this role need this trait every week?
Each competency needs interview questions behind it, tuned to what the candidate actually did rather than what they would hypothetically do. Our AI interview question generator drafts behavioral questions for a specific competency, and a framework for follow-up questions helps you dig until each score is defensible.
Part two · Skills
What you can’t teach fast, so check they arrive with it.
Skills are concrete technical, functional, or domain abilities, as opposed to the behavioral competencies above. The filter that decides what earns a spot is sharper than “is it a skill”: put in the ones a candidate has to already have, because you cannot teach them quickly. The reason to score a skill at all is to verify prior ability where ramping someone from zero would be too slow or too risky.
Run every candidate skill through one test: could a strong hire learn it in their first few weeks? Your codebase, your tools, your internal process, your product’s domain vocabulary: these almost always clear that bar, because you teach them to everyone in onboarding anyway. So don’t screen for prior mastery of them. Screen for the ability to learn them fast, which is a competency. Reserve the skills section for abilities that take years to build. For many roles the section ends up nearly empty, and that is fine; an empty skills section is a finding, not a failure.
Teachable — hire for the ability to learn these
- Your codebase and conventions — every engineer learns these in the first month, whatever you hire for.
- Your deploy pipeline and tooling — learnable in days; not a reason to pass on someone strong.
- Your product’s domain terms — absorbed in onboarding by anyone curious.
Hard to teach — verify they already have these
- Distributed-systems debugging in production — tracing a failure across services under time pressure takes years of reps to build.
- Zero-downtime data migrations — moving live data stores without an outage window is learned by doing it, badly at first, somewhere else.
- API design for external consumers — designing interfaces other companies build against, where mistakes are permanent.
Watch the line between a skill and an outcome, because the same theme can appear as both. “Has run zero-downtime migrations before” is a skill: have they done it. “Will ship our billing migration by month nine with no Sev1 incidents” is an outcome: will they do it here. Score the experience in the skills section and the target in the outcomes section, and don’t let a résumé line stand in for the goal.
Part three · Role outcomes
Role outcomes: predictions you can check.
Role outcomes are the results that would make you call this hire a success, written as forward-looking predictions. Start by brainstorming: across the first couple of years, at 3, 6, 12, and 24 months, ask “what would make me say this hire worked out?” and write down everything. That is the generative step; don’t edit yet.
Then make the hard cuts, and sharpen what survives. The test for a well-written outcome is checkability: is it specific enough that at six or twelve months you could settle it without anyone arguing about the verdict? Usually that makes it a number. Sometimes it is a clean yes or no, like “shipped the migration with no Sev1,” and that qualifies too. “Cut onboarding time in half” is checkable without a stated figure. What fails the test is the vague virtue: “build a strong pipeline” or “raise the engineering bar” tells an interviewer nothing about what evidence to dig for. Give each outcome the what and the how.
Example — outcomes across three roles
- Senior Backend Engineer: will ship the billing-service migration by month nine with no Sev1 regressions — by breaking it into independently shippable stages.
- Operations Manager: will cut new-hire onboarding time in half by month twelve — by rebuilding the checklist into a self-serve runbook.
- Content Marketer: will grow organic search into a third of qualified pipeline by month twelve — by building out the comparison-page cluster the team has never staffed.
Three different jobs, one format: a prediction specific enough to check. Every interviewer who reads one of these knows exactly what evidence to dig for, and every candidate for the role is being projected against the same targets.
Scoring
Write the anchors first.
Score every item, competency, skill, or outcome, on one scale with behavioral anchors: for each number, a written description of what that level of evidence looks like. The anchors are the load-bearing part, more than the range. When a 3 on “ownership” describes the specific behavior a 3 looks like, two interviewers hearing the same answer land on the same score. Without anchors, the same scale produces different numbers from different people, and the comparison the scorecard exists for quietly breaks. Yardstick scorecards run 0–4, and one anchor is worth stating outright: 0 means “not enough information,” not “bad.” If something never came up in the interview, score it 0 instead of guessing.
And don’t weight the criteria. Assigning “ownership is 30%, communication is 20%” feels rigorous, but it is false precision, and false precision is dangerous because it tempts you to let the score decide. Once “anything over a 3.5 is a hire” takes over, you stop having the debrief conversation, and that conversation is the one place you would discover a score was wrong or that the number hides a disqualifying gap.
So use the scores to make the conversation better, not to replace it: score independently before comparing notes, have the leader speak last so the room isn’t anchored, and then decide as a team. Where same-scale scoring pays off is the comparison: once every candidate is on one scale, you can lay them side by side in Yardstick’s candidate comparison, one scored grid across competencies, skills, and role outcomes with role-average baselines, and see the real differences instead of arguing from memory.
Put it into practice
Don’t start from a blank page.
Here is the real catch with everything above: on day one you have no performance data for the role, so you are writing your first scorecard from a hypothesis. That is the normal starting point, not a failure. But the blank page is why most teams never build the card at all.
That is the specific problem Yardstick’s AI scorecard generator is built for. It drafts a first-pass scorecard for your role: competencies, skills, outcomes, and the anchored scale, ready for you to react to instead of composing from nothing. The draft won’t be perfect. It is meant to be wrong in places and easy to fix. You are not asking the model to be right; you are asking it for something to edit, and reacting to a real draft is far faster than staring at an empty template.
If you already work with a coding agent like Claude Code or Codex, it can operate Yardstick through the yardstick CLI, and a general assistant can connect over MCP: drafting the scorecard, the matching question bank, and the interview plan for you to review, while every sensitive action waits for your approval. The agent prepares the work; you decide. Yardstick is a structured-interview ATS where your team builds job-specific interview plans, runs consistent interviews, and collects scorecards on one scale, with AI to get you off the blank page.
Generate a first-draft scorecard and edit from there.
Worked example
The same method, worked through a sales role.
To see the whole method applied end to end, read how to build a sales interview scorecard. It runs an Account Executive card through all three parts: which competencies survive the cut, which selling motions count as hard-to-teach skills, and how quota-shaped outcomes get written as predictions. The content changes with the role; the method on this page does not.
FAQ
Common questions about interview scorecards.
What are the three parts of an interview scorecard?
Competencies, skills, and role outcomes. Competencies are behavioral traits: how someone operates, like ownership or learning speed. Skills are the concrete abilities a candidate must already have because you cannot teach them quickly; this section is optional and often nearly empty. Role outcomes are the results that would make you call the hire a success, written as checkable predictions, like “will ship the billing migration by month nine with no Sev1 regressions.” Every item is scored on the same behaviorally anchored 0–4 scale.
How many items should each section of a scorecard have?
Three to four per section. If cutting to that number feels like losing something essential, the cut is working. A card with a dozen lines is a card no interviewer scores honestly, and it buries the traits that separate strong candidates from average ones.
What is the difference between a competency and a skill on a scorecard?
A competency is behavioral: how a candidate operates in any job, like resourcefulness or communication. A skill is a concrete ability the candidate must arrive with because you cannot teach it quickly, like debugging distributed systems in production. The test: could a strong hire pick it up in their first few weeks? If yes, as your codebase, tools, and process can be, score the ability to learn fast instead, which is a competency. If no, it is a skill worth verifying.
How do you write role outcomes for a scorecard?
Brainstorm across the first couple of years: at 3, 6, 12, and 24 months, what would make you say this hire worked out? Then cut to the few outcomes that define success and write each as a forward-looking prediction specific enough to check. The test is checkability, not the presence of a number: “will cut onboarding time in half by month twelve” and “shipped the migration with no Sev1” both pass, because at review time nobody would argue about the verdict.
Should you weight interview scorecard criteria?
No. Assigning percentages to each item feels precise but is false precision, and the danger of false precision is that it tempts you to let the score decide. Once “anything over a 3.5 is a hire” takes over, you skip the debrief conversation, which is the one place you would discover a score was wrong. Score each item honestly on an anchored scale and let a human team read the pattern.
What rating scale should an interview scorecard use?
The anchors matter more than the range. Yardstick uses 0–4, where each number carries a written description of what that level of evidence looks like, and 0 means “not enough information,” never “bad.” Anchored descriptions are what make two interviewers land on the same score; an unanchored 1–10 scale just produces ten flavors of gut feeling.
Can AI generate an interview scorecard?
Yes. Yardstick’s AI scorecard generator drafts a first-pass card for any role: competencies, skills, outcomes, and the anchored scale. You react to the draft and edit it, which solves the cold-start problem of building a card before you have performance data. The draft is meant to be a fast, imperfect starting point rather than a finished card; a human reviews and approves it.
Generate a first-draft scorecard.
Get a first-pass card for the role you are hiring now: competencies, skills, role outcomes, and the anchored scale, ready for you to react to and edit instead of building from a blank page. Then run the whole structured process in one system.
.webp?dpl=dpl_HN4igPZwjzxmoh3vVfcp8d1iYC5q)