Hiring process · Scored

The complete interview scorecard template.

Two of your interviewers just scored the same answer a 2 and a 4. Neither of them is wrong. The form they both filled in has Communication, Culture Fit and Technical Ability down the left, a one to five scale across the top, and nothing anywhere on it that says what a 3 looks like.

Below is a complete scorecard, free to copy, including the part almost every template leaves out: a written description of what each score means. There is a fully worked example for a Project Manager, and a pattern you can move to any role.

The structure

The three sections, and what goes in each.

A complete scorecard has three sections, and the difference between them is not cosmetic. Sorting a criterion into the wrong one is how a card ends up scoring the things that were easiest to write down.

Competencies. How they operate. Broad human qualities: how someone behaves, judges, collaborates, communicates, learns, adapts and leads. The test is transferability. If the quality would help someone succeed in many unrelated jobs, it belongs here. Name them in a word or three, and never name a tool, a method, a certification or a workflow. “Directness” qualifies. “Agile ceremony facilitation” does not.

Skills. What they have to arrive with. Concrete technical, functional or domain capabilities, tied to training, practice, tools, methods or certifications. Include one only if it normally takes meaningful prior education or career experience and could not be picked up in the first few weeks of your onboarding. Anything behavioral belongs in competencies instead. Knowing your CRM is not a skill on this card. Salesforce administration is.

Role outcomes. What a good hire will have done. Forward-looking predictions, each written starting with “Will,” each carrying a metric, a timeframe or a target, each one or two sentences long. “Improve delivery” fails. “Will deliver every committed project inside the window agreed at kickoff, across all four quarters of the first year” passes.

Three or four items per section. More than that and interviewers stop reading the card and start scoring from memory. In each of competencies and skills, mark one or two as must-haves, so the debrief knows which lines it cannot trade away.

That is what to write down. The harder question is what the numbers mean.

The scale

The anchors are the scorecard.

Here is one criterion, scored the way the rest of this page will score everything:

Influence without authorityMust have

Gets commitment from people who do not report to them, and keeps it without escalating.

Anchored 0 to 4 rating scale for Influence without authority
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Escalates to a manager as the first move when someone outside their line does not deliver. Owns the tracking, not the outcome.
2Chases people and keeps the item on a list. Gets movement while they are pushing, and the same handoff fails the same way next time.
3Gets commitments from people outside their reporting line and holds them to those commitments without needing a manager in the room. Fixes the handoff after it fails once.
4Is the person other teams tell first when something is going wrong, because working with them has been worth it before. Builds the standing arrangement that removes the class of problem, not only this instance of it.

The four written levels are the anchors. A scale without them is not a scale, it is a feeling with a number attached, and it is why two people who heard the same story write down different digits.

Three things about the scale itself are worth stating once, because they are where most cards go wrong.

0 means “not enough information,” not “bad.” You never write a 0 anchor, because it means the same thing on every criterion: the interview did not cover it. A 0 is a note to whoever runs the next conversation, not a mark against the candidate. Keeping “we did not cover it” separate from “we covered it and it was weak” is most of what stops a scorecard from turning into noise.

3 is the bar. The four levels you write are one clearly below expectations, two partially meets, three meets, four exceeds. So a 3 is not the midpoint of a range you hope people beat. It is what a person you would hire actually does, and a 4 is rarer than most interviewers score it.

Write the anchors before you meet anyone. Written afterwards, they describe the candidate you just talked to. That is the whole failure mode the scorecard exists to prevent.

Anchors take real work to write, which is why almost every template you can download stops at the criteria and hands you a bare one to five scale. Here is what a finished one looks like.

Worked example

A worked example: a complete Project Manager scorecard.

Project Manager, because the job exists at a surveying consultancy, a mechanical contractor and a software company alike. The criteria below are an example, not a standard. Yours should come from what winning actually looks like in your role, at your company, this year.

Competencies

Four, and every one of them would matter in a job with nothing to do with project management. That is the test.

Influence without authorityMust have

Gets commitment from people who do not report to them, and keeps it without escalating.

Anchored 0 to 4 rating scale for Influence without authority
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Escalates to a manager as the first move when someone outside their line does not deliver. Owns the tracking, not the outcome.
2Chases people and keeps the item on a list. Gets movement while they are pushing, and the same handoff fails the same way next time.
3Gets commitments from people outside their reporting line and holds them to those commitments without needing a manager in the room. Fixes the handoff after it fails once.
4Is the person other teams tell first when something is going wrong, because working with them has been worth it before. Builds the standing arrangement that removes the class of problem, not only this instance of it.

Directness

Says the uncomfortable thing while it can still change the outcome.

Anchored 0 to 4 rating scale for Directness
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Raises problems once they are undeniable. Bad news arrives at the deadline, described as something that happened rather than something they saw coming.
2Will say the hard thing when asked directly, and does not volunteer it. Softens the message enough that the listener can miss it.
3Raises a problem while it is still uncertain and still fixable, says plainly what they think is wrong, and does it without making it personal.
4Says it early to the person least likely to want to hear it, and does it in a way that keeps that person working with them afterwards.

Practical judgment

Decides with incomplete information, and can reconstruct the reasoning afterwards.

Anchored 0 to 4 rating scale for Practical judgment
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Waits for more information, or decides fast and cannot say why. Treats every open question as equally worth resolving.
2Reaches sound decisions on familiar ground, and stalls when the situation is new or the information is thin.
3Decides with what is available, names the assumption they are betting on, and can say what would make them change their mind.
4Tells the decision worth agonising over from the one worth making in five minutes, and spends their attention accordingly. The reasoning holds up to someone who disagrees with the conclusion.

Ownership

Treats the result as theirs, including the parts they do not control.

Anchored 0 to 4 rating scale for Ownership
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Reports on the work rather than owning it. When something outside their control fails, the account stops at whose fault it was.
2Owns their own tasks reliably, and treats the gaps between people as somebody else's to close.
3Takes responsibility for the outcome including the parts they do not control, and goes and closes the gaps rather than logging them.
4Picks up the problem nobody owns because it is in the way, and hands it back better defined than they found it, without turning it into a claim on territory.

Skills

Three, each one a body of knowledge someone arrives with: a method, a quantitative discipline, a commercial regime. The test is the first few weeks. If the person sitting next to a good hire could teach it to them by Friday, it is not a skill on this card, however much you want it done well.

Critical path schedulingMust have

Builds dependency-aware schedules, identifies the critical path, and re-baselines when work moves. Requires real practice with scheduling tools and dependency logic.

Anchored 0 to 4 rating scale for Critical path scheduling
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Reads a schedule as a list of dates. Treats every late item as equally urgent, because nothing in the plan distinguishes them.
2Can trace a dependency when it is written down, and misses the ones that were never captured. Builds the plan around who is available rather than around what unblocks what.
3Builds the plan around dependencies, names the critical path, and re-sequences the work when something moves.
4Sequences to reduce risk as well as to satisfy dependencies. Pulls the uncertain work forward so a bad answer arrives while there is still time to use it, and can say what the plan becomes if the riskiest assumption is wrong.

Earned value management

Measures cost and schedule performance against a baseline using earned value, and reads the resulting indices to forecast where the project lands.

Anchored 0 to 4 rating scale for Earned value management
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Reports percent complete and spend to date. Cannot say whether the project is ahead or behind in a way that survives a follow-up question.
2Can produce the figures when a template asks for them, and treats them as reporting output rather than as a forecast. Acts on a gut read instead.
3Baselines the work, tracks cost and schedule performance against it, and uses the indices to say where the project lands if nothing changes.
4Reads the trend early enough to change the plan while it is still cheap, and can explain to a sponsor what the numbers mean in money and dates without hiding behind the acronyms.

Contract administration

Works from the statement of work: prices scope changes, raises change orders, and holds the commercial boundary with clients and subcontractors.

Anchored 0 to 4 rating scale for Contract administration
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Treats the contract as paperwork somebody else owns. Absorbs added scope, and meets the commercial consequence once the budget is gone.
2Knows the scope is changing and raises it after the work has started, when the only options left are the cost and the date.
3Reads the statement of work as the operating document. Prices an addition in time and money, and takes a change order back to the client before the team starts on it.
4Sets the contract up so change is expected and cheap to process. Holds the boundary with a client and keeps the relationship, and tells the addition worth absorbing from the one worth pricing.

Role outcomes

Written as predictions, each with a number or a date in it. If this person is a success, these are true a year from now, and you could settle each one without arguing about it. They get anchors like everything else: what the evidence for that prediction looks like at each level.

Will deliver every committed project inside the window agreed at kickoff, or re-baseline it in writing before the original date passes, across all four quarters of the first year.

Anchored 0 to 4 rating scale for Will deliver every committed project inside the window agreed at kickoff, or re-baseline it in writing before the original date passes, across all four quarters of the first year.
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1No evidence of having owned a date that other people depended on.
2Has owned dates. The record is either dates moving late and quietly, or dates padded far enough that hitting them proves little.
3Can point to a run of projects where the committed window held, or moved openly and early, and can name the projects and the windows.
4The same, and they changed how dates get committed on the team they joined. Can describe what the process looked like before and after.

Will reduce projects that reach their due date carrying a slip leadership had not already heard about to zero by month nine.

Anchored 0 to 4 rating scale for Will reduce projects that reach their due date carrying a slip leadership had not already heard about to zero by month nine.
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Their examples of delivering bad news are all examples of it arriving at the deadline.
2Raises problems at the point they became undeniable rather than the point they became likely.
3Can walk through a specific problem they surfaced early, what it cost them to raise it before they were certain, and what got done with the time that bought.
4The same, and they built something that made early warning routine rather than personal: a check, a forum, a habit that outlasted their attention.

Will cut dropped handoffs between the two most failure-prone teams by half against a baseline recorded in month one, by month nine.

Anchored 0 to 4 rating scale for Will cut dropped handoffs between the two most failure-prone teams by half against a baseline recorded in month one, by month nine.
ScoreWhat that level of evidence looks like
0Not enough information gathered.
1Treats a dropped handoff as the other team’s problem.
2Fixes each dropped handoff as it happens, and the pattern does not change.
3Can name a specific recurring handoff failure they diagnosed and changed, and say what the two teams do differently now.
4The same, and the fix survived them. It kept working after they moved on to another project.

Ten criteria, forty anchors. That is one role.

Get this card drafted for your own role and edit from there. Yardstick writes the first pass, anchors included.

Your role

Moving this to your role.

The pattern transfers; the content does not. Copy the shape, not the criteria.

Start from what winning looks like, not from a list of traits. Write down what this person will actually have done in twelve months if the hire worked. Those are your role outcomes, and everything else follows from them: the competencies are how someone gets there, and the skills are what they cannot get there without.

Decide whether you are cloning a strength or covering a gap. The same job title needs a different card depending on the team it joins. A project manager arriving into a group that already sequences well but never says anything early is being hired for different criteria than one joining a team that communicates constantly and cannot hold a date. This is the input most cards skip, and it is worth reading the difference properly.

Write the anchors for 1 and 4 first, then fill in 2 and 3. The ends are easier to picture, and the middle stops being vague once they exist.

Ask the questions the anchors need. An anchor about what someone did when a date slipped only gets scored if somebody asks about a date that slipped. The card and the interview plan are one artifact, not two.

The full method, including how to break a role down before you write a single criterion, is on how to build an interview scorecard. This page is the finished object; that one is how you get to it. If the role you are hiring for is a sales rep rather than a project manager, there is a worked card for that one at sales interview scorecard.

Keeping it honest

Two rules that decide whether the card survives.

Two things reliably wreck a card after it is written, and both arrive looking like improvements.

Do not weight the sections. Assigning 40% to competencies and 30% each to the rest produces a number that looks precise and is not. The weights encode a judgment nobody debated, and then the arithmetic makes the call. Score every item on the same scale and let the pattern be visible.

Do not set a cutoff. There is no total that means hire. A candidate with two 4s and a 1 is a different proposition from one with straight 3s, and which you want depends on the team and the role, not on a threshold. The card exists to make the debrief better, not to replace it.

Two habits make the rest of it work. Score independently before anyone compares notes, because a scorecard filled in after the group conversation records the conversation. And have the most senior person in the room speak last.

None of this dismisses instinct. An experienced interviewer’s read on a candidate is real evidence, and the anchors are what let that read be stated in terms someone else can weigh. What the scorecard removes is not judgment, it is the situation where two people are certain they agree and are not.

What it gives you back is comparison. Once every candidate for a role has been scored against the same anchors, you can lay them side by side and see where they actually differ, which is what a candidate comparison is for.

The hard part

Why you will not finish this by hand.

Everything on this page is free to copy. If you are still interviewing off a bare one to five grid, it is probably not because you disagree with any of it. It is that the version above took an afternoon to write, for one role, and there are four more roles open.

That is the part Yardstick does for you. Point it at a role and it drafts the whole card: the competencies, the skills you cannot teach, the role outcomes written as predictions, and the anchors for every score. Then you edit it. The draft will be wrong in places, and wrong is a much better starting point than blank, because arguing with a written 3 takes a minute and inventing one takes an afternoon. The interview questions that make the anchors scorable come with it.

If you already work with a coding agent, you do not have to open the app at all. Claude Code or Codex can run the yardstick CLI to draft the scorecard and the interview plan for a role and bring them back for you to approve, and a general assistant connected over MCP can do the same. Yardstick is the system of record behind it. Nothing gets used until you approve it, and on a hire human review still matters.

Generate a first draft scorecard for your role. Your first 3 Jobs are free. After that it is pay-as-you-go on active Jobs, with no seats and no lock-in. See pricing.

Yardstick is a structured-interview ATS: the scorecard, the questions that feed it, and the comparison across candidates live in one place rather than in a form, a doc and someone’s inbox. If you want to see the scorecard generation on its own, that is interview scorecard software.

FAQ

Common questions about scorecard templates.

What should be on an interview scorecard?

Three sections. Competencies, the broad human qualities that would help someone succeed in many unrelated jobs. Skills, the concrete technical or domain capabilities a candidate has to arrive with because they take prior training or real practice. And role outcomes, forward-looking predictions written as “Will…” with a metric or a timeframe attached. Three or four items in each section, one or two of the competencies and skills marked must-have, and a written description of what each score means for every item.

What does each score mean on a 0 to 4 interview scorecard?

You write four levels per criterion: 1 clearly below expectations, 2 partially meets, 3 meets, 4 exceeds. The wording is specific to that criterion, not shared across the card, and that wording is the anchor. A 0 is available on every criterion and always means “not enough information,” so a gap in the interview never reads as a weakness in the candidate. Because 3 means meets, it is the bar rather than the midpoint, and a 4 is rare.

Should you weight interview scorecard criteria?

No. Weighting produces a number that looks precise while hiding a judgment nobody debated, and then the arithmetic makes the call instead of the team. Score every item on the same scale, and read the pattern rather than the total.

What is the difference between a competency and a skill on a scorecard?

Transferability, and how long it takes to acquire. A competency is a broad human quality that would help someone succeed in many unrelated jobs, so it is never named after a tool, a method or a certification. A skill is a concrete technical or domain capability that normally takes prior education or substantial practice, and could not be picked up in the first few weeks of onboarding. Salesforce administration is a skill. Learning a new system quickly is a competency.

How many criteria should an interview scorecard have?

Three or four per section, so nine to twelve in total. Past that, interviewers stop reading the card and start scoring from memory, which is the thing the card was meant to fix.

Who fills out the interview scorecard, and when?

Every interviewer who spoke to the candidate, independently, before the debrief. Filled in afterwards, the card records the group conversation rather than the interview, and the independence that makes the scores worth comparing is gone.

Written by

Lucas Price built the sales organization that took Zipwhip from under $1M to over $100M in ARR before its acquisition by Twilio, and hired the people who did it. Yardstick is the hiring system he wanted then.

Generate a first draft scorecard for your role.

The competencies, the skills you cannot teach, the role outcomes written as predictions, and the anchors for every score, drafted for the role you are hiring now and ready for you to edit. Then run the whole structured process in one system.