Interview Recipes

Part 3 of 3 · Preparing the interview

Interview rating scales: how to write behavioral anchors (BARS)

August 2, 2026 · 10 min read · Max Korpinen

A sales leader I supported used to debrief with me after he had interviewed the candidates I scheduled for him. The debriefs were short. One candidate was excellent on paper and sharp in the room, but she had small children, and he could not trust that the household would leave her enough time for the job. That was a no. Another looked promising enough that he flew to a different country to meet him, and the verdict that came back was about the way the man sat in his chair. Also a no.

Neither verdict contains a word about the work, and in much of the world, acting on the first one is illegal. Yet both were delivered with complete confidence, because nothing in the process asked for anything more: no criteria decided in advance, nothing written down, no definition of what a good answer would have looked like. Skip that definition and the interview still gets scored, just not by you. It gets scored by whichever impression happened to stick.

Rating scales are the preparation step that closes this gap, and the research keeps finding that they are among the most powerful parts of interview structure. This article covers what an anchored rating scale is, the evidence for it, and how to write one for any interview question in about ten minutes. No psychometrics degree required.

A 1-to-5 scale is not yet a rating scale

Most interview scorecards already have numbers on them. The interviewer meets the candidate, and afterwards clicks four out of five stars on "communication". This feels like measurement, but if nothing defines what a four is, the number is just a gut feeling wearing a digit. The previous article in this path opened with two interviewers reading one bear answer in confidently opposite ways; unanchored numbers have the same problem, with the added danger of looking objective in a spreadsheet.

The fix is old and unglamorous: anchor every point on the scale to a written description of what an answer at that level actually sounds like. Then scoring stops being "how did she strike me?" and becomes "which of these descriptions does the answer in my notes match?" The first question has no right answer. The second one does.

What is a behaviorally anchored rating scale (BARS)?

A behaviorally anchored rating scale, BARS in the literature, is a rating scale where each score is defined by a concrete description of observable behavior rather than by an adjective. Not "excellent communicator" but "let the customer finish before responding, separated the person's anger from the underlying problem, and named a concrete fix". The technique comes from performance appraisal research and moved into interviewing because it solves the same problem in both places: two raters, one performance, and no shared definition of good.

In a structured interview, the scale is attached to the question, not to the candidate. Each question gets its own anchors describing a weak, an acceptable, and an excellent answer to that question, and the interviewer rates the answer right after hearing it. This is exactly how the US federal government's structured interview guide builds its interviews: benchmark descriptions written for each level of each question, so that any trained interviewer scoring the same answer lands in the same place.

You do not need the federal apparatus. The lightweight version splits the work in two: for each question, two written descriptions in observable behavior, what a good answer contains and what the red flags sound like; across the whole interview, one shared four-point scale that pins every score to those descriptions. The anchors stay question-specific, where the research wants them, while the scale stays identical from question to question, so nobody is relearning the ruler mid-interview. That is the version this article teaches, and the one printed on every plan this site produces.

Why anchored scoring works

The evidence for structured interviews is unusually strong, but the headline validity of .42 (Sackett et al., 2022) is an average sitting on a wide spread. Some structured interviews predict performance about as well as any selection method ever has; others barely beat winging it. Where you land on that spread depends on how well you structure, and in the classic taxonomy of interview structure (Campion, Palmer and Campion, 1997), eight of the fifteen components are about evaluation rather than questions. Anchored rating scales sit at the center of that half: rate each answer, on a scale anchored in behavior, based on notes, independently.

The sharpest single finding comes from the same meta-analysis that established behavioral questions as the spine of the interview. Taylor and Small (2002) found that both behavioral and situational questions improved markedly when answers were scored against descriptively anchored scales. A well-chosen question scored on gut feel gives back much of its edge; the anchors are not garnish on the method, they are load-bearing.

Anchors also make interviewers agree with each other. Reviews of interview reliability consistently find that interviewers using highly structured formats, anchored scoring included, agree with one another far more than interviewers running free-form conversations. An interview where the score depends on which interviewer the candidate happened to get is measuring the interviewer.

And there is the fairness dividend, which my sales leader's debriefs illustrate better than any citation. Research on structured interviews finds meaningfully smaller demographic score differences than in unstructured ones, and the mechanism is not mysterious: an anchor that says "named a concrete fix and checked back afterward" leaves no line on the form where "small children at home" can become a score. Criteria written before the interview, tied to the job, are what make a decision defensible to a candidate, to a court, and to your own conscience.

How to write anchors in three steps

You need your job analysis and your chosen questions. For each question, the work looks like this.

Step 1: ask what the question is for

Every question you chose is there to test a competency for this role. Before writing any anchors, answer in one sentence: what would this role's version of a strong answer have to contain? Your job-analysis notes are the raw material, especially the who-was-good-at-this incidents. If your notes say the best support person you knew never escalated a problem without a proposed answer, that sentence is practically an anchor already.

Step 2: write what good looks like, and the red flags

For each question, write two short descriptions of the answer, not the candidate: one describing what a good answer contains, one naming the red flags. Two or three lines each is plenty, and two rules keep them honest:

Red flags deserve their own lines rather than being the mirror image of good, because weak answers fail in characteristic ways worth naming in advance: the story that is really about winning the argument, the plan recited with no why behind it. Name the failure mode and you will recognize it in the room instead of merely feeling it.

Step 3: score every answer on one shared scale

With the descriptions written, the scale itself can stay the same for every question in the interview:

ScoreAnchored to
1 · WeakThe answer mostly matches the red flags.
2 · MixedSome real evidence, but significant gaps.
3 · StrongThe answer matches "what good looks like".
4 · OutstandingBeyond good: a systemic fix, taught others, changed the outcome.

Two design choices in that table earn their keep. There is no unanchored midpoint to hide in: with four points, every score leans weak or strong, and "mixed" is a description you can defend, not a shrug. And "matches what good looks like" sits at 3, not at the top, which keeps headroom for the answer that goes beyond the criteria while keeping 4 anchored to behavior rather than to a halo. Expect most real answers to land on 2 and 3; a 4 should be rare enough to be news in the debrief. If half your candidates are scoring 4, the good-answer description has gone soft, not the candidates brilliant.

Before you interview anyone, run the two-interviewer test: imagine two colleagues scoring the same recorded answer with your descriptions and this scale. Would they land on the same number? If not, the wording that would split them is the wording to fix.

A worked example: scoring the angry-customer question

The job-analysis article built its example around a support lead role, with customer orientation flagged as the must-assess competency. Here is a question that slot might hold, as it appears in the pantry:

The card's "what good looks like" and red flags are the two descriptions from step 2, already written in observable behavior. Apply the shared scale to them and scoring the support lead's answers looks like this:

ScoreWhat the answer sounds like
1 · WeakMostly the red flags: the customer cast as unreasonable start to finish, a story about winning the argument, or the anger handled while the problem never got fixed. Or no specific incident at all.
2 · MixedA real incident with real evidence, but significant gaps: they calmed the customer, yet the underlying problem, the follow-up, or what it cost stays vague even after probing.
3 · StrongMatches what good looks like: let the customer finish, separated the anger from the actual problem, named a concrete fix or an honest limit, and checked back afterward.
4 · OutstandingBeyond good: everything in 3, and the answer changed something larger. The complaint became a fix for the next hundred customers, or the approach got taught to the rest of the team.

Notice what the scale never mentions: rapport, confidence, likability, posture. Every line is checkable against the notes you take in the room, and your follow-up probes exist precisely to fill the details the anchors ask about. Question, probes, anchors: one instrument, designed as a set.

Using the scale in the room

Four habits turn written anchors into reliable scores.

Score each answer, right after it is given. An overall rating formed at the end of the interview is an impression with a memory problem. Per-question scoring is what the anchors are shaped for, and it is one of the evaluation components the research keeps rewarding (Campion et al., 1997).

Same questions, same scale, every candidate. The scale's power is comparison. A 4 only means something when the candidate before scored 3 on the same anchors for the same question.

Score independently before you discuss. If your interview has multiple interviewers, everyone writes their numbers down before anyone speaks. A panel that talks its way to a shared score has one opinion with extra signatures.

Combine scores by arithmetic, not by debate. Add the numbers, weight the must-assess competency from your job analysis if you flagged one, and let the total carry the comparison. Kahneman tells the story in Thinking, Fast and Slow of replacing intuitive army interviews with trait-by-trait scoring and watching prediction improve; his conclusion was not that intuition is worthless, but that it earns its place only after the disciplined collection of objective information. Score first. Then, with the numbers on the table, your judgment has something honest to work with.

Anchors without the apparatus

The reason my sales leader could turn a chair into a verdict was not that he was an unusually bad judge of people. It was that nothing in the room asked him a better question than "what did you think?" A rating scale is that better question, printed in advance: here is what good looks like for this role, which of these did you actually hear?

Every question card in the pantry ships with its anchors built in: the "what good looks like" and red-flag sections are the two descriptions the scale scores against, pre-written in observable behavior, and every plan this site prints carries the four-point scoring scale next to them. Pick your questions, spend your ten minutes sharpening the good and red-flag lines for your role, and put the scale on the interview guide beside each question. Or let the planner assemble the plan for you, free, in about three minutes, no account needed.

Either way, decide what answers are worth before you hear any. It is the difference between running a measurement and collecting impressions, and it costs ten minutes.