The headline number for structured interviewing is a doubling: structured interviews predict job performance at roughly twice the rate of unstructured ones, .42 against .19 in the biggest modern meta-analysis (Sackett et al., 2022). The fine print is less quoted. That .42 is an average sitting on a wide spread, and the bottom of the spread reaches down into unstructured territory. Plenty of interviews that call themselves structured perform as if nobody had structured anything.
The spread is not a mystery. Every interview in that dataset had questions; many had scorecards. What varied is what the people in the room did with them. A guide that gets abandoned two questions in, a rating scale filled from a gut impression at the end of the day, probes improvised for the likable candidate only: each one quietly hands validity back. Structure is not a property of your documents. It is a property of behavior, and behavior is trained, not downloaded.
Everything earlier in this path happened at a desk: the analysis, the question types, the anchors, the probes, the guide that holds them. This article is where the path walks into the room. It covers what the evidence says training changes, what a full program contains, and a minimum viable version that a small team without a training department can run in one afternoon.
Training is a structure component, not an HR nicety
The canonical taxonomy of interview structure (Campion, Palmer and Campion, 1997) lists fifteen components that make interviews predictive, and interviewer training is one of them, sitting on the evaluation side alongside anchored scales, note-taking, and statistical combination of ratings. That placement is worth pausing on. The researchers did not treat training as onboarding trivia; they treated it as part of the instrument, because most of the other components only exist if the interviewer performs them. Same questions for every candidate is a decision made in the guide and kept or broken in the room. So is limiting ad-hoc follow-up. So is scoring right after the answer instead of from memory at dinner.
The US federal government's structured interview guide is blunter than the academics: "It is essential to train the person who will administer the structured interview. Interviewer training increases the accuracy of the interview." Essential is a strong word from a document that is otherwise very calm, and the federal practice matches it: panels get trained before anyone faces a candidate, not after something goes wrong.
If you have been following this path, one reframe makes the rest of the article simpler. You are not training people to interview. You are training them to run an instrument that already exists: the guide, its questions, their anchors, their probes. That is a much smaller and much more learnable job than "become good at interviews", and it is the reason the minimum viable version at the end of this article fits in an afternoon.
What the evidence says training changes
The most direct look at the interviewer's side of the equation is a meta-analysis by Huffcutt and Woehr (1999), which compared interview validity across studies where interviewer-related practices differed. Three practices came with better prediction: interviewers who were trained, interviewers who took notes, and the same interviewers being used across all candidates for a role. One widely assumed practice made no measurable difference: panel interviews were not more valid than one-on-one interviews.
That null result deserves more attention than it gets, because the panel is where anxious hiring processes reflexively spend money. Adding a third and fourth silent observer to the room feels like rigor; the evidence says it mostly adds calendar cost. Google reached the same conclusion from its own data with the Rule of Four: after four structured interviews, additional interviewers barely move the decision. If you have budget for rigor, the return lives in training the interviewers you already have, not in multiplying them.
Reliability tells the same story from another angle. Interviewer agreement rises with structure (Conway et al., 1995), and agreement is the ceiling on validity: a signal two of you read differently is not yet a signal. Training is the mechanism by which two interviewers become one instrument, which is the same job the shared guide does on paper.
Calibration is the part doing the heavy lifting
Most of what a training program contains is orientation: here is the process, here is the guide, here is what happens after. Necessary, cheap, forgettable. The part with real evidence behind it has a clumsy academic name, frame-of-reference training, and one core move: give every rater the same mental model of what good looks like, then make them practice using it until their scores converge.
A frame-of-reference session looks like this. Raters are shown the dimensions they will score and what each level of performance actually sounds like. Then they rate sample answers, compare their scores against a reference standard and against each other, and talk through the gaps. Then they do it again. The point is not the scores in the practice session; the point is replacing each rater's private theory of "a good answer" with the shared, written one. Rater training of this kind is the best-studied piece of the training puzzle, and it holds up: the original meta-analysis (Woehr and Huffcutt, 1994) found it improved rating accuracy, and an update nearly two decades later, with more than four times the studies, confirmed it, with the strongest gains exactly where interviews need them, in matching ratings to the right dimensions rather than to a general glow (Roch et al., 2012).
An honest caveat, in the spirit of this site's evidence policy: much of that research measures rating accuracy in controlled tasks rather than end-to-end hiring validity. The direction is consistent and the mechanism is exactly the one interviews depend on, but nobody should quote you a precise validity gain from a calibration hour.
If frame-of-reference training sounds familiar, it is because you have already built its materials. The behavioral anchors attached to every question are a frame of reference in writing: a weak, an acceptable, and an excellent answer, described before anyone is in the room. Google's version of the same idea is standardized rubrics "so that all reviewers have a shared understanding of what outstanding, solid, borderline, and poor response looks like", backed by interviewer training and calibration "so that interviewers are confident and consistent in their assessments". Writing the anchors was the first half of the work. Calibration is the second half: practicing against them until two interviewers reading the same answer land on the same score.
What a full program covers
For the complete version, the OPM guide ships an actual lesson plan, and its shape is worth stealing even at a fraction of the size. Three blocks. First, why: what structure buys in reliability, validity, and legal defensibility. Second, the materials: the competencies and the job they came from, the questions, the anchors, the rating forms, and the biases and rating errors that pull scores off target. Third, practice: critiqued run-throughs against recorded interviews before anyone does it live.
The error catalog in the second block is the one piece of theory worth everyone's time, because every entry is a way an untrained interviewer feels accurate while drifting:
| The error | What it sounds like in your head |
|---|---|
| First impression | "Knew within two minutes." The remaining fifty-eight confirm it. |
| Halo | "Great answer on leadership, so I nudged everything else up." |
| Contrast | "After yesterday's disaster, this candidate felt like a star." |
| Similar to me | "We just clicked." You share a hometown, not a competency. |
| Leniency / strictness | Every candidate a 4; or nobody ever earns one. |
| Central tendency | All 3s. The scale is a place to hide, not a measurement. |
Naming the errors does not cure them; awareness alone rarely does. The structure is the cure, and the errors are the argument for it: scoring each answer against its anchors, right after it is given, from notes, is what leaves first impressions and halos nowhere to operate. Train the catalog so people stop trusting the feeling of certainty, then train the procedure that replaces it.
The minimum viable version for a small team
You are not the federal government. Five people are hiring a support lead; nobody has run a training program and nobody is going to. Here is the version that respects that, built from the pieces above. Budget half a day in total, most of it inside work that was happening anyway.
Set one rule first: nobody interviews alone before training. Not as ceremony but as a license, the same way nobody touches production on day one. The rule is what makes the rest happen, because the afternoon it costs is now a requirement instead of a nice-to-have that loses to the calendar.
Pre-read, thirty minutes, solo. Each interviewer reads the interview guide for the role end to end, plus the case for structure if they have never met it. The guide is most of the curriculum: the questions, the anchors, the probes, the schedule, in the order they will be used. A first-time interviewer holding a good guide is already most of the way to defensible.
One ninety-minute session, the whole hiring loop together. Walk the guide once, so every interviewer knows which competencies are theirs and which belong to another stage. Spend ten minutes on the error table above, mostly so the words exist in the team's vocabulary when someone says "I might be haloing here". Then spend the rest on calibration, which works like this. Take one question card from the actual interview. One person answers it in character as a plausible real candidate, imperfections included; two minutes, no theatrics. Everyone else takes notes and scores the answer independently against the card's anchors, no talking. Reveal the scores. The gaps are the material: argue about what the anchors mean, not about who is right, and sharpen the anchor wording where it turns out to be readable two ways. Run it two or three times and watch the spread tighten. That convergence is frame-of-reference training, homemade, and it is the highest-value hour of the afternoon.
A card like this one is a self-contained calibration kit: a question that invites an imperfect, arguable answer, anchors to score it against, red flags to catch, and probes for the role-player to make necessary. Score the answer, compare, argue, repeat.
Shadow, then reverse-shadow. The new interviewer sits in on one live interview, watching how an experienced colleague runs the room, holds the script, and probes on deficiency rather than charm. Next round they swap: the new interviewer runs it, the experienced one observes, both score independently, and afterwards they compare notes and ratings against the anchors. That comparison is calibration again, on live material.
Re-calibrate a little, every loop. Training decays; scores drift. After each hiring round closes, when decisions are done and discussion is finally allowed, spend fifteen minutes comparing how the loop scored: any interviewer running notably hot or cold, any question whose anchors kept splitting readings. Fold what you learn into the guide, which is where this team's training actually lives.
The honest audit of that program: it covers the OPM lesson plan's three blocks at perhaps a twentieth of the cost, it keeps the best-evidenced piece, calibration, close to the form the research studied, and the license rule keeps it from evaporating. It will not certify anyone. It will make two of your interviewers score the same answer the same way, and that, more than any question you could buy, is where the .42 lives.
The cook, not the cookbook
Structure has a last mile, and the last mile is a person in a room making dozens of small decisions at conversational speed. Training is how those decisions get made the same way twice, which is all measurement has ever meant. The recipe was never the hard part; every kitchen that plates the same dish twice got there the same way, by practicing until the hands agree.
The materials are the good news. Every question card in the library already carries the anchors and probes this article keeps pointing at, so your calibration kit is written before the session is scheduled. Pull the cards into a guide, or let the planner assemble one for the role in about three minutes, free, no account needed until you print it, and you have the full curriculum for the afternoon: one guide, one error table, one teammate willing to play a mediocre candidate.
Next in the path: setting up the interview, the logistics that quietly shape validity before anyone says a word.