A few years ago I was interviewing a candidate for a sales executive role from my apartment. Ten minutes in, there was a commotion outside the building. Then my phone rang. I let it ring. It rang again. On the third call I apologized to the candidate and picked up, because whoever it was clearly had no plans to stop. It was Ikea. They had delivered my sofa, they said, and left it in the parking lot. In the middle of the parking lot, as it turned out, in several very heavy pieces, in the middle of a Finnish winter, blocking my neighbours' cars. I excused myself from a sales executive and went to carry a sofa.
I tell that story against myself, and it deserves it, but it is also this whole article in miniature. Nothing about my interview guide was wrong that day. The questions were good, the anchors were written, I knew what I was measuring. What failed was everything around the instrument: the room, the schedule, the phone, and a delivery booked without looking at my own calendar. Preparation on the interviewer's side is invisible when it works and only occasionally involves furniture. When it fails, it fails quietly, and the candidate pays for it in a score that was never really about them.
This article is about how to prepare for an interview as the interviewer. Everything earlier in this path was built at a desk: the analysis, the question types, the anchors, the probes, the guide that holds them. Training put the instrument into people's hands. Setting up the interview is the last step before anyone says a word, and it comes down to four decisions that quietly shape what the interview measures: who is in the room, what they have read before it, how the day is scheduled, and what the candidate has been told. Cooks call this mise en place, everything in its place before the heat goes on. Interviewers mostly call it admin, and that is the mistake.
Interviewer preparation is part of the instrument
The classic taxonomy of interview structure (Campion, Palmer and Campion, 1997) lists fifteen components that make interviews predictive. Read that list with a calendar in one hand and a surprising number of them turn out to be logistics. Use multiple interviewers. Use the same interviewers across all candidates. Control the ancillary information interviewers see, the résumé and the test scores. Hold candidate questions until the end. Make the interview long enough to ask enough questions. Do not discuss candidates between interviews. None of those is decided in the room. They are decided in the invitation, the calendar, and whatever lands in the interviewer's inbox the night before.
Structured interviews predict job performance at about .42 on average, roughly double unstructured ones, but that average sits on a spread wide enough that its bottom reaches unstructured territory (Sackett et al., 2022). The training article argued that the spread lives in behavior. Part of it lives one step earlier, in setup, because setup decides which behaviors are even available: an interviewer who has read the screening score before the call cannot un-know it, and a loop that seats a different colleague in the second chair for every candidate cannot be calibrated into consistency afterwards. The good news is symmetrical. These are the cheapest components of structure to get right, because each costs one decision and no talent at all.
Who should be in the room: panel size and the Rule of Four
The anxious instinct is to add people. The hire matters, so four colleagues crowd into the room, or a fifth round appears on the loop. The evidence is unusually clear that this buys less than it feels like. A meta-analysis comparing interviewer practices across studies found that panel interviews were no more valid than one-on-one interviews; what did come with better prediction was training, note-taking, and using the same interviewers across all candidates (Huffcutt and Woehr, 1999). Google reached the same place from its own hiring data with its Rule of Four: after four structured interviews, additional interviewers barely moved the decision (Google re:Work; Bock, Work Rules!).
The US federal government's structured interview guide lands on a practical number: "a panel of two or three interviewers may be better able to document and interpret the information". Two is my default for a competency interview, and the two have jobs. One leads, holding the guide, asking the questions and their probes. The other takes the notes and watches the clock. Both score every answer independently against the anchors before either says a word about the candidate. A third interviewer adds a little reliability and a lot of calendar; a fourth silent observer adds only the observer.
Bias is often the reason a panel gets built, so one finding is worth knowing: in structured interviews, race and gender similarity between interviewer and candidate has small or no effect on ratings, one-on-one or in panels of two, three, or four (Levashina et al., 2014). It is the structure doing the protecting, not the headcount. A diverse panel is good for other reasons, but it is not a substitute for the same questions, anchored scoring, and independent ratings.
The component that gets broken most often is the quiet one: the same interviewers for every candidate. The OPM guide puts it plainly: "when feasible, the same interviewers should be used (either in a panel or serially) across all candidates, to ensure consistency in ratings." In practice a loop is a fixed cast: the recruiter, then the hiring manager with a peer, then the skip-level with someone from the team the role serves. When a cast member cannot make a date, move the interview rather than the interviewer. Competencies belong to stages, two or three each, as the guide article laid out, so the cast is also a division of labor: each pair goes deep on its own competencies and stays out of the others'.
What to read before the interview, and what not to
Here is the component nobody likes: control ancillary information. The taxonomy names résumés and test scores specifically, and behind it sits a body of research on what interviewers do with an impression formed before the interview begins. In a field study of real employment interviews, interviewers who had formed a favorable pre-interview impression from the application and test scores behaved differently once the candidate was in front of them. They showed more "positive regard toward applicants", did more "'selling' the company and giving job information", and gathered less information (Dougherty, Turban and Callender, 1994). The candidates responded in kind, with warmer rapport and a different communication style. The interview had become a confirmation of a decision made by a piece of paper.
That is the résumé problem in its honest form. The trouble is not that the résumé is uninformative; it is that it is informative early, and early information sets the questioning strategy. A liked candidate gets sold to. A doubted candidate gets probed. Neither gets the same interview, which was the whole point of building one.
Two habits fix most of it. First, read the guide, not the résumé. The night-before preparation for a competency interview is the guide for the role: the questions, their probes, and the anchors you will score against. You do not need the candidate's history to run that instrument; you need the instrument. Second, if the loop needs somebody to walk the résumé, and it usually does, for dates, gaps, and what the person actually did in each role, give that job to one stage, typically the recruiter screen, and keep it out of the competency interviews. The federal guide goes as far as saying that a candidate's own supplemental documents, résumé included, are "for the candidate's reference only and should not be looked at by the interviewer during the interview".
The same logic applies to information from inside your own process. The screening score, the take-home result, and the previous stage's rating are all ancillary information for the next interviewer, and the fifteen components include not discussing candidates between interviews for exactly this reason. A loop where stage two hears "she was great" from stage one in the corridor has stopped being three independent measurements and become one measurement with two echoes. Route those signals to the debrief, where they belong, and let each stage score blind.
Levashina and colleagues add a subtler point about the first five minutes. Interviewers "may form early impressions based on the rapport building phase where nonjob-related information is exchanged", so the small talk about the commute is ancillary information too. Their suggestion is to standardize it, and the OPM guide's opening script does exactly that: welcome the candidate warmly, describe the job briefly, explain the process the same way every time, say that notes will be taken, and hold substantive candidate questions until the end, which is another of the fifteen components. Write the opening into the guide so the rapport phase has the same length and shape for every candidate. What happens once the first real question lands is the next article's subject.
Scheduling interviews: the calendar is part of the instrument
My other recurring preparation failure was more mundane than a sofa. In my recruiting years I would sometimes stack seven interviews into one day, because the calendar allowed it and the pipeline demanded it, and by the fifth I was greeting candidates with the previous candidate's name. That is embarrassing. What the research says happens by the fifth interview is worse, because nobody notices.
A study of more than 9,000 MBA admissions interviews over ten years found that an interviewer's score for a candidate depended on the candidates they had already seen that day (Simonsohn and Gino, 2013). Interviewers behaved as if each day should produce roughly the expected spread of scores, so one who had already recommended three strong applicants in a morning became reluctant to recommend a fourth, however strong. The effect was not trivial: by the authors' estimate, an applicant following a run of high scores needed the equivalent of about 30 more GMAT points, or nearly two more years of experience, to land the score a quieter day would have given them. Contrast effects sit alongside this. In the classic laboratory demonstration, an average candidate rated after two strong candidates was marked down and after two weak ones marked up, and the distortion was largest for exactly the middling candidates most hiring decisions hinge on (Wexley et al., 1972).
Neither of these is a character flaw. They are what a tired human does with a day of judgments, and the countermeasure is scheduling.
The rules that follow are cheap:
- Cap the day. Three interviews per interviewer per day is comfortable; four is the ceiling. Past that you are no longer measuring the fifth candidate. You are measuring your afternoon.
- Schedule the scoring, not just the interview. Block fifteen minutes after every interview for each interviewer to finish their notes and score every answer against its anchors while it is fresh. A scale filled in from memory at dinner is a scheduling failure before it is a discipline failure.
- Keep the slots identical. The federal guide's rule is that "all candidates should be allotted the same amount of interview time". Same length, same question set, same order, the same short buffer at the end for the candidate's questions. A forty-minute interview for one candidate and a sixty-minute one for another are two different instruments.
- Give interviewers notice. OPM again: interviews "should be scheduled far enough in advance to provide adequate preparation time". A guide that arrives ten minutes before the call gets skimmed, and a skimmed guide gets abandoned by question three.
- Put the debrief after the last candidate. No corridor comparisons between interviews; ideally the calendar does not even make them possible. One debrief, once every candidate's independent scores are in, is where discussion belongs.
Length is a scheduling decision too, which is why "longer interviews, more questions" is on the taxonomy's list. Two or three competencies per stage, two questions each with their probes, at six to eight minutes a question, is forty-five minutes of questions before the opening and the candidate's own questions. Book an hour. A thirty-minute slot for that guide does not produce a shorter structured interview. It produces an unstructured one that started out structured.
The room, the video call, and the medium
The OPM guide's requirements for the setting read like common sense until you count how many I have broken. Interviews "should be held in a quiet, non-threatening, and private place". "Seating arrangements should be the same for all candidates." The room and facilities "must be accessible to candidates with disabilities". Waiting candidates get somewhere separate to wait, away from the candidates who have just finished. Add one item from experience: the phone goes on silent, and the delivery window goes on another day.
Most interviews I run now happen on video, and the evidence on that has moved. The first meta-analysis of technology-mediated interviews found that candidates interviewed by phone or video received lower ratings (d = −.41) and reacted less favorably (d = −.36) than candidates interviewed face to face, across twelve studies (Blacksmith, Willford and Behrend, 2016). A much larger update, 31 samples, published in 2026, paints a more usable picture: interview ratings no longer differed meaningfully by medium, but candidates still found the organization less attractive after a technology-mediated interview and were somewhat less inclined to pursue the job, and both effects were worse for asynchronous, one-way video than for a live call (Moon, Park and Park, 2026).
The practical reading is that a live video interview, run well, measures about as well as a room. Three cautions follow. First, pick one medium per stage and use it for every candidate; the 2016 authors warn that varying the medium across applicants "could lead to fairness issues". Second, structure matters more on video, not less. Levashina and colleagues note that these formats deny interviewers the full range of cues and "may benefit from structuring even more than face to face interviews", and their own content analysis found video interviews using far fewer structure components than face-to-face ones. A video call is where the guide earns its keep. Third, keep a human in the competency interview. One-way, pre-recorded formats are where candidate reactions fall furthest, and a candidate who withdraws is a data point you never get to score. If you record the call, ask in advance; the rules differ by country and by US state.
What to tell candidates before the interview
Sackett and colleagues, the authors behind the validity numbers this site is built on, are careful about what those numbers do not settle. Their follow-up paper on selection design says that "integration of multiple outcomes, such as validity, cost, time constraints, testing volume, subgroup differences, and applicant reactions, is needed for an informed decision" (Sackett et al., 2023). Applicant reactions are on that list on purpose, and structured interviews have a known weakness there: applicants find them harder than unstructured conversations, react a little less warmly, and can come away with a slightly less positive view of the organization (Levashina et al., 2014). Rigor reads as cold when nobody explains it.
The fix is cheap and, surprisingly, it improves the measurement. In the same review, giving applicants advance knowledge of the questions produced significantly higher perceptions of fairness. And in two studies of transparent interviews, where interviewees were told which dimensions they were being assessed on, interview performance improved, construct validity became satisfactory where the nontransparent version's was not, and criterion-related validity did not fall (Klehe et al., 2008). Telling candidates what you are measuring does not spoil the test. It gets more of them taking the test you designed instead of the one they guessed at.
So the invitation does real work. Send, in writing, before every stage:
- The shape. How long, who they will meet (names and roles), and how many stages the loop has.
- The format. That it is a structured interview, that they will be asked for specific examples from their own experience, and that every candidate gets the same questions.
- What you are measuring. The competencies for that stage, in plain language. Not the anchors; the anchors are yours.
- The mechanics. That notes will be taken, whether the call is recorded, and that there is time at the end for their questions.
- The practicalities. Where, which link, what to bring, and how to ask for an accommodation.
That line about anchors deserves a look at an actual card, because a card is where transparent ends and leaked begins:
The competency it sits under and the general territory of the question are safe to share, and sharing them is what the transparency studies did. The rest of the card, what a strong answer contains and which flags are red, stays in the guide. A candidate told "we will ask about decisions you made under time pressure" arrives with a real example instead of improvising one on the spot, which is exactly the answer you want to score.
There is a recruiting return on this too. In Greenhouse's 2026 survey of about 3,000 candidates, 38% said they had walked away from a hiring process because it included an AI interview and another 12% said they would, with pre-recorded video scored by a machine among the most-cited triggers; in the previous year's report, a fifth of candidates had turned down an offer over a poor interview experience. A prepared human interviewer who told the candidate what to expect and then did exactly that has become a differentiator. Structure, explained, is a candidate experience, not a tax on one.
The interviewer's preparation checklist
Everything above, compressed to what fits on the back of the guide.
The week before
- Fix the cast: the same interviewers per stage for every candidate, competencies assigned to stages, one lead and one note-taker per room.
- Send the invitation: shape, format, competencies, mechanics, practicalities.
- Book identical slots, at most three or four per interviewer per day, a fifteen-minute scoring block after each, one debrief after the last candidate.
The day before
- Read the guide, not the résumé: your competencies, questions, probes, and anchors.
- Confirm the room or the link, the recording agreement, any accommodation, and how notes will be taken.
Ten minutes before
- Guide open, note space ready, phone silent, door closed or camera on and notifications off.
- Opening script in front of you: welcome, the job in two sentences, the format, the notes, any questions, begin.
Everything in its place
I did carry the sofa, eventually, and the candidate was gracious when we resumed. I have no idea whether she was any good. Nobody who spent half an interview thinking about a parking lot could tell you, and that is the real cost of a setup failure: not the embarrassment, but a measurement that never happened. Mise en place is the least glamorous part of any kitchen, and it is the reason the plates go out on time.
The materials are the easy part. Every question card in the library carries the competency you can share with the candidate and the anchors you keep for yourself, and the planner turns a role into one focused interview, its questions, probes and anchors in a timed running order, in about three minutes, free, no account needed until you save it. Once it is saved, the job page proposes how the rest of the analysis splits across the other interviews in the loop. Fix the cast, send the invitation, book the hour and the fifteen minutes after it, read the guide instead of the résumé. Then close the door.
Next in the path: running the interview, how to stay structured once a live human starts talking.