The phrase means fifteen things. Most of us do three of them.
Over the last few months I have asked a fair number of hiring teams whether they run structured interviews. Almost all say yes without much hesitation. Ask what they mean by it and the confidence drops. Usually it comes down to everybody getting the same questions. Sometimes there is no clear answer at all.
In one of the organisations I worked in, four of us would interview a candidate and then each fill in a form HR had designed. Five point scale, one score per parameter. Printed, uniform, filed. It looked like structure.
The form asked for a rating and never asked why. No evidence, no example, no space to write what the candidate had actually said that produced a 3 rather than a 4. A number with no reasoning behind it is not a rating. It is an opinion wearing a uniform.
The second problem was worse. We filled the forms after interviewing all the candidates, not after each one. The scores were never assessments of individuals at all. They were a comparison we had already made in our heads, written down afterwards in a format that made it look deliberate.
That process would have called itself structured. Of the fifteen components the research describes, it was failing at least six.
Fifteen components, in two groups
Campion, Palmer and Campion reviewed the interview literature in 1997 and did something more useful than recommending structure. They took it apart.
Seven components structure the content: base questions on a job analysis, ask everyone the same questions, use situational or past-behaviour questions rather than general ones, limit improvised follow-ups, ask more questions rather than fewer, control what other information the interviewer sees, and hold candidate questions until the end.
The other eight structure the evaluation, and this is the half that gets skipped: rate each answer separately rather than forming an overall impression, use scales with behavioural anchors, take notes as evidence rather than as reminders, use multiple interviewers and keep them consistent across candidates, do not discuss candidates between rounds, train the interviewers, and combine the ratings by a rule rather than by discussion.
Run your own process against that list. Most organisations that call themselves structured have done three or four items from the content group and very little from the evaluation group. Question bank, yes. Anchored scales, no. And the combination step is a conversation, always.
I have since run that list against processes I have been responsible for, including our own. Nobody comes out clean. The content side is usually respectable. The evaluation side is where everyone, including us, has work to do.
One skipped component is worth naming, because it costs nothing. Giving an interviewer the CV, the test scores and the previous round's feedback before their own round contaminates that round, since the interviewer stops assessing and starts confirming. I did it for years. I would open the file the evening before and tell myself I was preparing. What I was doing was deciding early and then spending the interview looking for support.
What structure actually buys
Levashina and colleagues reviewed decades of this work in 2014. Beyond the expected gains in reliability and validity, two findings stand out.
Structured interviews are harder to game. Rehearsed impression management works beautifully on a free-flowing conversation and much less well on a question tied to a defined competency and scored against an anchor. If you have wondered why polished candidates outperform their eventual job performance, that is where I would look first.
And candidates mind structure far less than interviewers assume. The resistance is almost entirely internal. We tell ourselves candidates want a natural conversation. Mostly we want it, because it is more pleasant to conduct.
Though I am not sure that second finding travels as well as it reads. A good backend engineer here is running four or five processes at once and is evaluating you at least as hard as you are evaluating her. A rigidly structured interview can read as a company processing people rather than recruiting them, and I have seen candidates disengage for exactly that reason. Structure the evaluation ruthlessly. But someone, usually the hiring manager, has to spend unstructured time selling the role, and that conversation should not pretend to be an assessment.
The numbers moved in 2022
For twenty years the field ran on Schmidt and Hunter's 1998 meta-analysis, which put general mental ability at the centre of selection at a validity of about .51.
Then Sackett, Zhang, Berry and Lievens showed that a range restriction correction had been applied inappropriately across decades of meta-analyses, and re-estimated the lot. General mental ability fell to about .31. Structured interviews landed at about .42, which puts them at the top of the table, above cognitive tests and work samples.
The comparison that matters most is the internal one. Structured interviews, around .42. Unstructured interviews, around .19. Roughly half the predictive power from the same hour with the same candidate, decided entirely by how the hour was organised.
The number behind the number
Sackett's team also reported how much these estimates vary, and structured interviews vary enormously. The 80 percent credibility interval runs from about .18 to .66. Their own phrasing is .42 plus or minus .24.
At the top of that range, a structured interview is the most powerful instrument available to us. At the bottom, it is barely distinguishable from a chat.
So the label is doing no work at all. What decides where you land is which of the fifteen you have implemented and how well. Saying you run structured interviews tells a candidate, a client or an investor nothing. Naming the seven you have implemented tells them a great deal.
What to change
Audit yourself against the fifteen. An hour with the list, each component marked done, partly done, or not done. Free, and more informative than any vendor demo.
Anchor your scales. A five point scale with no anchors measures the interviewer's generosity, not the candidate's ability. Write out what a 2 and a 4 look like for this role, in observable terms.
Demand evidence with every rating. If an interviewer cannot write the sentence that produced the score, the score should not count.
Rate immediately, never at the end of the batch. Scoring five people against each other from memory records a comparison, not an assessment.
Stop sending the panel the full file in advance. Each interviewer gets the role definition and their own competencies. Everything else circulates after ratings are in.
What the label is worth
The structured interview is not a compliance exercise adopted for legal cover. On current estimates it is the most predictive instrument we have, which is a surprising place for the field to have landed after decades of chasing better tests.
But it earns that number only when it is built properly. A five point scale on a printed form, filled in at the end of the week, is not a structured interview. I know, because I filled in plenty of them.
This is why Sainterview conducts the interview rather than merely recording one. Same competencies for every candidate, anchored scales instead of bare numbers, each answer rated separately and against evidence, at the time rather than at the end. An attempt to sit nearer the .66 end of that range than the .18 end, which is a distinction the word "structured" is far too generous to make on its own.
Next issue: What research tells us about variability in hiring decisions
The research behind this article: Campion, Palmer & Campion (1997), A Review of Structure in the Selection Interview, Personnel Psychology 50(3), 655 to 702. Levashina, Hartwell, Morgeson & Campion (2014), The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature, Personnel Psychology 67(1), 241 to 293. Schmidt & Hunter (1998), The Validity and Utility of Selection Methods in Personnel Psychology, Psychological Bulletin 124(2), 262 to 274. Sackett, Zhang, Berry & Lievens (2022), Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range, Journal of Applied Psychology 107(11), 2040 to 2068.