We structure everything in hiring except the one conversation that decides the outcome.
Look at the shape of a typical hiring process. A screened CV. An online coding test. Four interviews with prepared questions. A scorecard mapped to competencies. Hours of effort, most of it reasonably well designed.
Then thirty minutes are blocked for the debrief, it finishes in eleven, and somebody asks: so, what do we think?
That final meeting has no structure at all. No agreed criteria, no order of speaking, no record of who disagreed or why. It is the one step where we abandon method and go back to pure conversation. It is also the step that produces the offer.
That is the argument of this issue. Whatever structure you build into your hiring process, it gets undone at the last step, because the last step is the one nobody has bothered to structure.
Two things rush in to fill that vacuum.
The first is noise
A senior backend developer, call her R, was interviewed in one of my earlier employer companies. Strong coding test, four rounds, all of which ran past their slot. At the debrief the engineering manager spoke first, as the senior-most person usually does, and said he liked her. Clean thinker. The second interviewer had not seen enough depth on system design. The third was not convinced she would gel with the team. The fourth had the scorecard open and pointed out that her machine coding score was among the best that quarter.
Four capable engineers, one candidate, four incompatible versions of her.
Nothing about the evidence changed between 10:00 and 10:11. What changed was who spoke first, how confidently they spoke, and what each person happened to remember.
Researchers call this noise, meaning unwanted variability in judgement. If you weighed the same object on four scales and got four different readings, you would not sit down for a philosophical discussion on the nature of weight. You would get the scales serviced. When four interviewers produce four different readings of the same person, we call it judgement and put it in an offer letter.
The second is rank
Some years ago, in an organisation I no longer work for, I sat on a committee deciding promotions for a batch of officers. One name belonged to a team whose business head was several rungs senior to me. She arrived with her view already formed and no visible interest in testing it.
I held a different view of that officer and I had my reasons. I got about ninety seconds. What followed was not an argument, because an argument needs two sides. It was a position, stated with the weight of rank, followed by a silence that everyone correctly read as the end of the discussion. I stopped talking. The promotion went through.
I am not sure I was right, and that is rather the point. What stayed with me is what the room did with my disagreement. It did not examine it. It did not record it. It absorbed it.
Notice where this can happen and where it cannot. Nobody bulldozes a scorecard. They bulldoze a conversation. Rank can only operate in the unstructured part of the process, which is precisely why it shows up at the debrief.
I should add that I have been on the other side of this. In the last article I wrote about my habit of deciding on a candidate in the first few minutes. What I left out is that I do not merely hold that early view. I argue for it, fluently and with seniority, and the room quietly moves towards a conclusion I had reached before I had any evidence.
The finding nobody in hiring wants to hear
There is a question that was asked seventy years ago and answered fairly conclusively, and our industry has been politely ignoring it ever since.
Paul Meehl asked in 1954 whether expert judgement or a simple formula does better at combining evidence into a prediction. Grove and colleagues settled it with a meta-analysis of 136 studies in 2000. Mechanical combination was about ten percent more accurate on average. It was substantially better than expert judgement in a third to a half of the studies, and substantially worse in only six to sixteen percent.
The detail that should give every hiring manager pause is what did not matter. The advantage held regardless of the judgement task, the type of data, and the amount of experience the judges had. Experience did not close the gap.
Now consider what a hiring debrief actually is. Four people arrive holding separate pieces of evidence and combine them by talking until a conclusion emerges. That is textbook holistic combination, which is the method that loses.
This is not a fringe position in the selection literature either. When Campion, Palmer and Campion catalogued the fifteen components of a structured interview in 1997, the last one on their list was combining ratings by a rule rather than by discussion. It is also, in my experience, the single component that organisations are most likely to skip. We are happy to standardise the questions. We are extremely reluctant to standardise the verdict.
None of this says a formula should make the hire. Meehl himself allowed for what he called the broken leg case, where a human knows something the formula cannot see and should override it. The problem, borne out repeatedly since, is that we believe we have spotted a broken leg far more often than we actually have. The discipline is not in refusing to override. It is in requiring the override to be stated as a reason and written down.
Which suggests a different sequence for the eleven minute meeting. Combine the ratings first, mechanically, before anyone speaks. Then let the panel look at what came out and argue with it. The discussion becomes an audit of a result rather than the method for producing one.
What to change in that meeting
Collect ratings and evidence before the discussion opens, with no exceptions for senior panelists. This is the one that does most of the work. Independent judgement first, debate second.
Open with the combined result, not with opinions. Compute the aggregate from the submitted ratings and put it on screen before anybody speaks. The meeting then starts from a number and a set of competency-level gaps rather than from whoever is most senior.
Spend the time on disagreement, not on agreement. If all four rated coding strongly, there is nothing to discuss. The entire value of the meeting sits in the dimensions where the ratings diverged.
Junior-most person speaks first on each dimension. If the EM's view lands first, everything that follows is confirmation.
Require overrides to be reasoned. Anyone arguing against the combined evidence should be able to say what they know that the evidence does not capture. Sometimes that is real and important. Often, saying it out loud is enough to reveal that it is not.
Find a way to keep dissent on the record. How formally you do this needs thought, since records of this kind can be put to uses nobody intended. But a room that leaves no trace of its disagreements cannot learn from them either.
The point
Disagreement in a hiring panel is not a problem. It is usually the most valuable thing the panel produces. The problem is a room that cannot say why it disagreed, and has no way of placing a senior view next to a junior one and asking which of them the evidence supports.
Structure is not there to remove judgement, or to reduce a person to a score. It is there to make the evidence behind a judgement clear enough that it has to be argued with.
We have spent twenty years structuring the interview. The decision is still a conversation.
This is not an abstract interest of mine. It is why we built FitSuite around the handover rather than around any single instrument. HireMatrix, the layer that consolidates the evidence, exists to do the thing the research has been pointing at since Meehl: combine the ratings before the discussion begins, hold every piece of evidence against the competencies the role actually requires, and show where that evidence is strong, where it is thin, and where two evaluators looked at the same answer and read it differently.
It does not make the decision. I would be suspicious of any system that claimed otherwise. What it does is ensure the panel walks into that eleven minute meeting with the evidence already combined, so that the conversation becomes an audit of a result rather than the method for producing one.
Next issue: almost every company says it runs structured interviews. Almost none of them mean the same thing by it, and the research suggests the difference is worth about half your predictive power.
The research behind this article: Meehl (1954), Clinical Versus Statistical Prediction, University of Minnesota Press. Grove, Zald, Lebow, Snitz & Nelson (2000), Clinical Versus Mechanical Prediction: A Meta-Analysis, Psychological Assessment 12(1), 19 to 30. Campion, Palmer & Campion (1997), A Review of Structure in the Selection Interview, Personnel Psychology 50(3), 655 to 702. Highhouse, Nye, Zhang & Rada (2023), Improving Workplace Judgments by Reducing Noise, Annual Review of Organizational Psychology and Organizational Behavior. Kahneman, Sibony & Sunstein (2021), Noise: A Flaw in Human Judgment, William Collins.