62%. That’s how much of the variance in a large set of manager ratings came from the people doing the rating, in a study published in the Journal of Applied Psychology in 2000. The actual performance of the managers being rated explained 21%.
I think about that number every time a hiring manager tells me the candidates aren’t good enough. “The quality just isn’t there” is a sentence I’ve heard more times than I can count in more than 20 years of recruiting, in markets where talent was scarce and in markets where it supposedly wasn’t. Early in my career I took that sentence at face value, as a report on the market. Over the years I’ve come to read it differently: as a report on how the judging is set up, and on how far any reviewer’s bar can move from one week to the next.
It’s a hard point to raise with someone whose req you’re trying to fill, so most recruiters skip it and source harder. They send another batch, collect another round of rejections, and slowly start to believe the market is broken.
More volume can make it worse, since every new batch gets judged by the same moving bar, and a bigger pile of rejections starts to look like proof.
Sponsored by Omni Sidekick
Stop retyping the same tone every time you write online. Omni Sidekick is a Chrome extension that helps you write and rewrite your content and replies, in your voice, not generic AI-speak. It also lets you screenshot anything on a page and hide distracting images and videos as you scroll (LinkedIn included), so you can actually focus on what matters. Get instant help right where you’re typing, no tab-switching required.
62% of the score
The study is by Steven Scullen, Michael Mount and Maynard Goff. They took two large samples of managers, each person rated by several bosses, peers and subordinates, and split the variance in those ratings into its sources. The rater’s own tendencies (how harsh they were, what they paid attention to, how they used the scale) explained 62% of the variance in one sample and 53% in the other. The performance of the person being rated explained 21% and 25%.
Those were performance ratings, and I’m stretching the finding by carrying it over to hiring. I know that. The raters in that study had watched these managers work, often for months or years. A hiring manager reading a resume has two pages and maybe 90 seconds of attention.
I haven’t seen a study that measures the rater’s share for resume screens specifically. My guess is that it gets bigger, because the less evidence there is on the page, the more of the verdict has to come from the person reading it. That’s a guess, and I’d be happy to see data that proves me wrong.
What moves a hiring manager’s bar is usually ordinary. Two people resign from their team in the same month, and every profile starts to look like another way things could go wrong. The last hire didn’t work out, so anything that looks like that person gets marked down. A strong profile lands at the top of the batch and the next five look thin next to it. None of this shows up in the ATS. The notes just say “not senior enough” or “not a fit,” and those notes get read later as facts about the candidate.
A small aside about the word “senior.” Having hired across three regions, I’ve stopped trusting it completely. A senior in one country has four years of experience, and in another she has twelve and would find the four-year version slightly insulting. Job titles drift the same way inside companies, where “senior” often means “has been here a while.” I bring it up because hiring managers lean on the word more than on almost anything else in their feedback, and it carries less information than they think it does.
AI makes the word even shakier. When tools handle much of the recall and speed that used to signal seniority, years of experience start measuring exposure more than judgment, and a resume shows the first far more clearly than the second.
Two weeks ago, I spoke with a recruiter friend who was looking for a coaching session about why his hiring managers were rejecting candidates. The first thing I asked him to do was pull a month of rejection notes for a single req and read them in one sitting. Date order, no candidate names. The pattern showed up fast.
A handful of phrases kept repeating (“not senior enough,” “lacks depth,” “not a culture fit,” “just okay”), and none of them told him much about any particular person. Mixed in were a few specific notes that told him a lot. Those were the useful ones. The repeated phrases mostly reflected what kind of week the hiring manager was having when they wrote them, and sure enough, they clustered in the same stretch of days.
When a hiring manager tells me the candidates are weak, I ask a boring follow-up question: weak compared to whom? The answer is almost never a person who is on the market right now.
The candidate who lives in the hiring manager’s head
Usually it’s someone they used to work with. The best engineer from their last company, remembered at her peak, carrying six years of context about a codebase that nobody could bring into an interview. Or it’s the person who just resigned, which is worse, because the hiring manager remembers everything that person knew and none of the months it took them to learn it.
The job description gets written from that memory, and then it grows. Someone adds a tool the team plans to adopt next year. Someone else adds “experience in a regulated industry” because of one painful audit. By the time the ad goes live it describes a person who probably exists, but who is employed, well paid and not reading job ads.
Then a recruiter or software filters the applications against that list.
88%. That’s the share of employers in the Hidden Workers study by Harvard Business School and Accenture who agreed that qualified high-skills candidates get screened out because they don’t match the exact criteria in the job description. For middle-skills roles the figure was 94%. These are employers describing their own process, which is the part I find hard to get past.
The authors estimated that about 27 million people in the US alone fall into this “hidden” group. Companies that said they were open to hiring from it reported being 36% less likely to face talent shortages, though that’s what the companies reported, and it’s a correlation, so I wouldn’t treat it as a promise.
So the hiring manager ends up looking at a pool that survived a filter built from a memory, and judging it against the same memory. Of course it looks thin.
I used to think a better intake meeting fixed this. For years my advice to recruiters was to take a full 45-60 minutes for intake and push back on every must-have, and I still think both are worth doing. I’ve changed my mind about how much they help, though.
An intake meeting produces words, and words like “senior” or “strong communicator” mean something different to every person at the table. Everyone nods, and each of them pictures a different person. You only find out what the hiring manager means when you put a real profile in front of them and watch where they slow down.
Calibration can go wrong too.
The same forces act on recruiters. We get tired, we get anchored by the first strong profile of the day, and we screen differently on a Friday afternoon than on a Monday morning. That’s why I’d apply the check to recruiter screening as well as to hiring manager feedback. The packet works best when both sides rank it and compare notes, because then it reads as a shared standard instead of a test aimed at one person.
I’m also not sure this argument holds in thin markets. If you need an embedded engineer with a specific safety certification who speaks the local language, there may be a dozen plausible people in the country, and most of them are happy where they are. A perfectly calibrated hiring manager still ends up with a shortlist of two.
Calibration gets people to agree with each other, and a group can agree beautifully around the wrong picture of the job. When that happens, the rejections stop being random and start being consistent, and I can’t tell which of those is worse for the candidates on the receiving end.
Then there are occasional interviewers, the architect or finance manager who interviews twice a year and no longer remembers what a normal answer sounds like, so their bar is set by whoever they happened to see last.
Six blind profiles before sourcing starts
The check you can run on any hard‑to‑fill technical or product role takes one hour in the first week and only a few minutes in the third.
Before sourcing anyone, build a packet of six real profiles. Pull them from earlier pipelines for similar roles and strip out names, photos and school names, because school names carry more weight in a quick read than most reviewers intend. Choose them to cover a spread: two you think are strong, two borderline, one clear no, and one wildcard from a different industry.
In week one, sit down with the hiring manager. They rank the six and write one sentence per profile explaining the decision, before you discuss anything. Copy their exact words, because those words become your screening notes. “Has owned a migration from start to finish” is something you can screen for. “Feels senior” gives you nothing to work with.
Pay attention to where they get stuck. The wildcard is usually where it happens, because a profile from logistics or banking doesn’t fit any shape they know. Write down the questions they ask out loud while they’re deciding. Those questions tend to be the real requirements, the ones nobody put in the job ad.
Print the packet if you can. People skim on screens and read paper slowly, which is a weak reason, but it’s a real one.
In week three, two of the six go back into the pile, mixed in with real new candidates. If the hiring manager judges them the same way, you’re aligned and can stop worrying. If they flip, you talk about what changed. Sometimes it’s their week. Sometimes the job has moved and nobody told you.
Of course, this is one example of an activity you can try, but it doesn’t work with every manager. It’s a starting point-you can experiment with it or devise your own approach.
Consistency is most of what makes a hiring judgment worth anything. When Paul Sackett and three colleagues re-ran decades of selection research in 2022, correcting a statistical adjustment that had inflated older estimates, structured interviews came out as the strongest predictor of job performance, with a mean validity of .42. Cognitive ability tests, which had topped those rankings for decades, landed at .31. Structured, in that research, means the same questions, scored against the same rules, by people who agreed beforehand what a good answer sounds like. A SIOP summary of the paper also noted a fair amount of spread around that .42, so a structured process run carelessly can still disappoint.
I’d only use the packet for technical and product roles. I don’t know what the sales version looks like, and I suspect it’s different.
Once the search closes, keep the packet. The six profiles and the hiring manager’s sentences are the most accurate record you’ll have of what “good” meant for that role, which is more than the job description ever captured. When a similar req opens six months later, start from that packet. Ask the hiring manager whether the old sentences still hold, and you’ll usually find out in ten minutes what a fresh intake meeting would take an hour to surface.
The packet is also useful when someone new joins the interview panel. Have them rank the same six before they meet a single candidate, then compare their sentences with the hiring manager’s. Gaps show up quickly, and it’s much easier to talk them through on anonymous profiles than after a live debrief.
There’s upkeep. Profiles age, and a packet built two years ago can describe a market that no longer exists. In the EU there’s also data retention to think about, so strip anything identifying and check your company’s policy before keeping any material past the end of a search.
If data retention is a concern, a safer option is to have AI generate six synthetic profiles modeled on the spread above (two strong, two borderline, one clear no, one wildcard).
The part above ended with the week-three check and the admission that it can make a hiring manager feel audited. That feeling mostly comes from how the conversation is opened, so below is the exact script for when someone flips on a packet profile, how to tell a real flip from a change in priorities, the four answers you’ll get and what to do with each, and the one-page sheet. It’s a separate thing from the packet itself.
The Week-Three Conversation, Word for Word
The week-three check is easy to set up and easy to ruin. What decides whether it helps is the conversation you have when a hiring manager flips on one of the packet profiles, so here’s the script, the lines to avoid, and what to do with each kind of answer.





