Flowmingo Logo
Blog/AI In HR

Interview bias: how big the main ones are and what helps

Flowmingo Editorial TeamFlowmingo Editorial Team5 mins readOct 06, 2026
Interview bias: an orange ring-bound notebook with a cream see-saw tipped to the right on its cover, 2 cream chairs facing each other on it with the right one sunk lower, and 3 cream divider tabs sticking out to the right named First impression…

You finish the third interview of the day and already know who you want. That is interview bias at work, and 87% of US hiring decision-makers say the first conversation tells them who will succeed (Express-Harris poll). Yet when all you know is that 1 candidate interviewed better, the odds that person is the better hire are only 56% to 61% (NPR).

None of this makes you unfair: your mind takes shortcuts to cope with 5 interviews a day. Most guides to hiring bias name the shortcuts but few size them, and the popular fix, awareness training, has little evidence of changing behavior.

The size tells you where to spend Monday. A strong first few minutes tracks the final score about twice as closely as the job offer.

This guide to interview bias sizes the effects and gives you 3 moves with research behind them, mostly from lab and review studies. Write the criteria before you meet anyone, and score each answer on its own against a written rubric. Then let a fixed rule add up the scores before anyone debates them.

Key takeaways

  • How much do the first few minutes decide? They track your final rating about twice as closely as the job offer, so score each answer separately and decide last.
  • Which one change should I make first? Write the criteria before you meet anyone. In 1 experiment with male volunteers, doing so removed a gender gap of about 1.5 points in scores.
  • Do structured interviews really reduce interview bias? In the scores, mostly. One study of nearly 20,000 applicants found no gender or race gap, but a review still found a race gap, under a third of the one for ability tests.
  • Do blind hiring or bias training fix the problem? Neither on its own. A review of 492 studies of methods to change implicit bias found trivial changes in behavior, and in 1 French trial of volunteer firms anonymous CVs widened the minority interview gap from 2.4 to 13 percentage points.
  • What should I do when my gut and the scores disagree? Let the total win unless you can name a job-related criterion it missed, and write that down. A fixed scoring rule predicted job performance more than 50% better than experts weighing everything in their heads.

1. How much does interview bias change who I hire, and which biases should I fix first?

Interview bias is real and measurable but not 1 number, so you cannot rank the biases by size. Fix them with 3 moves: lock the criteria, score each answer on its own, and let a fixed rule add up the scores.

Bias pushes scores in 1 direction, and noise is random disagreement. The table sizes parts of both, so do not add them up.

Bias How it shows up Size found What should help (our judgment; only criteria-first tested)
First impression An early view colors the rest Tracks the final rating at 0.42 and the job offer at 0.22 (Barrick et al.) Score each answer before the next
Daily quota (narrow bracketing) After several high scores you hold back, as if each day had a quota In MBA admissions interviews, the next score fell about 0.075 on a 1 to 5 scale (Simonsohn and Gino) Score against written anchors, not the day's tally (our inference; no fix was tested)
Similarity You warm to shared hobbies or style Shared culture often outweighed productivity in 120 interviews with elite-firm employers (Rivera, foundational 2012) Score a written working-style criterion
Moving goalposts You decide what matters after seeing who applied Male raters gave the male applicant 6.06 and the female applicant 4.53; no significant gap once criteria came first (Uhlmann and Cohen) Lock criteria before the first CV
Overconfidence A chat makes you surer, not more accurate Better performer picked 62% with tests plus an unstructured interview, 69% with tests alone (132 hiring decision-makers, airline ticket-agent applicants, Kausel et al.) Treat the interview as 1 weighted input
Halo effect 1 trait colors the rest No reliable size in the sources checked Rate each criterion on its own evidence

In a foundational 2010 study of internship interviews, impressions from the opening minutes tracked the interviewer's final rating at 0.42. They tracked the actual offer at only 0.22 (Barrick and colleagues). This means your first impression tracks your own rating twice as closely as the offer, and part of it is a real read (section 8).

The risk is letting those first minutes set the score before the evidence arrives.

In a foundational 2013 study of 9,000+ MBA admissions interviews, high scores earlier in the day pulled the next score down (Simonsohn and Gino). The authors call this narrow bracketing and found contrast an unlikely explanation. The next applicant lost about 0.075 points after a 0.75-point rise in earlier scores, counted from the third interview of the day.

That is small for 1 person but can swing a close call. Which to fix first is our judgment, not a study's result.

Interview bias across a hiring cycle: screening, first minutes, scoring order, interviewer agreement and the decision, each with the size found in research

Where interview bias enters: screening, first minutes, scoring order, interviewer agreement and the final decision, each with the size found in a foundational study.

2. How can I check my own interview scores and hires for bias without a data team?

Run 4 spreadsheet checks: interviewer agreement, how scores track hires, pass rates by group and what your notes say. A small team cannot prove interview bias with so little data, but it can spot big gaps.

  1. Do your interviewers agree? List each interviewer's score per criterion for the same finalists, then count the score pairs 2 or more points apart. Say 3 of 6 pairs do: by our rule of thumb, the anchors need rewriting. Across 125 reliability estimates, interviewers who met the same candidate separately agreed at 0.44 on a 0 to 1 scale (Huffcutt and colleagues, foundational 2013). Panel members, who share 1 interview, agreed at 0.74.
  2. Do scores track how hires performed? Line up each hire's interview score with a manager rating after about 6 months (our rule of thumb). You only see the people you hired, so the link looks weaker than it is.
  3. Do pass rates differ by group? At each cut, divide each group's pass rate by the highest group's rate. If 30 of 50 men and 12 of 40 women pass, the women's rate is 50% of the men's. Below 80% is the US rule of thumb for adverse impact by race, sex or ethnic group (29 CFR 1607.4). It is a cheap habit, not a legal line, and local rules may limit collecting group data.
  4. What do your notes say? Count vague phrases, for example "great energy" and "not a fit", which hide bias in friendly words and cannot be audited.

Interviewer bias and noise: separate interviewers agree at 0.44 on a 0 to 1 scale and panel interviewers at 0.74

Interviewer agreement: separate interviewers agree at 0.44 and a panel at 0.74 where 1 is perfect, across 125 reliability estimates and 32,428 people (foundational).

3. Do structured interviews reduce interview bias, and which one change should I make first?

Mostly, for the part you can measure: scores. In 1 study of nearly 20,000 applicants, highly structured interviews showed no score gap by gender or race, or by interviewer and applicant matching.

It covered 1 managerial job in 1 large organization and measured ratings, not hiring outcomes (McCarthy and colleagues, foundational 2010).

Berry and colleagues put the Black-White score gap in structured interviews at 0.24 standard deviations, under a third of the 0.79 for ability tests. Our own numbers below are AI ratings, and ratings against a fixed rubric still move together, so check yours.

Locking criteria shows why structure helps. In 1 experiment, male volunteers who set the criteria after seeing the applicant rated the male applicant 6.06 and the female applicant 4.53.

Those who committed to criteria first rated the male and female applicants 5.07 and 5.31, with no significant gap. They were volunteers, not HR professionals (Uhlmann and Cohen, foundational 2005).

Hiring bias experiment: male volunteers rated the male applicant 6.06 and the female applicant 4.53 when criteria came after, and 5.07 and 5.31 when criteria were set first

Criteria first: male volunteers rated the male and female applicants 6.06 and 4.53 when criteria came after, and 5.07 and 5.31 when set first (foundational).

Flowmingo data · AI interviews scored against recruiter-set criteria · 9 Jul to 4 Oct 2026

  • 0.51 typical correlation between 1 criterion's rating and another's inside the same hiring project, where 0 means unrelated and 1 means always moving together (median of 204 projects with 30 or more scored candidates, 27,599 AI interviews, 120 companies)

Ratings that move together can mean genuinely related skills or an overall impression carried across criteria, and the data cannot tell which. Scores are Flowmingo's AI scores, not job performance.

3.1 Which one change should I make first?

Write 4 to 6 criteria, and what a 1, a 3 and a 5 look like for each, before you meet anyone. Then score every answer against them.

This change comes first because it stops the criteria drifting toward the person in front of you, and every other step needs it.

The table shows 1 illustrative criterion for a customer support specialist, not customer data: handles an upset customer.

Score What the answer shows
5 The situation, what they said and did, the result, and what they would change
3 What they did, but no result, or "we" throughout
1 The customer being difficult and no action of their own

Our structured interview guide has example questions.

4. Does blind hiring work, or does hiding names just move the bias into the interview?

Sometimes. Hiding names helped in a famous orchestra study but backfired in a French trial, so test it on your own shortlist first.

People reach for blind hiring because field experiments keep finding callback gaps at the CV stage. Quillian and colleagues reviewed 28 field experiments (foundational 2017). In the 24 run since 1989, white applicants received on average 36% more callbacks than African American applicants.

The famous orchestra study is weaker than its reputation. A screen between auditioners and the jury raised by 50% the chance a woman advanced from some rounds (Goldin and Rouse, foundational 2000). The authors concede large standard errors, and an audition is a performance, not a conversation.

In a French trial of anonymous CVs, the interview gap for minority candidates widened from 2.4 to 13 percentage points (Behaghel et al., foundational 2015). Only 62% of invited firms with over 50 employees took part, and the researchers cite self-selection and lost context.

Blind hiring trial: the interview gap for minority candidates grew from 2.4 to 13 percentage points when CVs were anonymized

Anonymous CVs in France: the minority interview gap grew from 2.4 to 13 percentage points, and only 62% of invited firms took part (foundational).

Have 1 colleague shortlist about 20 recent CVs with names, photos and dates of birth hidden, then compare their list with yours (our suggestion). Differences can reflect taste as well as names, so use them as a prompt. Our candidate shortlisting guide suggests hiding names where you can.

Blind hiring does nothing for the interview itself, where you see the person, so the interview still needs structure to cut interview bias.

5. Does interviewer training reduce interview bias, or is it a waste of my managers' time?

Not on its own. The evidence finds awareness training may raise awareness briefly but has not been shown to change behavior, so practice scoring instead.

A UK government evidence summary found no evidence that unconscious bias training changes behavior or improves workplace equality (GOV.UK, foundational 2020). A review of 492 studies of methods to change implicit bias found mostly weak effects and trivial changes in behavior (Forscher et al., foundational 2019).

What does help is skill training on the rubric. In a foundational 2011 lab study of videotaped interviews, frame-of-reference training and anchored rating scales each improved rating accuracy (Melchers et al.).

In practice, frame-of-reference training can look like this 1-afternoon session, which is our own adaptation, not a tested protocol. Everyone scores the same 3 recorded answers alone against the rubric. Then the hiring lead shows reference scores written in advance, and the group rewrites any anchor where 2 people were 2 or more points apart.

If a veteran interviewer resists, show the Kausel result from section 1. Adding an unstructured interview rating to test scores made hiring decision-makers pick the better performer less often.

6. Interviewers write 'good vibes' or disagree on one candidate. How do we score against a rubric?

Score each answer on its own against written 1 to 5 anchors before anyone compares candidates, then add the scores by a fixed rule.

A feeling such as "good vibes" is where interview bias hides, because nobody can audit it or compare it across candidates. The table rewrites 4 vague notes as evidence (illustrative, not customer data).

Vague note Evidence note Criterion
"Great energy, really clicked" "Said they tested 2 reply macros and cut reply time by a third" Improves how work is done
"Strong culture fit" "Described disagreeing with a manager and what changed" Working style
"Not senior enough" "Has run a team of 3; the role needs 10" Scope of past work
"Reminded me of myself" Delete: not evidence None

Use the 1, 3 and 5 anchors from section 3. Then add the scores by a rule you fix before the first interview, for example equal weights. Add a must-have floor too, where any must-have below 3 means no hire (interview scorecard template).

6.1 Whose score do I trust when 2 interviewers disagree?

Neither is the verdict. A gap of 2 or more points on 1 criterion, in our rule of thumb, means the evidence or the anchor differs. First check that both used the same version of the criteria.

Next talk through the evidence, and anyone who missed something changes their score in writing. If you still differ, a third person scores the recording alone, with consent, and you average the independent scores.

In our data, criteria are fixed for everyone scored after the last edit, but earlier candidates are not rescored, as the card shows.

Flowmingo data · about 38,700 candidates · 797 companies · 9 Jul to 4 Oct 2026

  • 14.3% of AI interview reports lack at least 1 criterion that is on the project today, mostly because criteria were added after that candidate was scored (6,042 of 42,147 since 9 Jul 2026)

A recruiter adding a criterion is not evidence of bias, only of criteria drift. Scores are Flowmingo's AI scores, not job performance.

How Flowmingo helps

Flowmingo's AI Interviewer works from the questions and criteria you set up before applicants start. The overall score is a published formula over your weights (how scoring works). Structure supports consistency but does not make a score correct or free of bias (AI ethics).

7. How do I get interviewers to score alone before the debrief, so no one voice wins?

Have every interviewer submit scores in writing before anyone speaks, then read the totals out and start with the biggest gaps. Whoever speaks first can pull the room toward their view, and any interview bias in it spreads, so scores go first.

Kahneman calls independence "the real deep principle of what we call decision hygiene" (Issues in Science and Technology). He admits part of decision hygiene lacks research support.

In over 10 million panel interviews, interviewers who peeked at colleagues' scorecards were 3.6% more likely to give the exact same rating (Greenhouse, foundational 2023). It is vendor data from 1 platform.

Use this debrief script, which is our practice, not a tested protocol:

  1. Every interviewer sends the chair scores plus 1 line of evidence per criterion before the meeting, so nobody sees a colleague's scores first.
  2. The chair reads the totals, then opens on the criteria with the biggest gaps.
  3. Each interviewer says what the candidate said or did that earned the score, in reverse seniority order.
  4. A score changes only with a written reason.
  5. The formula gives the total, and any override needs a written, job-related reason and a second person's sign-off.

For example, a senior person who loved the candidate's energy overrules the scores, so step 5 applies to them too.

8. My gut and the scores disagree. Is 'culture fit' just bias, and who wins?

Treat your gut as a prompt to look for evidence, not as a vote. Let the total win unless you can name a job-related criterion the scores missed.

In section 1's Barrick study, competence impressions carried over to a separate interviewer, but liking did not. So when you like a candidate, ask which answer earned it and write that down.

A fixed formula beats holistic judgment even for experts. Combining candidate data by formula tracked job performance at 0.44, against 0.28 for combining it in the head (Kuncel et al., foundational 2013). That is more than 50% better.

Unconscious bias in decisions: a fixed formula tracked job performance at 0.44 against 0.28 for holistic judgment

Gut versus formula: holistic judgment tracked job performance at 0.28 and a fixed formula at 0.44, more than 50% better (foundational).

When the gut and the scores split, use this rule:

  1. Name the criterion your gut is reacting to.
  2. If it is already on the rubric, rescore it with the evidence.
  3. If it is not on the rubric, ask whether it is job-related. If so, add it for every candidate and rescore everyone, never for 1. Write the date and why, with a second person's sign-off, because criteria added mid-search are the moving-goalposts risk from section 3.
  4. If neither applies, the total wins. Write the reason either way.

8.1 What should I score instead of culture fit?

Score written, anchored working-style criteria that you ask of every candidate.

For example, replace "fit" with how the person takes feedback, how they disagree with a manager and how they handle unclear instructions.

9. Can interview bias get me sued, and what must I write down and keep?

Most interview biases, such as a first impression or a daily quota, are not unlawful in themselves. In the US, a hiring decision influenced by a protected trait is unlawful, even if it is not the only reason.

This is general information, not legal advice, and the interview is part of the hiring decision in the US, UK and Singapore. Written, job-related scores are what you can show if a rejected candidate asks why. For example, scores with evidence tied to the job answer that, while an unexplained "not a fit" leaves nothing to show.

Where The rule What it means for an interviewer
US Title VII bars refusing to hire because of race, color, religion, sex or national origin, even as 1 motivating factor; it covers employers with 15 or more employees for 20 or more calendar weeks Keep scores and notes with job reasons
UK (Great Britain) Equality Act 2010 s.39(1)(a) bars discrimination in choosing whom to offer work; s.60 limits health questions before an offer; Northern Ireland has its own laws Ask everyone the same job-related questions; candidates can request your notes (ICO draft guidance)
Singapore Workplace Fairness Act passed in 2025, not yet in force: expected end-2027, no gazetted date as at 4 Oct 2026 Use objective criteria consistently (voluntary Tripartite Standard)
New York City Local Law 144, enforced 5 Jul 2023 (foundational), covers automated tools Not human interviews: see our guide to fair AI interview rubrics

9.1 What must I write down, and for how long?

For each candidate, write the criteria, scores with 1 line of evidence each, who scored and when, the total and the decision's reason. Keep notes factual and job-related.

Keep them at least 1 year in the US, and longer once a charge is filed (29 CFR 1602.14). Singapore's voluntary Tripartite Standard also asks for at least 1 year.

Treating scorecards as hiring records is our reasonable reading, not explicit text. Some US states ask for longer, as the California example in our interview scorecard template shows.

10. What should I do before, during and after every interview to keep bias out?

Before the interview, write the criteria, anchors and questions, and during it score each answer before the next. After it, score alone, open the debrief on the gaps and let the total decide.

Each row of the checklist maps to a section above, from criteria and anchors in sections 3 and 6 to the debrief in section 7.

Before During After
Write 4 to 6 criteria and lock them before any CV Ask every candidate the same questions in the same order Submit scores in writing before the debrief
Write 1, 3 and 5 anchors for each criterion Score each answer before the next Read the totals, then open on the biggest gaps
Fix the questions, their order and the time Write what the candidate said or did Add scores by a fixed rule; log any override and its sign-off
Choose the scoring rule and who may override Hold the overall judgment until the end Keep scores and notes at least 1 year; check your numbers

For example, if you only have 20 minutes this week, write the criteria and anchors for the next role before you read any CV. Then score right after each answer and submit your scores before you talk to anyone.

If you are the whole panel, the independence has to come from time, because you cannot score alone beside a colleague. Hold the overall judgment for a day (our suggestion) and, with consent, ask 1 colleague to score a recording blind.

Structure shrinks bias and does not remove it. None of the effects above was measured on your team, so check your own numbers after each hiring round. No study we found measures the extra minutes structure costs per candidate, so we give no figure.

How Flowmingo helps

Flowmingo's AI Interviewer covers the criteria and scoring parts of this checklist. You set up the questions and weighted criteria before applicants start, and every report scores each criterion from 0 to 10 with a written justification. The decision stays with you. Same questions in the same order and independent scoring are practices for your own team. The Free plan includes unlimited AI interviews at Flowmingo. Each score comes with a written reason your team can question if it looks like interview bias.

11. Sources

Every study, law and quote in this guide links to a source below, and Flowmingo figures come from Flowmingo's own platform data.

Sign up and start in about 60 seconds. No card, no call

Sign up for free
Flowmingo Editorial Team

Flowmingo Editorial Team

We write practical guides for recruiters and HR teams who want to hire faster and more fairly. Each guide draws on hiring research, employment rules and Flowmingo's own data from real interviews, and lists its sources.

LinkedInXYouTubeFacebook

Oct 06, 2026