In 2014, Amazon set out to automate the hardest part of high-volume hiring: screening thousands of applicants down to a shortlist. Three years later, the company quietly scrapped the project. The reason behind that failure is the most important hiring-AI lesson of the decade, and most companies are still getting it wrong.
Amazon's tool wasn't broken in some exotic, technical way. It did exactly what it was built to do: it studied a decade of the company's own hiring decisions and learned to copy them. However, because most of Amazon's past engineering hires were men, the model taught itself that men were preferable. It downgraded resumes containing the word "women's" and penalized graduates of two all-female schools. Amazon could not guarantee the system would stop inventing new ways to discriminate, so it rightly shut the project down.
Here is the lesson to take away before anything else: an AI trained on your history will reproduce your history, bias included. And a shortlist you can't explain is a shortlist you can't defend. That was true in 2017 - in 2026, with AI now embedded across hiring, and regulators actively watching -it matters even more.
Learn what actually happened at Amazon, why the same mistake is happening at scale now, and five things your company should do differently to protect your brand, reputation and your candidates.
Amazon has long run on automation, from its recommendation engine to the packing machines in its warehouses. So as the company scaled from 11 employees in 1995 to hundreds of thousands two decades later, automating recruitment looked like the obvious next step. A team of engineers in Edinburgh was given a single goal: build a tool that could read resumes and identify top talent automatically, freeing recruiters from the grind of manual screening.
The intent was reasonable. The execution is what turned it into became a cautionary tale for Talent Acquisition teams everywhere.
Amazon borrowed a familiar idea from its own storefront. Just as shoppers rate products from one to five stars, the hiring tool would rate candidates from one to five stars, surfacing the best and discarding the rest.
To build it, the engineers created roughly 500 models trained to recognize about 50,000 terms that appeared in past applicants' resumes. They taught the system what "good" looked like by feeding it the resumes of the company's existing top performers, then matching new applicants against that pattern. According to Reuters, the algorithms even learned to favor verbs like "captured" and "executed" that appeared more often in male engineers' resumes.
The problem was quite literally baked into the data. Most of the resumes from past successful hires belonged to men, so the model concluded that being a man was itself a hiring signal. It began downgrading resumes that included the word "women's," as in "women's chess club captain," and filtered out candidates from all-women education systems.
Two failures compound here, and both are still common:
First, the training data carried the company's historical bias, and no one checked for it before the model went to work. Second, the tool produced a score with no explanation attached. A number from one to five tells a recruiter nothing about why a candidate ranked where they did, which means bias can hide inside it undetected. Amazon abandoned the project in 2017 and, according to The Guardian, was using a watered-down version as late as 2018.
It would be comforting to file Amazon's story under "early mistakes from the experimental era." But sadly the evidence points to that same failure now playing out across the market, only worryingly - faster and at far greater scale.
AI in hiring is no longer the exception: LinkedIn's Global Talent Intelligence reporting puts AI use somewhere in the recruitment process at roughly two-thirds of organizations, and analysts Fosway found 56% of HR teams already using it in recruiting outreach and 36% for sourcing recommendations. Screening at scale without AI is rapidly on track as the minority position.
The bias problem never went away: In 2024, University of Washington researchers testing the text-embedding models that power many resume screeners found they favored white-associated names in 85.1% of cases, and disadvantaged Black male candidates in 100% of the test cases (study). In October 2025, Stanford researchers found AI resume-screening tools systematically rated older and female candidates lower than younger male candidates, even when the resumes were otherwise identical. Amazon's mistake, made in 2017, is being reproduced by tools sold and deployed today.
Volume is making the temptation to cut corners worse: Job seekers now submit close to 11,000 applications per minute on LinkedIn, a 45% jump in a single year, much of it AI-assisted. CNBC reported in late 2025 that recruiters are "drinking through a fire hose," with 70% of hiring teams saying fewer than half the applications they receive even meet the role's criteria. When the pile is that big, an opaque tool that just spits out a ranking looks appealing, which is exactly the trap Amazon fell into.
And now the stakes are legal, not just reputational: Under the EU AI Act, AI used for recruitment and candidate evaluation is classified as high-risk. Organizations deploying these systems face obligations including bias testing, technical documentation, human oversight, transparency to candidates, and record-keeping, with penalties reaching EUR 15 million or 3% of global annual turnover. A hiring tool that can't show its work is no longer just risky; it's becoming non-compliant.
The through-line from Amazon to today is simple: bias enters through the data, hides inside an unexplainable score, and only surfaces once the damage is done. The way to avoid it hasn't changed either.
1. Assess skills, not your hiring history
Amazon's model failed because it was trained to find people who resembled past hires. If your past hires skew one way, so will your shortlist. The fix is to evaluate what a candidate can actually do against the requirements of the role, rather than pattern-matching them against who you hired before. Structured, skills-based interviewing gives every applicant the same competency-based questions and assesses the answers, not the pedigree. At NSS Group, this approach produced a 50% increase in hires from candidates who would never have passed traditional CV screening.
2. Reject the black box; demand an explanation for every score
A one-to-five rating with no reasoning behind it is where bias hides. Insist that your screening tool can show, for every candidate, exactly why they scored the way they did and against which criteria. Hubert's assessment uses deterministic AI models rather than a probabilistic black box: the same candidate giving the same answers gets the same score every time, and every score ties back to a specific response. That is what makes a shortlist explainable rather than just fast.
3. Test for bias continuously, not once
Amazon discovered its bias by accident, after the fact. Bias testing has to be built into the process and repeated, because models drift and data changes. This is no longer optional in the EU, the AI Act requires ongoing bias monitoring and documentation for hiring systems. Consistency is what makes real auditing possible. If a system produces a different result for the same input on different days, you can't audit it; if it's deterministic, you can.
4. Keep a human making the final decision
The goal is to augment recruiters, not replace their judgment. AI should do the heavy lifting of consistent, structured screening and then hand a clear, evidence-backed shortlist to a person who makes the call. Hubert is built this way on purpose: the final decision always stays with your team. Automating the screen is not the same as automating the hire, and conflating the two is how organizations sleepwalk into the Amazon problem.
5. Build for defensibility from the start, not as a retrofit
Amazon tried to bolt fairness checks onto a system that was never designed to be transparent, and it couldn't be salvaged. Legal defensibility and candidate fairness have to be design principles, not features added later. The encouraging part: this does not cost you speed. At OKQ8, structured AI interviews delivered an average time-to-shortlist of 57 minutes and a 9/10 candidate satisfaction score, with every candidate assessed against the same criteria. Speed and fairness are the same outcome when the system is built correctly, not competing priorities.
Amazon didn't fail because AI can't help with hiring. It failed because it pointed a powerful, opaque model at biased data and let it score people in the dark. The companies that get this right in 2026 will be the ones that treat explainability, consistency, and human oversight as requirements rather than nice-to-haves.
That is the entire premise behind Hubert: structured, skills-based AI interviews that give every candidate a fair, consistent assessment and give recruiters shortlists they can actually stand behind, in front of stakeholders and regulators alike. The benefits of AI in high-volume hiring, without repeating its most famous failure.
Want to see what defensible AI screening looks like in practice? Book a demo