Your applicant tracking system does what it was built to do. It collects, stores, tags, and moves people between stages. Ask it who is good and it has nothing to say, because judgment was never its job.
Someone still reads 400 applications. Usually late, usually fast, usually a person who started the week with three other requisitions.
Parsers and keyword filters read the CV and rank on what is written there. Fast, tidy, and biased toward whoever writes well about themselves.
The cost shows up in who never reaches the list. When NSS Group moved off document-led screening, hires from candidates who would have been filtered out at the CV stage rose by 50%. Those people were always qualified. Their CVs were just worse than their skills.
This is the newest option and the one worth slowing down for. A candidate answers questions, a large language model reads the answers, and a score comes out with a paragraph explaining it.
It reads beautifully, but the catch is repeatability. Run the same answers through that kind of model on Tuesday and again on Thursday and the ranking can move, because the model is producing a likely response rather than applying a fixed rule. Research on hiring use specifically has found this.
So when a hiring manager asks why candidate 41 beat candidate 12, you get an explanation that sounds right without being the actual arithmetic. Nobody notices until someone challenges a rejection.
The fourth option talks to all 400 people at once.
Softwares like Hubert run the same competency-based interview with every applicant, by chat or voice, in over 30 languages, on the phone most of them are holding. Nobody waits for a slot. Coop Östra went from application received to screening finished in under 1.5 hours.
Scoring is where it splits from option three. The conversation is genuinely conversational; the assessment underneath it is not improvised. Fixed rules, fixed weights, versioned and held steady for the length of a hiring cycle, so the last candidate through on Friday is measured exactly like the first one on Monday. Feed it identical answers twice and it returns an identical score twice.
That is a duller sentence than "AI-powered," and it is the reason the shortlist survives being questioned. The reasoning a recruiter reads is the reasoning that produced the number, not a justification written afterward.
The assumption is that candidates dislike being interviewed by AI. The numbers say otherwise: 96% complete the interview, and satisfaction averages 9 out of 10.
If your applications arrive in tens, read them yourself. You will do a better job than any tool on this page.
At 400, or 4,000, the choice is not really between vendors. It is between screening the document, screening on a model's mood, or screening the conversation. Everything else in a comparison table sits downstream of that, implementation included: Aleris were interviewing candidates five working days after signing, inside the Teamtailor setup they already had. The slow part is no longer the software, it is deciding what you want measured before ten names go to a hiring manager on Friday.