An AI interview is a job interview conducted by software rather than a person, through either chat or voice. Candidates can complete it on any device, on their own schedule and recruiters receive scored results, usually inside their existing ATS.
Not every AI interview is structured, and this is the first quality marker to check. Some tools improvise the conversation, so questions drift from one candidate to the next. In a structured AI interview, every candidate applying to the role answers the same competency-based questions, assessed against the same criteria. Decades of selection research show structured interviews are among the strongest predictors of job performance; unstructured ones are not, whether a human or an AI is asking.
When a candidate applies, the agent invites them to an interview they can complete on any device, at any time. During the interview, the agent asks follow-up questions and gives candidates space to explain their experience in their own words. Afterwards, responses are scored and the recruiter sees a ranked list of candidates. How that scoring happens varies significantly between tools, and it is the single most important thing to understand before buying one.
There are two fundamentally different approaches, and vendors do not always volunteer which one they use.
Some tools use large language models (LLMs) to score responses. LLMs are probabilistic, which means the same answer can receive different scores on different runs. Research has documented this inconsistency in hiring contexts. An LLM can also generate a convincing explanation for a score after the fact, but that explanation is a plausible narrative, not necessarily the logic that produced the score.
Other tools use deterministic models: the same answer always produces the same score, and the explanation shown to the recruiter is the actual scoring logic. This distinction matters most if you ever need to defend a decision, to a rejected candidate, an internal auditor, or a regulator.
A useful question for any vendor: if the same candidate gave the same answers twice, would they get the same score both times? If the answer is anything other than an unqualified yes, ask how they would defend that in an audit.
In the EU, AI used in recruitment is classified as high-risk under the EU AI Act. That brings concrete obligations: the system must be transparent enough for humans to interpret its outputs, it must maintain an appropriate level of accuracy, and there must be effective human oversight. In the US, local rules such as NYC Local Law 144 add bias-audit requirements.
For buyers, the practical implications are: automated rejections without human review are a liability, explanations must be genuine rather than generated, and you need an audit trail covering scores, overrides, and recruiter actions. Ask where candidate data is stored and whether it is used to train third-party models; both questions have regulatory and candidate-trust consequences.
The evidence suggests yes, when the experience is well designed. The two strongest signals to look at are completion rates and satisfaction scores: an interview candidates abandon, or resent, is telling you something. As a reference point, published figures from deployments at scale show completion rates in the 88–96% range and average candidate satisfaction around 9/10. Hubert publishes its candidate reviews openly if you want to see the verbatim feedback behind those numbers.
Two design factors drive this type of acceptance. Flexibility: candidates can respond at their own pace, outside office hours, on a phone. And relevance: candidates consistently rate interviews higher when the questions clearly relate to the actual job. Get either wrong, with generic questions, no follow-ups, and no feedback, and the effect reverses... a bad AI interview doesn't just lose one candidate, it damages the employer brand with everyone they tell.
Five questions cover most of the risk:
How is scoring done? Deterministic or probabilistic, and can they prove it.
Is the science real? Ask what the assessment methodology is grounded in and how models were validated. Structured interviewing frameworks such as STAR with behaviorally anchored rating scales are the established standard.
Who makes the decision? The final call should stay with your team, with no automated rejections. This is both good practice and, under the EU AI Act, a compliance requirement.
Does it work in your ATS? Screening data that lives outside your workflow creates admin, not efficiency.
What happens when candidates use AI? Candidates increasingly draft answers with ChatGPT. Ask whether the tool detects AI-generated responses and, crucially, what happens next; flagging for human review is defensible, auto-rejecting is not.
A chatbot answers questions or collects basic information, such as availability or license checks. An AI interview agent conducts a full structured interview: competency-based questions, follow-ups, and scoring against explicit criteria. The output is a ranked, explainable assessment, not a form submission.
It should not. Under the EU AI Act's human oversight requirements, and as a matter of defensibility, the agent should screen and rank while a human makes every decision. If a tool auto-rejects, ask the vendor how that holds up when a rejected candidate challenges the decision.
Some try. The reasonable response is detection plus human judgment: tools such as HubertDetect™ flag likely AI-generated responses for recruiter review rather than rejecting anyone automatically. Treat a flag as information for a human, not a verdict.
Far less than most enterprise HR software, because the agent plugs into an existing ATS rather than replacing it. Documented deployments range from five working days (Aleris) to two to three weeks (OKQ8, with no formal training required). If a vendor quotes months, ask what the time is actually spent on.