top of page

Nobody Audits the Rejects

Why the AI doom loop in recruitment is a measurement failure, not a behaviour problem


The term doing the rounds is the AI doom loop. Candidates use AI to generate applications at volume. Employers use AI to filter them. Both sides optimise against an algorithm rather than against the job. Reported application volumes now sit in the hundreds per role, and most practitioners will tell you privately that the published averages understate it.

The diagnosis is accurate. The prescriptions are not.

What follows the diagnosis, almost universally, is advice aimed at behaviour. Candidates should apply with more care. Recruiters should look harder for authenticity. Everyone should build a network before they need one. None of this is wrong, exactly. It is simply not a mechanism. You cannot ask a market to voluntarily de-optimise, and you cannot fix a systemic failure by appealing to the good intentions of the participants.


There is a mechanism available. It has been sitting in plain sight for twenty years in a completely different industry.


What actually broke


A resume and a cover letter used to be a costly signal. Not because the artefact itself was valuable, but because producing a tailored one took an hour. That hour was the information. It told the employer that the applicant had read the role, understood it well enough to write about it, and cared enough to spend the time. The document was a proxy. The cost was the message. Generative AI drove the marginal cost of a credible-looking application to approximately zero.

The artefact survived. The signal inside it did not.

This is a well-understood failure mode. When the cost of producing a signal collapses, the signal stops carrying information. The predictable response from the receiving side is to filter harder, which produces a stronger optimisation target, which drives applicants to optimise more precisely against it. Each cycle degrades the channel further. Nobody in the loop is behaving irrationally. That is what makes it a loop.


Framed this way, two things become obvious. The first is that “be more authentic” is unenforceable, and “use less AI” is unilateral disarmament that penalises the person who complies. The second is that only one side of this market can change the equilibrium, because only one side holds the gate.


The measurement that does not exist


Recruitment is a well-instrumented function by the standards of most corporate operations. Time to fill. Cost per hire. Source of hire. Offer acceptance rate. Application volume, endlessly. Funnel conversion at every stage.


Now ask a different question. What is the false reject rate at the screening gate?


In most organisations the honest answer is that nobody knows, nobody has calculated it, and no system in the stack is capable of producing it. The screening decision is the highest-leverage decision in the entire process, it is now substantially automated, and it is the one decision in the funnel with no error measurement attached to it.


Consider what that means structurally. An automated control is operating at scale on a population, it is producing an irreversible binary outcome, and there is no feedback loop connecting its decisions to reality. Successful hires are attributed to the process that selected them. Rejected candidates disappear and are never scored against anything. The control cannot be wrong, because being wrong has not been defined as a measurable state.


I have spent twenty six years in contact centre operations, much of it at head of contact centre level, accountable for the whole operation and for the quality and platform architecture inside it. From that seat, the question is never whether an individual evaluation was correct. It is whether the scoring system as a whole can be trusted to make decisions on your behalf at volume, and what instrument tells you when it cannot.


The answer in that world is not exotic. When automated scoring is introduced against interactions, whether that is speech analytics or an auto quality model, you build a standing calibration process around it. A sample of the machine’s output is re-scored blind by human evaluators, and the rate at which the two disagree is measured and tracked over time. That disagreement rate is the override rate. It is the control on the control. When it drifts, you know the model has decoupled from the operation before the operation feels the consequences.


That design requirement is not a nice-to-have you get to defer. An automated scoring model without a published override rate is not a control. It is an assertion, and I would not sign off on running an operation on one.


Recruitment has deployed automated scoring at a scale that dwarfs contact centre QA, on decisions with far higher consequence for the individual, and it has not built the equivalent instrument.


This is a system failure


There is a useful taxonomy for root cause that separates System, Skill and Will. Most commentary on the doom loop implicitly blames Skill or Will. Candidates are being lazy. Recruiters are not looking hard enough. Everyone has lost sight of the human element.


The evidence does not support that reading. Both populations are responding rationally to the incentives the architecture creates. Volume is rewarded on the supply side because the cost of an application approached zero and the return on any individual application approached zero at the same time. Filtering is rewarded on the demand side because throughput is what the tooling optimises and what the function is measured on. Change the individuals and you get the same outcome.


That is the definition of a System failure. And it is the same failure I see in AI deployments across contact centres, which is where the position I keep returning to comes from: AI performs to the level of the architecture it is deployed into, the same as any other platform. Recruitment did not get a bad technology. It deployed a capable technology into a process that already had no attribution model, and the technology executed the existing incentives faster than the process could absorb.


The hidden market is the symptom, not the cure


The most common closing move in commentary on this topic is to point candidates toward the hidden job market. The best roles are never advertised. Hiring happens in DMs and Slack channels and internal referrals. Build the network now, before you need it.

As a description of current reality, this is correct. As a recommendation, it should worry everyone in the profession.

Controls that are not measured lose trust.

That is a reliable pattern in any operation. When people stop believing that a control produces sound decisions, they do not campaign to improve it, they build a shadow process beside it and route the work through that instead. Referral hiring and quiet networked placement is precisely that shadow process. Its growth is a diagnostic signal that the formal channel has failed measurement, and the signal is being read as a career tip.


Look at what the shadow channel gives up:


  • No audit trail.

  • No consistency of assessment.

  • No defensible position under a discrimination complaint.

  • No mechanism to detect that a hiring manager’s network is systematically narrower than the labour market.

  • No attribution, which means no way to determine whether referred hires actually outperform screened ones or whether the profession has simply retreated into a channel where nobody can check.


The formal gate has poor measurement. The informal one has none by design.


Advising candidates to work the hidden market is reasonable individual advice in a broken system. Presenting it as the answer normalises the retreat.


What instrumenting the gate looks like


Three mechanisms, in ascending order of cost and descending order of how easily they can be dismissed.


Audit the rejects. Take a role that has been filled. Pull a random sample of applications the screening stage rejected, strip identifying detail, and have a human panel score them blind against the actual role requirements. Compute the disagreement rate. That number is your false reject rate and your override rate in one figure. It costs a day of work per role, it requires no new technology, and most organisations running automated screening have never produced it once. Until it exists, any claim that the screening model is working is unfalsifiable.


Change what the gate asks for. Replace the volume-tolerant artefact with a small, role-specific, verifiable task drawn from the work itself. Expensive to fake across a hundred applications, cheap to assess, and it restores cost to the signal without requiring anyone to behave more nobly than their incentives allow. The point is not to defeat AI use. It is to make the signal informative again whether or not AI was used.


Cap the funnel rather than filtering it. Close applications at a fixed number, or run short timed windows. Volume is the disease and harder filtering is symptom management. This one is rarely attempted because throughput is what the tooling is sold on, which tells you something about whose problem the tooling was built to solve.


None of these are difficult. The first is close to free. The reason they are not standard practice is that the function has never been asked to demonstrate the accuracy of its screening decisions, only the speed of them.


The uncomfortable version


If a hiring team cannot state how often a human reversed the machine’s reject decision, it is running an unmeasured control on the most consequential judgement in the process, and it is asking candidates to be more deliberate than the system they are applying into.

The doom loop is not a candidate behaviour problem.

It is what an unattributed control looks like at scale, in a function that has not yet been asked to prove its own accuracy.


Start with the rejects. Everything else is commentary.


Built on Rigor. Engineered for Scale.

bottom of page