top of page

The Human Who Always Agrees


Why your AI oversight layer has a failure rate, and why nobody is measuring it


A large enterprise recently described its AI hiring system to the press. The system screens hundreds of thousands of applicants a year and recommends a shortlist. Human recruiters then review those recommendations before anyone is hired.

To demonstrate that humans remained in control, the organisation offered a statistic:


Recruiters hire around 85 per cent of the candidates the AI recommends.

That figure was presented as evidence of oversight. Read it again. It is evidence of the opposite.


When a reviewer agrees with a machine five, six, seven times out of eight, the reviewer is not exercising judgment. They are confirming an outcome that has already been decided. The decision rights have moved. The accountability has not. What remains is a person whose signature makes an automated decision defensible without making it any more scrutinised.


This is not a criticism of one organisation, and the point is not that the tool is bad. The tool may be excellent. The point is that the governance wrapped around it is asserted rather than measured, and that this is now the normal condition of AI deployment across every industry I work in.


Oversight is not a checkbox. It is a control.


Every AI governance framework in circulation requires meaningful human oversight. The EU AI Act requires it. Australian regulators are moving toward it. Every internal AI policy I have reviewed in the last two years contains some version of the phrase.

Almost none of them define what it means operationally.


So the requirement gets satisfied at the design stage, a human reviews the output before action is taken, and it is never tested again. The organisation reports that human oversight is in place. The auditors accept it. The risk register records it as mitigated.


But oversight is a control. Controls have failure rates. A fire suppression system that has never been tested is not a control, it is an assumption with a budget line. We would not accept an untested control anywhere else in a high-consequence process. We accept it here because the control is made of people, and people are assumed to be inherently reliable in a way that machines are not.


The evidence runs the other way. Automation bias is one of the most robust findings in human factors research. When a system produces a confident recommendation, reviewers systematically defer to it, and they defer more, not less, as the system becomes more accurate. Reliability breeds trust, trust breeds deference, and deference is indistinguishable from oversight on an org chart.


Add the operating conditions most reviewers actually face. Volume targets. No time budget for disagreement. No structured process for recording a challenge. No consequence for agreeing and a real cost, in time and in friction, for disagreeing. Under those conditions, agreement is not a judgment. It is the path of least resistance, and the system was designed to make it so.


High-reliability organisations have language for this. Deference to expertise means authority flows to the person closest to the problem, not the loudest signal in the room. Reluctance to simplify means resisting the tidy answer. Both fail the moment the reviewer's rational strategy is to accept whatever the model produced.


Here is what makes this tractable rather than merely uncomfortable. The failure is measurable, and the data already exists in systems organisations are running today.


Override rate. How often does the human reviewer reach a different conclusion from the machine? Not the aggregate accuracy of the system, the rate at which the control actually fires. An override rate near zero does not mean the model is perfect. This indicates that the review layer is inactive.


Time on review. How long does a reviewer spend on each decision, and how does that distribute? If the median review takes eleven seconds, the reviewer is reading a recommendation, not evaluating a case.


Override outcome quality. When reviewers do disagree, are they right? This is the test that separates a functioning control from noise. If overrides produce better outcomes than the model's original call, the human layer is adding value. If they produce worse ones, you have a training problem rather than an oversight problem, and you should know which.


Disagreement distribution. Does override behaviour cluster in a small number of reviewers? If three people out of forty account for most challenges, you do not have an oversight function. You have three conscientious individuals and thirty-seven approvers, and you are one resignation away from having none.


None of these require new instrumentation.


They require deciding that the oversight layer is a system component with performance characteristics, rather than a policy statement.


This is not a hiring problem

I use the recruitment example because it was public. The pattern is not confined to recruitment and in contact centre operations it is already pervasive.


Automated quality scoring is the clearest case. AI evaluates a sample of interactions, or increasingly all of them, and produces scores against a rubric. A QA analyst reviews and calibrates. Within a quarter, the calibration sessions are no longer about whether the model is right. They are about aligning the human scorers to the model, because the model is consistent and the humans are not, and consistency is easier to defend than judgment. The control has quietly inverted. The machine now audits the auditor.

Agent assist follows the same arc. Suggested responses are accepted verbatim under handle time pressure. Acceptance rate gets reported as an adoption success metric, which it is, and never as an oversight risk metric, which it also is. The same number means both things and only one gets to the steering committee.


Automated dispositioning and summarisation close the loop. Wrap-up codes and interaction summaries generated by a model, approved in bulk at end of shift, then fed back as the ground truth for the next round of workforce planning, root cause analysis and model training. Errors do not stay errors.


They become the historical record.

In each case the organisation can honestly state that a human reviews the output. In each case the review has degraded into approval, and the degradation is invisible because nobody instrumented it.


When these failures emerge, the natural reaction is to question the model: Is it accurate? Is it biased? Is it explainable?

Those are necessary questions and they are not the exposed ones. Vendors have been forced to answer them, and the serious ones now do, with documentation and fairness testing and audit trails.


The unexamined surface is the review layer and that layer is not the vendor's product.


It is yours.

It is built out of workflow design, reviewer capacity, escalation paths, incentive structures and the question of whether disagreeing with the machine is a career-neutral act on a Tuesday afternoon with a queue backing up.


An AI system performs to the level of the architecture it is deployed into. That principle is usually invoked to explain why good models produce poor operational outcomes. It applies with equal force to governance. A well-built model inside a review layer that never fires is not a governed system. It is an ungoverned system with a compliance narrative attached, and the narrative will hold right up until the moment somebody asks for the override rate.


So ask for it. Before the regulator does, before an adverse outcome makes it a legal question, before the number gets quoted in a press interview as proof of something it disproves.


If your human oversight layer agrees with the machine ninety-five per cent of the time,

you do not have oversight.


You have a witness.




Built on Rigor. Engineered for Scale.

Comments


bottom of page