When the machine does the work, what is the human for?

Human oversight in AI systems

Most of the anxiety about AI and particularly AI agents is focused on which roles in which sectors will disappear. In my mind, this is the wrong framing. I do not think the impact will be defined as much by sector or role as by the changes cutting across the jobs that remain. The more important question is what sort of work we are leaving for people as agents take over more of the routine execution.

Recent research from Glean’s Work AI Institute offers an early indication of what those roles could look like. It found that while AI is saving digital workers around 11 hours a week, employees are also spending 6.4 hours managing it, including checking outputs, debugging errors and cleaning up AI-generated work. The findings point to an uncomfortable possibility. We may automate much of the routine execution without removing the burden from employees, instead leaving them to supervise the systems, deal with failures and carry the responsibility when something goes wrong.

“Human-in-the-loop” is the phrase doing most of the reassuring here. It appears in every governance policy, every vendor deck, and every board paper that wants to sound responsible. And most of the time it means almost nothing. A human is in the loop in the sense that a name sits next to an approval button when the work stops until a person clicks on it. Whether that person understands the decision they are approving, has the time to interrogate it, or could realistically overturn it, is a separate matter, and usually an unexamined one.

I have started to think of this weak version of the idea as accountability without authority. We put a person at the end of an automated process, hand them responsibility for the outcome and give them none of the means to change it. They are expected to approve decisions without seeing how they were reached, while being measured on throughput in a way that discourages them from pausing to challenge the system. The model is right often enough that questioning it feels faintly ridiculous, until the day it is confidently wrong and the person overseeing it discovers their role was to absorb the blame, not to prevent the error.

Anyone who has worked with public sector organisations will recognise the shape of this. A caseworker “reviews” an eligibility decision produced by a system whose logic they were never shown, on a screen that offers approve or refer, under a service-level target that assumes approval. The result is a form of liability transferred to the person at the end of the process. An IBM training manual from 1979 put the problem plainly: “A computer can never be held accountable, therefore a computer must never make a management decision.” Close to fifty years later, many organisations have built precisely the arrangement it warned against, placing a human at the end of the process to carry the accountability the computer cannot.

So the more useful design question is where humans should remain involved and what authority they need to exercise meaningful oversight.

The “where” is easier than people imagine or pretend. Human judgment is wasted on the routine cases that make up most of the volume, and it is indispensable on the edge cases, the novel situations, the decisions where the cost of being wrong is measured in someone’s liberty, health or livelihood rather than a refund. Good design does not sprinkle a human evenly across every transaction. It concentrates human attention where the machine is least trustworthy and the stakes are highest and lets the machine run where it is genuinely more reliable than we are. Deciding which is itself expert work, and it is work most organisations have not done.

A better operating model can give people more meaningful work and clearer authority over the systems they oversee.

Authority is the harder half because it is cultural rather than technical. For an override to be real, four things have to be true at once. The person needs enough visibility into the system to form an independent view. They also need sufficient time to question the decision properly and the standing to say no without putting their career at risk. And those around them have to treat a challenged decision as the system working, not as an insubordinate employee gumming up the automation. Take any of those away and you are back to “AI theatre”. Most human-in-the-loop controls fail on the last two, which is precisely why they are the ones nobody audits.

This changes what we should be training people for, and it is the part I find genuinely encouraging. The centre of gravity of a good job is moving away from carrying out a workflow, because the workflow is exactly what is being automated, and towards governing, improving and overseeing the systems that carry it out. What emerges is a more demanding and more valuable job than the one it replaces.

It also already has a shape. In consultancies and AI labs, the emerging name for it is the forward-deployed engineer. This is someone embedded in the messy reality of a live environment, who takes a system from a messy set of requirements through a promising pilot to something that actually runs, owns it once it does, watches it in production and keeps deciding which problems are worth the engineering effort and which are not. Palantir coined the idea years ago and the rest of the industry has quietly converged on it. OpenAI, Anthropic, Databricks and the larger consultancies all now field some version of the role because it turns out the hard part of enterprise AI was never the model’s capability. The real challenge lies in everything around it, including the integration, the edge cases, the judgement calls and the person willing to stand behind the output.

Forward-deployed engineers offer a useful preview of how work may change as machines take on more of the execution. A similar principle applies whether organisations are hiring specialists or reskilling the people they already have. Success in this kind of work depends on being able to reason about a system, question its outputs, recognise when it is operating beyond the conditions it was built for and decide which issues need to be escalated. The people currently carrying out these workflows are often well placed to move into governing and improving them. Developing that capability internally is likely to be more sustainable than hollowing out their roles and relying entirely on scarce external talent.

A technical precondition rarely makes it into the governance conversation, and it is worth addressing directly. Meaningful oversight depends on visibility into how a system operates. Genuine authority over an automated system also requires it to be inspectable, adaptable and capable of running on an organisation’s own terms. However, much of enterprise AI is moving in the opposite direction, towards opaque models accessed through APIs, whose behaviour can change without notice and whose reasoning remains off-limits by design. In those circumstances, a human may sit alongside the system without having any real control over it.

This is why the open-model argument matters beyond the ideology it often attracts. Recently, a broad coalition of companies, including some of the largest names in the industry, made the public case for open-weight models. Beyond the language around American sovereignty, the central argument is one of control. Organisations can inspect and adapt a model, run it on their own infrastructure, avoid dependence on a single provider and retain the knowledge they develop through using it.

The safety argument should also matter to anyone responsible for governance. A system that can be examined, tested and corrected by many people may be easier to defend than one whose safety depends on keeping its inner workings hidden. Complete openness will not suit every use case, and most organisations are likely to rely on a combination of models. The challenge lies in choosing the right model for the right task at an appropriate cost. For humans to exercise genuine authority over these systems, transparency cannot be treated as a minor procurement consideration. It determines whether meaningful oversight is possible.

This brings us back to the work that organisations now need to do. Moving AI from the pilot phase into live operations depends on getting the operating model right. Leaders must decide which decisions can be handled by machines, which require human judgement and what authority people need when they challenge an automated outcome. A better operating model can give people more meaningful work and clearer authority over the systems they oversee. Without one, businesses risk automating the most engaging parts of a role while leaving employees accountable for decisions they have little power to influence. Many are already making these choices by default, without fully considering the jobs they are creating in the process.

Ash Gawthorp

Ash Gawthorp is Chief Technology Officer at Scale Factory (formerly Ten10), a UK AI enablement consultancy specialising in quality engineering, cloud, data, and agentic automation, with all delivery carried out by onshore UK teams. With more than 25 years of experience in software engineering and technology leadership, Ash works closely with organisations to modernise systems, improve operational resilience, and adopt emerging technologies, including AI, in practical and scalable ways.

Author

Scroll to Top

SUBSCRIBE

SUBSCRIBE