An AI agent told Zebra it had made $10 billion. That's the mistake its CIO isn't worried about.
Matt Ausman governs AI agents across a global operation. The errors that worry him are the ones too small to catch — and nobody has decided who carries them.
An automated reporting agent inside Zebra Technologies produced a number. Ten billion dollars of revenue, it reported, in a single business segment.
That figure, as Zebra’s Chief Information Officer Matt Ausman pointed out when he told me the story, is more than the entire company’s revenue. So nobody believed it. The agent had invented the number, presented it in the ordinary way, and been caught inside a morning.
Which is where most stories like this stop, on a tidy moral about checking the robot’s homework. Ausman went somewhere else.
“It’s the big hallucinations that I’m not worried about,” he said. “The big stuff we catch well. The small stuff scares me.”
That inverts most of what is currently said about artificial intelligence going wrong. The disasters we are being warned about are the ones an organisation is already built to catch. A machine that claims your warehouse business earned more than your entire company is not a hard problem; it is a self-reporting one. The hard problem is the machine that is wrong by five per cent, quietly, for eight months.
The mistake that tells on itself
Ausman runs technology for a company most people have never heard of and almost everyone depends on. Zebra makes the scanners, handheld computers, label printers and machine-vision cameras that run the world’s warehouses, shop floors and delivery rounds. If a parcel reached your door this week, a piece of Zebra equipment almost certainly tracked it. He came up through General Electric, across power turbines, healthcare and banking, so he thinks like someone who has watched a crane lift a hundred tons off a turbine rather than someone who has read a paper about it.
His example of the error that worries him is deliberately unglamorous. A pallet gets put away in the wrong place. Not dumped in the middle of an aisle, where anyone walking past would see it. One rack over. One row across. Close enough that the system says the job is done and nobody has reason to look.
“Those are the ones that really worry me the most,” he said. “It’s the close that a human may not even catch it either.”
Nobody pays for that error on the day. It is paid a fortnight later, when a picker goes to the rack, finds nothing, glances left and right, gives up, and the order ships late or not at all. Multiply by a few thousand pallets and you have a number that surfaces in a quarterly review as a mystery, with no incident and no alert behind it.
Both of us reached for the same film, independently, about ninety seconds apart: Office Space, and the scheme to skim fractions of a penny off every transaction on the theory that nobody audits a rounding error. The plot works because the fraud is built to sit below the threshold at which anyone looks. An AI agent does not need to be malicious to produce the same effect. It only needs to be slightly wrong, consistently, inside your tolerance.
The number that ends the checking
There is a fix Ausman has been testing, and it is better than it first sounds.
Show people the confidence score. When an agent produces an answer, have it say how sure it is. “What if we start saying, or we have AI say, I think this is 85 per cent correct?” In his experience this improves behaviour: at fifty per cent, the human looks hard. At ninety-nine, they wave it through.
And there is the trap, which he names himself.
“That’s what scares me though — when it does get to be 99 per cent, you’re no longer checking it. And when it does fail, how do we make sure that humans catch it?”
Psychologists have a name for this, and so, it turns out, does the British regulator. Automation bias is the documented human tendency to accept a machine’s output because it came from a machine. The Information Commissioner’s Office — the UK’s data protection regulator, and the body that would investigate if an automated decision harmed you — puts it almost that plainly in its own artificial intelligence audit framework: “non-meaningful human review is caused by automation bias or a lack of interpretability.”1 The warning is not about people who ignore the machine. It is about people who agree with it too readily.
Here I want to push past where either of us took it on the recording, and mark this as mine rather than his.
When the confidence score reaches ninety-nine per cent, two things become true at once. You have stopped checking. And what you were checking was probably the wrong thing in the first place.
A confidence score is the machine marking its own homework. It reports that the step was performed correctly, by the machine’s own reckoning. It says nothing about whether the business got what it wanted. An operation can run at ninety-nine per cent on every process measure it owns and still lose the pallet, miss the delivery, refund the customer and carry the cost, because the measure was watching the machine rather than the result.
This is not theoretical. MIT‘s NANDA initiative studied more than three hundred enterprise deployments of generative AI and found roughly ninety-five per cent producing no measurable acceleration in revenue.2 Read that alongside the confidence-score problem and a pattern falls out. Those projects were not failing their own tests. They were passing their own tests and failing to move any number anybody outside the project cared about.
So the instruction is uncomfortable and cheap. Stop auditing the machine’s confidence; start measuring the outcome. Did the customer get the thing. On the day. Undamaged. At the price. Were there fewer lost pallets this quarter than last. Those are numbers no agent can score itself against, which is exactly what makes them worth having.
The one-page version of this argument
Every agent you run, five questions, ten minutes. Which loop it sits in. What happens when it is wrong. What you measure today. What you should measure instead. And who owns the one per cent.
The detection test sits at the foot of it: if this agent were wrong by five per cent, every day, for eight months, which number would move first — and who is looking at that number?
Download the agent accountability worksheet
One page. No email needed.
In the loop, on the loop, and the alarm clock you never watch
“Human in the loop” is the phrase every organisation reaches for when asked whether its AI is safe, and it has been worn smooth by overuse. Ausman does something useful with it. He takes it apart into three.
Human in the loop is you, doing a step yourself, alongside the machine. You are the check, every time, by design.
Human on the loop is your manager. Nobody inspects each item; somebody watches a dashboard, reads the exceptions, audits a sample. Every factory floor has had a version of this for decades.
Human out of the loop is your alarm clock. You set it, you sleep, you do not lie awake monitoring it. Fully autonomous, and nobody thinks that reckless, because the task is narrow, the failure is obvious and the cost of getting it wrong is a late breakfast.
The sorting question is not how clever the system is. It is what happens when it is wrong. Dispensing the wrong medicine can kill someone, so you want four independent sensors and a human. Putting a parcel on the wrong doorstep annoys someone, so you do not. “Do you need to be 99.9999 per cent?” he asked — the standard he wants from the engine of the aircraft carrying him over the Atlantic — “or are you okay with an 80 per cent accuracy rate?”
What the taxonomy exposes is the organisation that believes it is in the loop and has drifted, unannounced, out of it. Nobody signs that decision off. The confidence score just kept getting better.
Your data will never be clean, and that is fine
The standard advice for any organisation approaching AI is to fix its data first. Clean the records, sort the governance, assign the ownership, then start. It sounds reasonable, and I have watched organisations disappear into it for years and come out with nothing.
Ausman’s answer is close to heresy, and it is the most commercially useful thing in the episode.
In the back office, he agrees, you can chase clean data and largely get it. At the front line you cannot, because human beings are involved. Someone picks a jar off a supermarket shelf, changes their mind three aisles later, and puts it down wherever they happen to be standing. A child grabs something. A driver puts a pallet one row over. “You will have bad data as you get closer to the front line,” he said, and no governance programme fixes it, because the environment keeps moving.
So the strategy changes shape. You stop trying to make one record perfect and start checking the same fact several ways. He describes asking a colleague which technology would win, machine vision or radio tags. The answer was neither: it is both. A barcode says one thing, a radio tag another, a camera reads the label, the weight says a fourth. When three agree and one does not, you have found something worth a human’s attention, without ever having achieved a clean database.
That is a different way to spend the budget. Fewer consultants on a data quality programme. More sensors, and an agent whose job is to notice when they disagree.
Reading this far?
Subscribe to The Control Layer for one piece a week in this register — AI, cybersecurity, sovereignty, and the geopolitics of the technology stack. Free.
Who owns the one per cent?
Ausman put the question better than I could, and put it as a question because he has no answer either.
“If something is 99 per cent good, who owns that one per cent risk? What if it’s 95 per cent?”
His way in is the self-driving car. Today, if you crash, it is your fault and your insurance pays. When the car drives itself, who pays? The manufacturer? The company that wrote the software? You, for having bought it?
Then he brings it back to the floor and it stops being an abstraction. A worker follows a procedure an agent gave them. The procedure was wrong. That worker did not write the agent, test it, choose it, or have any way of reading its code.
“Are they going to be reprimanded because they followed a procedure that AI gave them from an assistant, and it was the wrong procedure? Whose fault is that?”
This is the argument I want to make here. Accountability does not evaporate when a human steps back from a decision. It goes looking for an owner. And where nobody has decided in advance where it should land, it lands by default on the least powerful person in the chain, the one on a shift holding the handheld, who did what the machine told them.
His larger fear is worse. A single agent inside your own business you can watch. What about your agent talking to your supplier’s agent talking to your carrier’s agent? He compared it to theft inside a company: two people you can control, five in collusion you cannot. “Three different agents getting together and coming up with something you never expected is much, much harder to control, spot and identify.”
He is a self-declared science fiction obsessive, Asimov first and Foundation above all. So when I asked whether we are governing today’s autonomous systems with a rulebook written for yesterday’s assistants, he went to the Three Laws and straight past them. Asimov’s robots were rarely villains. They followed the rules exactly and produced outcomes nobody wanted, because a rule interpreted at machine speed stops behaving like the rule you wrote. “I interpret the rule a little bit differently,” is how Ausman put it, in the voice of the machine. He raised Skynet, half-joking, then gave the serious version: what worries him is the speed and scale at which a misread rule propagates before anyone notices.
Which returns us to the small errors. A fast, confident, slightly wrong system is a harder governance problem than a spectacular one, and we have built our controls for the spectacular.
Europe moved the date. Britain went the other way
There is a legal clock running on this, and it moved three weeks ago in a direction most people missed.
The European Union’s AI Act was due, on 2 August 2026, to make human oversight a hard legal requirement for high-risk systems, including by name those used to monitor and evaluate the performance and behaviour of workers.3 Then on 24 July the EU published Regulation (EU) 2026/1744, the Digital Omnibus on AI, in force from 27 July, moving that deadline to 2 December 2027.4 The transparency rules — telling people they are talking to a machine, labelling synthetic images and video — arrived on schedule.5 The requirement to have a human meaningfully in charge slipped by sixteen months.
Hold that against something Ausman mentioned in passing. Zebra built an agent whose job was to audit its own people’s work on order entry, after years of resistance to automating that process. The agent found the mistakes the humans had made, trust followed, and with it the automation they had been arguing about for years. Excellent change management, by the side door.
It is also, on the face of the text, exactly what the AI Act calls high-risk: a system that monitors and evaluates the performance of workers. In Europe, the duty to keep a human meaningfully in charge of it now begins in December 2027.
The United Kingdom went the other way, and almost nobody covered it. Section 80 of the Data (Use and Access) Act 2025 replaced the old blanket restriction on automated decisions with new Articles 22A to 22D of the UK GDPR, in force since 5 February 2026.6 The new regime allows significant automated decisions more widely than before, tightening only around sensitive data such as health and biometrics, and demands safeguards in exchange: tell the person, let them object, let them demand a human, let them contest the outcome. Everything turns on one phrase — meaningful human involvement. If a human is meaningfully involved, the decision is not solely automated and the restrictions fall away.
Which is where the ninety-nine per cent problem stops being operational and becomes legal. Someone who has stopped checking because the score is high is still, on paper, in the loop. Whether they are meaningfully involved is a question the ICO’s own audit framework answers badly for anyone hoping to wave it through: a reviewer needs the “knowledge, experience, authority and independence to challenge decisions”, and non-meaningful review is what automation bias produces.1
Two jurisdictions, opposite directions, same unanswered question. A right to contest a decision is worth nothing against an error nobody detected.
Predictive judgement
Ausman’s own prediction, volunteered, is that within ten years we will each have a digital twin: an agent that sees what we see, reads what we read, answers our email and takes our meetings. He is not entirely comfortable about it. “When do we stop interacting with humans as much?”
I am not going to adopt it, for the same reason I would not adopt any ten-year call. Nobody can be held to it, which makes it an opinion rather than a judgement. Here is mine, narrower and dated so that you can.
Prediction: By 2 December 2027, the day the EU’s deferred human-oversight duty finally applies, at least one named organisation or regulator will have published an account of an AI agent incident in which the stated root cause is an error that ran below the detection threshold for weeks or months, rather than a single visible failure. Not a hallucination somebody spotted on the day. An accumulation nobody spotted at all.
Signals to watch: the phrase automation bias appearing in a UK or EU enforcement notice or reprimand; agent-specific language entering the operational risk sections of FTSE 100 and Fortune 500 annual reports; and continuous outcome monitoring, rather than model accuracy, appearing in enterprise procurement questionnaires for agentic systems.
What would falsify it: if by that date published incidents remain dominated by the visible kind — the invented figure, the leaked record, the obvious hallucination — then the detection problem is smaller than Ausman and I both think, and existing controls are catching more than either of us credits them for.
The publication that calls its predictions in writing.
Every Control Layer piece ends with a falsifiable prediction and a list of signals to watch. Subscribe to track them. One email a week. Free.
The bottom line
Gartner expects more than forty per cent of agentic AI projects to be cancelled by the end of 2027, blaming escalating costs, unclear business value and inadequate risk controls.7 Ausman is relaxed about that, and his reasoning is worth borrowing: if every project succeeded, you were not taking enough risk. From the same forecast comes the number that matters more. Gartner expects at least fifteen per cent of day-to-day work decisions to be made autonomously by 2028, up from none at all in 2024.8
Fifteen per cent of the decisions in your working life, made without you, inside two years.
For fifty years the deal was different. On 26 June 1974, in a supermarket in Troy, Ohio, a cashier scanned a ten-pack of Wrigley’s chewing gum and the barcode era began. Barcodes are now scanned more than ten billion times a day.9 For half a century that machine did one thing. It told you the number. A human decided what to do about it. The scanner never had an opinion and never surprised anyone.
This year the machine started deciding. The accountability for those decisions did not vanish when we stepped back from them. It went looking for an owner, and it has not been given one.
You can automate the decision. You cannot automate the answering for it.
Listen to the full conversation
The Control Layer — “When the Tool Starts Deciding”, with Matt Ausman, Chief Information Officer of Zebra Technologies. Watch on YouTube, or listen on Apple Podcasts and Spotify.
And if you take one thing from it, take the worksheet: the five questions to ask about every agent you run, on one page.
Amer Altaf is Founder and CEO of Arkava, a UK and European sovereign AI agentic automation business, and Managing Editor of The Control Layer.
The Control Layer publishes weekly. Subscribe free.
Decision-grade analysis on AI, cybersecurity, technology sovereignty, and the geopolitics of the technology stack — written for the board paper, not the timeline. By Amer Altaf, Founder & CEO of Arkava and Managing Editor of The Control Layer.
One email a week. No paywalls on the analytical pieces. Unsubscribe in one click.
References
Unattributed quotations and accounts are from the recorded conversation, The Control Layer, published 19 August 2026. Descriptions of Zebra’s own operations, agents and internal practice are the company’s own account and are attributed as such throughout. The output of Zebra’s reporting agent is quoted as Matt Ausman described it and is not a statement about Zebra’s financial results.
Information Commissioner’s Office, AI and data protection audit framework — human review. The ICO requires that “human reviewers have appropriate knowledge, experience, authority and independence to challenge decisions”, and identifies that “non-meaningful human review is caused by automation bias or a lack of interpretability”. https://ico.org.uk/for-organisations/advice-and-services/audits/data-protection-audit-framework/toolkits/artificial-intelligence/human-review/ 2
MIT NANDA initiative, The GenAI Divide: State of AI in Business 2025, 18 August 2025. Analysis of more than 300 public enterprise AI deployments, 150 leader interviews and a survey of 350 employees; approximately 95 per cent of pilots produced no measurable revenue acceleration.
Regulation (EU) 2024/1689 (the AI Act), Annex III, point 4(b): AI systems intended to be used “to monitor and evaluate the performance and behaviour of persons in such relationships” are classified as high-risk. Article 14 sets the human oversight requirement for high-risk systems. https://artificialintelligenceact.eu/annex/3/
Regulation (EU) 2026/1744, the Digital Omnibus on AI, published in the Official Journal on 24 July 2026 and in force from 27 July 2026. Chapter III high-risk obligations, including the Article 14 human oversight duty, deferred from 2 August 2026 to 2 December 2027 for standalone Annex III systems and to 2 August 2028 for Annex I systems embedded in regulated products.
Article 50 transparency obligations, the general-purpose AI provider obligations in force since August 2025, and the Article 5 prohibited-practices regime in force since February 2025 were unaffected by the deferral and remained on the original schedule.
Data (Use and Access) Act 2025, section 80, inserting Articles 22A–22D into the UK GDPR; in force 5 February 2026. Article 22A defines meaningful human involvement; Article 22B sets the tighter rule for decisions involving special category data; Article 22C sets the required safeguards, including the right to make representations, to obtain human intervention and to contest the decision. https://www.legislation.gov.uk/ukpga/2025/18/section/80/enacted
Gartner press release, 25 June 2025: “Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls.” Analyst: Anushree Verma, Senior Director Analyst. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
Gartner, same release: “at least 15% of day-to-day work decisions will be made autonomously through agentic AI by 2028, up from 0% in 2024.”
GS1 US press release, 26 June 2024: the first UPC barcode was scanned on 26 June 1974 on a 10-pack of Wrigley’s chewing gum at a Marsh Supermarket in Troy, Ohio; barcodes are now “scanned more than 10 billion times daily”. https://www.gs1us.org/industries-and-insights/media-center/press-releases/gs1-us-celebrates-50-year-barcode-scanniversary


