SEP 8, 2026
GOVERNANCE

An AI agent can "succeed" at its task and still produce the wrong outcome — conventional monitoring doesn't catch it

Published September 7, 2026 by Callum Turner on TheNextWeb ("AI agent reliability requires a new model of observability"), the article carries the argument of Robert Hommes, founder of the observability startup Moyai: an agent can make a valid request, receive a valid response, and still make the wrong business decision — the technical infrastructure records a success while the business experiences a failure, a blind spot conventional observability tools, built for deterministic systems and known error codes, simply were not designed to catch.

Hommes' starting point is concrete, not theoretical: "The most dangerous example of an AI agent is one that successfully completes a task while actually producing the wrong outcome." He gives the example of a procurement agent instructed to buy a specific type of coffee bean — the request goes out, the response comes back, the task closes as successful, but "from the business perspective, however, the agent is steadily producing the wrong outcome." The article cites the same logic in an airline scenario: an agent tells a stranded passenger their flight has been rebooked, the underlying reservation never actually goes through, and the passenger only finds out upon arriving at the airport. "We do not have an error code that says, 'I reached the endpoint, I queried it with the wrong parameter, and I received something different from what I needed.' Technically, nothing is failing, but it is not working," Hommes summarizes — a 200 response can mask an invalid decision just as easily as a correct one.

The most useful part of this article, for us, isn't the description of the problem — it's the honest limit Hommes points out in his own proposed fix. He credits human-in-the-loop architectures, which require an employee to approve a consequential action, with real value — exactly the principle AppManager has been built on since its first module. But he adds a nuance that can't be brushed aside: "Once you have material impact, you will see it, but you are already too late [...] It had to get worse before you noticed it." His proposal — detect what's different from normal behavior first, then check whether that deviation is actually a problem, rather than stacking a new rule for every failure already observed (what he calls "whack-a-mole") — tackles a distinct and harder problem than simply getting a human to approve an action already flagged as risky.

For AppH

  • The article puts a precise name on a principle AppH already applies without labeling it this way: a clean technical run doesn't guarantee a correct business outcome — which is why an AppManager automation is never closed out by itself; a human confirms the actual result, not just the absence of an exception.
  • AppManager's mandatory approval click before any consequential action is exactly the kind of human-in-the-loop architecture Hommes credits with real, useful protection — validated here by an industry expert with no connection to AppH.

Against / the honest limit

  • Hommes' own critique of human-in-the-loop applies partly to AppH too: an approval screen is only as trustworthy as the summary an automation shows before the click — AppH does not today have a dedicated behavioral-anomaly-detection layer like the one Moyai proposes, only explicit human oversight on the actions we ourselves have defined as consequential.
  • Robert Hommes is the founder of Moyai, a startup that sells exactly the observability product category he describes as missing — a credible, well-sourced argument, but also one made by a founder explaining why the market needs what his company sells.

What stopped us in this piece wasn't the general warning about AI agents — we read plenty of those — it was the line about the 200 response. A system that correctly answers a badly framed request signals nothing abnormal to whoever is watching it the conventional way. That's exactly why, at AppH, we never let an automation close itself out without a human confirming the actual result. That said, Hommes is right that human approval alone has its own limit — it protects the action we thought to route for approval, not the one we didn't identify as risky in time. We're not claiming to have this solved: AppH does not today have behavioral-anomaly detection like the kind Hommes describes, only explicit human oversight over what we ourselves have defined as a consequential action. Saying that plainly seems more useful to us than letting anyone believe the approval click, on its own, settles this for good.

Reviewed by an AppH human
← Previous article (older)Next article (newer) →

← Back to news

Want us to walk you through how this applies to a real case?

Talk to AppH

Get new posts by email

One email when we publish new analysis — never spam, unsubscribe in one click.