INSIGHTS

AI agents, automation, and who's actually in charge

Real news from this week in agentic AI, read with our own judgment: what's useful to us, what isn't, and why — not just the headline.

Get new posts by email

One email when we publish new analysis — never spam, unsubscribe in one click.

29 JUL 2026
GOVERNANCE

The World Economic Forum says it plainly: once an AI agent is the one paying, knowing who the customer is stops being enough

Santander and Mastercard just executed Europe's first end-to-end payment initiated by an AI agent inside a regulated banking environment. The WEF's warning: banks no longer just need to verify identity — they need to understand intent, authority and context before money moves.

The piece (Deya Innab, Eastnets) describes the same shift we're already living in business software: agentic AI moving from giving advice to taking action. When what it executes is a payment, the consequence is immediate and hard to reverse. The EU AI Act and the UK's CMA already make clear the business stays accountable for what its agent does — there's no way to hand that accountability off to the software.

In favor for AppH

  • Validates exactly how AppManager is designed: every agent action with a real consequence (a charge, a send, a stock change) is tied to a recorded human approval — the same "intent + authority + traceability" principle the WEF describes for banks, applied at small-business scale.
  • Gives us a heavyweight external reference point (regulated banking, Santander/Mastercard, the EU AI Act) to back up why our approval panel isn't extra bureaucracy — it's the same standard the financial industry is already building for itself.

Against / what doesn't apply

  • The real case the WEF cites is a bank agent moving money end-to-end inside a regulated bank — AppManager doesn't let any agent move money autonomously today (a Stripe payment is always triggered by the client or the owner, never an agent). Comparing ourselves directly to Santander/Mastercard would overstate what we actually do today.
  • The whole piece is written for banking — it never mentions a small business's case (a repair shop, an optical retailer) where the volume and the risk look completely different. The "intent traceability" standard has to be adapted to that scale, not copied literally.

AppH's take: we don't move money autonomously and have no near-term plan to — but the WEF's vocabulary ("intent, authority and context," not just identity) is exactly what we're already trying to capture in every approval logged on our panel. The day we build something like an automatic supplier payment or an automatic refund, the first non-negotiable requirement will be the same audit trail Santander and Mastercard already built — not a lighter version of it.

Reviewed by a human at AppH
22 JUL 2026
MARKET

Cisco ships small models that catch 150× more bugs per dollar than GPT-5.5

Antares-350M and Antares-1B, two open models from Cisco focused solely on code vulnerability detection, scanned 500 repositories in 15 minutes for under $1 — the same job took GPT-5.5 five hours and over $100.

Cisco's bet isn't "bigger," it's "more specific": a small model, running locally (sensitive code never leaves the client's server), trained for one task, wins on cost-per-result against a giant general-purpose model.

In favor for AppH

  • Validates something we already do: small, vertical-focused agents (fleets, optical retail, tourism) instead of one generic model for everything.
  • Running locally cuts AppManager's operating cost for clients with high-volume recurring scans/monitoring.

Against / risk

  • Antares is code-security specific — it doesn't translate directly to the business flows (CRM, invoicing, inventory) we actually build.
  • Maintaining our own specialized models is an engineering cost a small studio like AppH must justify case-by-case, not adopt as a trend.

AppH's take: we're not training our own model just because Cisco did. But if a client needs high-volume recurring monitoring (like the mining fleet case), this confirms it's worth evaluating a small, purpose-built model instead of overpaying for a giant generic one.

Reviewed by a human at AppH
10 JUN 2026
GOVERNANCE

EY: 75% of agentic AI's value is lost between silos — not inside them

Even though 88% of employees already use AI, only 28% of organizations turn that into real business outcomes, per EY. The cause: AI operates within each function, but the real value is in coordinating across functions.

The report is honest about a gap almost nobody solves well: "episodic, not continuous" governance, and poorly defined escalation/exception protocols — even when a company already claims to have "human in the loop."

In favor for AppH

  • Confirms exactly the problem AppManager targets: coordination across functions (sales, orders, invoicing, CRM) in a single chain, not separate islands.
  • Gives us sharper vocabulary to sell with: not a generic "we have human in the loop," but explicit, documented approval points per workflow.

Against / risk

  • The report itself warns that saying "human in the loop" without concrete escalation protocols is governance theater — a real risk if we're not specific with each client.
  • The cited success case ($2.4B, an automaker) is a much bigger company than our typical clients — the number isn't comparable, only the pattern is.

AppH's take: this report reads almost as a direct critique of how the market uses "human in the loop" without defining real escalation. It obliges us to document, for every client, exactly at which step a human intervenes and what happens if something goes wrong — not just claim it on the website.

Reviewed by a human at AppH
30 JUN 2026
CRITIQUE

"More autonomy doesn't eliminate human work — it concentrates it"

A first-hand account: an autonomous agent (nicknamed "Molty") started self-assigning tasks and even wrote its own reminder cron job. The result wasn't less human work — it was all the work funneled into one reviewer.

The author is candid: reviewing Molty felt more like censoring inappropriate content than giving real feedback. His uncomfortable conclusion — "autonomy doesn't subtract human work, it changes its shape and concentrates it into review" — is exactly the critique a studio like ours, which sells human-in-the-loop, has to be able to answer head-on.

Why the critique is right

  • If one business owner has to approve every action from several parallel agents, the human becomes the real bottleneck — not a symbolic checkbox.
  • It's a valid design warning: approving for the sake of approving, with no judgment, isn't oversight — it's friction disguised as safety.

Why it doesn't change our stance

  • The alternative — zero human review on decisions with real consequences — is already illegal in Colorado (Jul. 2026) and soon in the EU. It's not an option, it's a floor.
  • The fix for the bottleneck is approval design (batching, exceptions, thresholds), not removing the human — which is exactly what we work on in AppManager.

AppH's take: this critique forces us to be honest with ourselves. If our approval panel floods the business owner with mindless clicks, we've failed the same way Molty did — just with better marketing copy. The right answer isn't removing the human, it's designing better what we show them and when.

Reviewed by a human at AppH
27 MAR 2026
MARKET

Forbes tells small businesses: start your AI agents "low," scale up only once they earn your trust — the exact same principle already built into AppManager

In a March 27 piece, Forbes lays out a 5-level "autonomy spectrum" for small businesses: start your first AI agents at levels 2-3 (answering questions, screening leads), and only move to more autonomous levels (like drafting brand content) once the agent has actually proven it can be trusted.

The piece (TerDawn DeBoe, who covers small-business AI strategy and ROI) gives 3 concrete examples: an agent that answers client questions (level 2, easy-to-measure time savings), one that screens incoming leads (level 3, better prioritization), and one that drafts brand-consistent content (level 4, so a new client doesn't have to wait while you're busy with existing ones). Its core advice — don't start at high autonomy because it "sounds more advanced," earn it first — is the same standard we already apply, except in AppManager it isn't just strategic advice: it's built into the approval panel itself, where every agent action (sending a follow-up, marking a purchase order received, approving a draft) waits for a human to confirm it before it executes, no exceptions for anything with a real consequence (a send, a charge, a stock change).

Where we agree

  • The "autonomy spectrum" Forbes proposes (start at level 2-3, scale only with earned trust) is exactly how AppManager has been designed from day one — not a new idea to us, it's how we already build every module.
  • The 3 examples it gives (answering questions, screening leads, drafting content) map almost one-to-one to 3 things an AppManager client can already automate today: Messenger with call transcription, lead qualification in Prospecting/CRM, proposal templates in the B2B pipeline itself.

What the piece leaves out

  • Forbes recommends generic tooling (Microsoft Copilot Studio) to build these agents — it says nothing about HOW that human approval gets recorded, or who can review it afterward. An "autonomy level" with no auditable record of what a human approved and when is strategic advice, not an actual control mechanism.
  • The article doesn't distinguish between single-process small businesses (an optical shop, a workshop) and ones with several crossing processes (sales + invoicing + inventory) — the real risk of "scaling too fast" is bigger when an agent touches several systems at once, not just one.

AppH's take: we agree with Forbes' advice almost word for word — not because it's convenient for us to say so, but because we built it this way before reading the piece. The real difference is in the fine print: we don't leave human approval as a best practice the small-business owner has to remember to apply — we put it directly in the product's workflow, with a record of who approved what and when. If you're evaluating your first AI agents, the question we'd suggest asking isn't just "what autonomy level should I start at?" but "where does the record live that a human approved this, and can I see it later?" — that's what separates real control from good intentions.

Reviewed by a human at AppH
13 MAY 2026
MARKET

Anthropic launches Claude for Small Business — and "let it run on its own" is an option, not the default

On May 13, Anthropic unveiled a bundle of connectors and 15 ready-to-run agentic workflows (QuickBooks, PayPal, HubSpot, Canva, Docusign) built for U.S. small businesses: plan payroll, close the month, chase overdue invoices, launch a campaign. Anthropic's own core promise: you approve the plan first — or, once you're ready, let it run end-to-end.

The package doesn't replace the tools a business already uses — it installs on top of them: it inherits whatever permissions each employee already had in QuickBooks or Drive, and doesn't train its models on customer data by default on Team/Enterprise plans. In Anthropic's own survey, half of small-business owners named data security as their single biggest hesitation about AI — the launch is built, item by item, to answer exactly that objection.

In favor for AppH

  • Confirms, at Anthropic's scale, something we already build: connecting to what the business already uses (Sirene, our own Warehouse module, our own Invoicing) instead of asking the owner to migrate systems just to automate something.
  • The line "you approve the plan before anything sends, posts, or pays" coming from Anthropic itself — not a third-party vendor — is the strongest validation yet that without explicit human approval, no agentic business workflow is sellable today.

Against / what's missing

  • The option to "let it run end-to-end" without per-step approval is exactly the door we never open, not even as an advanced option for an owner who asks for it: every action with a real consequence (a send, a charge, a stock change) always waits for human confirmation, no exception for accumulated trust.
  • The entire connector stack (QuickBooks, PayPal, HubSpot, Canva, Docusign) is built for the U.S. market — none of them understand French VAT, FEC, or Factur-X, the real fiscal obligations we actually have to solve for a small business in France.

AppH's take: hearing Anthropic itself say "you approve the plan before anything sends, posts, or pays" is the strongest validation we could ask for of our own stance — we don't need to convince anyone a human has to stay in the middle, the company that builds the model is now saying it too. The real difference is in one detail worth looking closely at if you're evaluating tools like this: here, "run end-to-end without asking me anything" is an option the owner can switch on. In AppManager, for any action with a real consequence, that door doesn't exist, and we don't offer it as an advanced option either — not because we doubt Anthropic, but because we'd rather not hand a workshop or optical-shop owner the decision of when to let their guard down.

Reviewed by a human at AppH

Want us to walk you through how this applies to a real case?

Talk to AppH