Guide · AI
Why agentic AI gets cancelled.
Agents are the part of AI a board kills first. Gartner expects more than four in ten agentic projects to be cancelled by 2027, on cost, unclear value, or weak controls. Here is why they die, and the five controls that get one past the board and keep it there.
This guide, in 6 parts
Agents are arriving and getting cancelled at the same time.
Both numbers are true at once. Task-specific agents are forecast to land in a large share of enterprise apps within a year, and more than four in ten agentic projects are forecast to be cancelled by 2027. The capability is spreading; the projects keep dying.
The gap is not the model. It is scope and governance. Agents that survive are scoped to one task with a number and governed to match what they can touch. Agents that die were a demo of autonomy that no one owned, priced, or controlled. That difference is the whole guide.
« An agent with no number is the first line item cut. »
Why agentic projects get cancelled.
Gartner names cost, unclear value and weak controls. On the ground, those show up as five failure modes. Each is avoidable, and visible before launch.
It was a demo of autonomy, not a use case with a number
Most agentic projects start as a proof of concept driven by the hype, not by a metric. When the novelty fades, nobody can name the cost, delay or backlog it was meant to move, so finance has nothing to defend at budget time. An agent with no number is the first line item cut.
The cost grew faster than the value
Agents call models repeatedly, retry, and chain steps, so the bill scales with every task and every loop. Gartner names escalating cost as a lead reason projects die. A pilot that looked cheap on a handful of runs becomes hard to justify once it runs at the volume that would have made it worth doing.
No human was in the loop, so one bad action ended it
An agent that can act on real systems can also act wrong on real systems: a refund issued, a record changed, an email sent. Without a point where a person reviews the exceptions, the first visible mistake reaches a customer, and the project is paused after the incident rather than before it.
Nobody owned the running of it
A demo belongs to whoever built it. An agent in production needs someone who owns it day to day, watches what it does, and is accountable when it drifts. With no named owner, there is no one to fix it when it breaks and no one to defend it when it is questioned, so it quietly stops.
It could not pass governance, so it never left the lab
Without decision logging, traceability and limits on what the agent may touch, it cannot pass internal review, the auditor or the regulator. Gartner also warns that applying one blanket policy to every agent fails on its own; the control has to match how much the agent is trusted to do.
How to ship an agent that survives.
Five moves, in order. None of them is the model, and the first one is the one most teams skip.
Pick one narrow task with a number on it
Not a department, not a role, one bounded task: a single queue, a single decision, a single document type, with a cost or a delay attached. A narrow first use case is the difference between an agent you can prove and a platform you can only promise. If you cannot put a number on the task, it is not the first one to ship.
Keep a human in the loop, by design
Decide up front where a person reviews the agent's work, and on which actions it may never act alone. Start with the agent proposing and a human approving, then widen its autonomy only as the numbers earn it. The human in the loop is not a lack of ambition; it is the thing that lets the agent touch real systems at all.
Put a cost line on it before you scale it
Estimate the per-task cost at the volume you actually intend to run, not the pilot volume, and set the point at which it stops being worth it. Cap the spend, cache what repeats, and watch the bill as a first-class metric. The projects that survive are the ones where someone can show value rising faster than cost.
Name an owner who runs it
One person accountable for what the agent does, with the alerts, the runbook and the authority to pause it. The owner watches for drift, signs off on widening its autonomy, and answers for it in a review. A capability with no owner is a liability waiting for the incident that ends it.
Build governance to match the agent's reach
Decision logging, traceability, and hard limits on what the agent may access and act on, sized to how much it is trusted. A read-only agent and one that can move money are not governed the same way. This is what lets the system pass internal review, the auditor and the regulator, and what lets leadership sign it off without flinching.
The five controls a board signs.
A board does not approve an agent because it is clever. It approves the agent it can see being controlled. Here is what each control has to show.
A narrow first use case
One bounded task with a cost or delay attached, not a platform.
A human in the loop
A named point where a person reviews exceptions and high-risk actions.
A named owner
One accountable person with the alerts, the runbook and the authority to pause it.
A cost line
Per-task cost at real volume, a cap, and value shown rising faster than spend.
Governance that fits its reach
Decision logging, traceability and access limits sized to what it may touch.
The survives-the-board checklist.
An agent is ready to put in front of the board when it can tick all five. Until then, it is a demo waiting to be cancelled.
- One narrow task with a number on it, not a department or a platform.
- A human in the loop on every action the agent may not take alone.
- A named owner who runs it, watches it, and can pause it.
- A cost line at real volume, with a cap and value shown beating spend.
- Governance sized to the agent's reach: logging, traceability, access limits.
Frequently asked.
Why do most agentic AI projects get cancelled?
Gartner expects over 40% to be cancelled by the end of 2027, and names three drivers: escalating cost, unclear business value, and inadequate risk controls. Underneath those, the pattern is the same. The project started as a demo of autonomy rather than a task with a number on it, the cost grew faster than the value, no human reviewed the agent's actions, no one owned the running of it, or it could not pass governance. None of those is a model problem, and all are visible before launch.
How do we scope an agentic AI project so it survives?
Start narrow. Pick one bounded task with a measurable cost or delay, keep a human in the loop on the actions that matter, put a cost line on it at the volume you actually intend to run, name an owner who is accountable for it, and build governance sized to what the agent can touch. A narrow first use case you can prove beats a broad one you can only promise, and it gives finance and the board something concrete to keep funding.
Do we still need a human in the loop for AI agents?
For anything that acts on real systems, yes, at least until the numbers earn the autonomy. The fastest way to get an agentic project paused is a visible wrong action reaching a customer. Start with the agent proposing and a person approving, decide which actions it may never take alone, and widen its autonomy only as the results justify it. Human oversight is what lets the agent touch production at all, not a sign of low ambition.
How is this different from getting an AI pilot to production?
It is adjacent, not the same. Getting any AI pilot live is about wiring it to real systems and data and putting an operational owner on it. Agents add their own failure modes on top: they act, not just answer, so cost scales with every loop, one wrong action can end the project, and governance has to be sized to how much the agent is trusted to do. This guide is about those agent-specific controls and the governance that keeps one alive.
Should we wait for the technology to mature before building agents?
Waiting on everything is not the answer, because task-specific agents are forecast to sit in a large share of enterprise apps within a year, so the capability is arriving whether or not you build it. Waiting on the wrong project is. The honest move is to ship one narrow, well-governed agent on a task with a number, learn what it costs and where it needs a human, and let that earn the next one. Breadth without those controls is how you join the cancelled 40%.
You can scope much of this yourself. Most teams can pick the narrow task, name the owner, and read the checklist without help.
Bring us in for the hard part: choosing the one agent worth shipping first, sizing its governance to what it may touch, and getting it live and signed off. Our Hub Map is how we find the task with a number behind the hype. See the Hub Map, and how we get agents to production in AI in production.
Have an agent the board is hesitating on?
No deck, no gate. A working session on the real use case, with the operators who would get it live and keep it there.