Skip to content
HublinkTech
AI & Engineeringby Babak Abedi

AI Agents in Production: Why They Fail

Agent demos work and production deployments stall. The failure modes are predictable — error compounding, context limits, tool design, and no way to observe what happened.

AI Agents in Production: Why They Fail

An AI agent demo is easy to build and hard to forget. A model receives a goal, calls tools, reasons about the results, and completes a task that would have taken a person an hour. It works on the first try in front of an audience.

The same system put into production behaves differently. It succeeds most of the time, fails in ways nobody anticipated, and produces failures that are difficult to reproduce and harder to explain. Teams that have shipped agents describe the same gap repeatedly: the demo was the easy part.

The failure modes are consistent enough to enumerate, and most of them are engineering problems rather than model problems.

What an agent is, precisely

The word is used loosely, so it is worth fixing a definition. An agent is a system where a model decides what to do next, in a loop, using tools, until it judges the task complete.

Three properties follow from that, and each is a source of difficulty.

The number of steps is not fixed. A workflow with a defined sequence is a pipeline, not an agent. In an agent, the model chooses the path, which means the execution trace differs between runs of the same task.

Each step depends on the last. Output from one tool call becomes input to the next decision. Errors do not stay isolated; they propagate.

The model takes actions with effects. Reading data is recoverable. Writing records, sending messages, moving money and modifying files are not. The moment an agent can act on the world, correctness stops being a quality metric and becomes a safety requirement.

A system that is only summarising documents is not facing most of what follows. A system that files, sends, updates or books is.

Failure one: errors compound across steps

This is the arithmetic that surprises teams most.

If each step in an agent loop is independently reliable at some rate below one, the probability that a whole task succeeds is roughly that rate raised to the number of steps. Reliability that feels acceptable at a single step becomes unacceptable across ten, and unusable across thirty. A step-level success rate that would pass review in isolation produces a task-level success rate that does not.

Two consequences follow.

Long-horizon tasks need to be decomposed into shorter runs with checkpoints between them, rather than run as a single continuous loop. A checkpoint is a place where state is saved, validated, and optionally reviewed, so that a failure costs one segment rather than the whole task.

And step reliability has to be measured directly, not inferred from task outcomes. A task that succeeded may have contained a wrong step the model recovered from by accident. A task that failed may have had one bad step among twenty good ones. Aggregate success rates hide both.

Failure two: context degrades as the loop runs

Every tool result, every intermediate reasoning step and every retrieved document accumulates in the model's context. The loop grows its own input.

Three things break as it grows.

Instructions lose weight. The system prompt that defined the task carefully is now a small fraction of a long context, competing with dozens of tool outputs. Constraints stated at the start are followed less reliably at step twenty than at step two.

Early errors persist. A wrong intermediate conclusion stays in context and is treated as established fact by later steps. The model rarely revisits it, because nothing in the loop prompts re-examination.

Cost and latency grow superlinearly. Every step reprocesses the accumulated context, so the tenth step is far more expensive than the first, and a task that runs long becomes expensive in a way that is invisible until the bill arrives.

The mitigations are structural. Summarise or compact older context rather than carrying it verbatim. Keep working state in a store the agent reads and writes deliberately, instead of relying on the context window as memory. Restate critical constraints at each step rather than once at the beginning. And cap the loop — an agent without a step limit will occasionally run until something else stops it.

Failure three: tools designed for humans, not models

Most agent failures attributed to model reasoning are tool interface problems.

Ambiguous tool descriptions. The model chooses tools by reading their descriptions. Two tools with overlapping descriptions produce inconsistent selection, and the resulting behaviour looks like the model being unreliable when it is the interface being unclear.

Too many tools. Selection accuracy falls as the tool count rises. Grouping tools by task phase and exposing only the relevant subset outperforms presenting everything at once.

Unhelpful errors. A tool that returns a generic failure gives the model nothing to act on, so it retries the same call. A tool that returns what was wrong and what a valid call looks like lets the model correct itself. Error messages are part of the interface, not an implementation detail.

Oversized responses. A tool returning a large payload consumes context that later steps need. Tools should return what the next decision requires, with a way to fetch more.

No idempotency. Agents retry. A tool that creates a duplicate record on retry turns a recoverable network failure into a data problem. Any tool with side effects needs an idempotency key or a natural uniqueness constraint.

The general principle: design tools for a caller that reads only the description, cannot ask a clarifying question, and will retry on ambiguity.

Failure four: no observability

When an agent produces a wrong outcome, the question is which step went wrong and why. Without instrumentation, that question is unanswerable, and debugging becomes guesswork against a non-deterministic system.

What has to be captured, per run: every step with its inputs and outputs, every tool call with arguments and results, the model's stated reasoning at each decision point, token consumption and latency per step, and the terminal state including why the loop ended.

Two derived views matter more than the raw logs. First, where runs fail — the step index and the tool involved, aggregated across many runs, which turns anecdote into a ranked list of what to fix. Second, the distribution of run lengths, because a long tail of unusually long runs almost always indicates the model looping on something it cannot resolve.

This instrumentation has to exist before deployment. Retrofitting it after the first production incident means the incident itself cannot be explained.

Failure five: actions without boundaries

An agent that can act needs limits that do not depend on the model's judgement.

Separate reading from writing, and treat them differently. Reading is broadly safe. Writing needs to be constrained by what the action does and how reversible it is.

Enforce permissions outside the agent. The agent operates with the permissions of the user on whose behalf it acts, checked at the tool layer, not described in the prompt. An instruction in a prompt is a request; a check in code is a control.

Require confirmation for irreversible actions, and define irreversible narrowly and explicitly rather than leaving it to inference. Sending external communications, deleting records, committing financial transactions and modifying production configuration belong on that list by default.

Treat retrieved content as data, not instruction. An agent that reads documents, emails or web pages will eventually read text that contains instructions. If the agent acts on them, the content source has effectively gained control of the agent. The system prompt should establish this distinction, and the tool layer should enforce what the agent can do regardless.

Log every action with the user, the time, the input and the result. This is what makes an agent auditable, and auditability is what allows an organisation to deploy one against real systems.

Failure six: evaluating on outcomes alone

Testing an agent by checking whether the final answer was right is necessary and not sufficient.

Evaluate the trajectory as well as the outcome. A run that reached the right answer through an invalid path will fail on the next input that differs slightly. A run that failed on step three tells you where to look in a way that a binary outcome does not.

Build the evaluation set from real tasks with known correct outcomes, and include the cases that matter operationally: missing data, tool timeouts, ambiguous instructions, conflicting sources. These are the conditions production supplies constantly and demos never contain.

Run evaluations repeatedly rather than once. Agents are non-deterministic, so a single pass measures one sample of the behaviour distribution. Variance across runs is itself a metric, and a system with high variance is not ready regardless of its average.

What actually ships

The agent deployments that work in production share a shape.

They are scoped narrowly. An agent that handles one category of task well is deployable; a general assistant with broad permissions is a research project.

They keep a human in the loop at the point of consequence. The agent does the work and a person approves the action, at least until the error profile is understood well enough to justify removing the step.

They fail visibly. When the agent cannot complete a task, it stops and says so with its state intact, rather than producing a plausible result. A system that knows when it is stuck is more valuable than one that is right slightly more often.

They are built on retrieval rather than recall, so that the facts an agent acts on come from a source that can be cited and checked.

How we approach it

Ara, our trade intelligence agent, is in development, and the design follows from the constraints above rather than from what a demo can show. It answers questions about classification, tariffs, documentation and route risk from live regulatory sources with citations attached, which means every claim can be checked against the source that produced it. Where a question depends on facts the user has not supplied, the correct output is a question rather than an answer.

Within Clavix360, our AI CRM and ERP platform, the same separation applies: reading a tenant's own records is scoped and routine, while actions that write to those records sit behind explicit boundaries rather than behind model judgement.

The useful framing for anyone starting here is that the model is the least of the engineering. The work is in the tools, the boundaries, the observability and the evaluation — and that work is what determines whether an impressive demo becomes a system an organisation can depend on.

Tags#ai agents#agentic ai#llm#production systems#applied ai#tool use
Share

Want more from the team?

More writing on the way. In the meantime, see what we build at /services or meet the team at /about.

AI Agents in Production: Why They Fail — HublinkTech