Agents in Production: The Systems Work That Makes Them Reliable
A production agent is not a clever prompt with tools. It is a durable workflow with boundaries, evidence, observability, and a safe route for uncertainty.
An agent demo produces an impressive answer. An agent in production produces a dependable outcome when networks fail, data is messy, users cancel, and the model is uncertain.
The model is only one component. The dependable part is the system around it: a job that survives a restart, narrow tools, an output contract, limits on parallel work, and enough evidence that a human can understand what happened.
Production agent loop
Make the workflow durable, observable, and reviewable.
Start with a job, not a chat request
Long-running agent work should not live inside a single HTTP request. A user request should create a durable job record containing the input, owner, status, attempt count, timestamps, and an idempotency key. A queue then delivers that job to a worker.
That one decision changes the operating model. Work can resume after a deploy, be retried without creating duplicate side effects, be cancelled deliberately, and expose meaningful progress in the UI. The user sees collecting inputs, analysing batch 4 of 12, or preparing review instead of a spinner with no truth behind it.
request → validate → persist job → enqueue → worker → persist result → notify
↑ │
└──── cancellation / status ┘
State is a product feature
Use explicit states such as queued, running, awaiting_review, completed, failed, and cancelled. They make support, retries, and user communication much less ambiguous.
Idempotency is non-negotiable
Give every side-effecting action a stable key. A retry must not send a second email, create a duplicate ticket, or overwrite a newer result.
Give agents small, typed units of work
One giant prompt over an entire corpus is difficult to observe, expensive to retry, and vulnerable to context limits. Split the input into coherent batches, then give a bounded pool of agents those batches. Keep the concurrency deliberately conservative: upstream APIs, rate limits, databases, and spend are shared resources.
The agent's response should be structured data, not prose that downstream code has to guess at. Define a schema with required fields, enums, size limits, and evidence references. Validate the result before it can affect the rest of the workflow.
type Finding = {
severity: "critical" | "medium" | "info"
title: string
evidence: { sourceId: string; quote: string }[]
recommendation: string
confidence: number
}
Design the failure path before the happy path
Every external call needs a deadline. A timeout is not necessarily a failure of the whole job; it is often a signal to preserve what is known, mark the affected unit for review, and continue. This is preferable to quietly inventing a result or throwing away a successful hour of work because one batch stalled.
Good defaults include:
- A timeout per tool call and per agent batch.
- A bounded retry policy with exponential backoff and jitter.
- A cancellation check between meaningful steps and before side effects.
- A fallback result that clearly says manual review required rather than pretending to be complete.
- A dead-letter or review queue for work that has exhausted safe retries.
Make evidence a first-class output
An agent's conclusion is rarely enough on its own. Store the source excerpt, source identifier, tool inputs and outputs, model/version, prompt version, latency, token usage, and the decision taken. This gives people a way to verify the result—and gives engineers a way to debug regressions.
When batches overlap, the same issue will often appear more than once. Consolidate before presenting results: normalize the evidence, group equivalent findings, retain every affected source, and keep the highest applicable severity. The report gets shorter without becoming less truthful.
Why did the agent say this?
Can we reproduce it?
Did it finish?
What needs a person?
Observability is how you operate an agent
Logs are useful, but a production workflow also needs metrics and traces. At a minimum, measure queue age, completion rate, retry count, timeout rate, model latency, cost per completed job, tool failures, and how often humans override the result.
A heartbeat is particularly valuable for long jobs. If a worker stops updating it, a reconciliation process can safely identify and recover stranded jobs. That is much more useful than discovering a stuck task from a customer email.
- Alert on stale heartbeats, not merely on errors.
- Alert when retries or fallback-to-review rates change materially.
- Sample completed work for evidence quality, not only JSON validity.
- Review cost by successful outcome, not raw token totals.
Tool safety is the boundary that matters
The moment an agent can call a browser, database, filesystem, or third-party API, it becomes a security boundary. Never let a model decide its own permissions. The application should enforce allowlists, scoped credentials, input validation, rate limits, and a confirmation step for irreversible actions.
For web-facing tools, validate URLs and block private network targets to prevent server-side request forgery. For database tools, use purpose-built operations rather than arbitrary query execution. For anything that sends, deletes, pays, or publishes, prefer a human approval gate—or an exceptionally narrow policy with an audit trail.
The practical architecture
There is no universal multi-agent template. Most systems become more reliable when they begin with one worker and add specialization only when it improves a measurable constraint: throughput, domain context, tool isolation, or review quality.
Use a single agent when
The task is short, the tools are few, and one coherent context produces the best result. Add durable state and validation before adding more agents.
Use a worker pool when
Inputs can be split independently, a failed subset should be retried alone, or throughput matters. Keep an explicit cap on concurrency.
Use a reviewer when
Consequences are high, answers must cite evidence, or the first pass is intentionally broad. A reviewer should verify a clear rubric—not just rephrase the first agent.
Use human review when
The cost of a wrong action exceeds the cost of a short delay. Ambiguity is a routing condition, not an embarrassment.
A launch checklist
- Persist the job before performing model or tool work.
- Define allowed state transitions, retry behavior, and idempotency keys.
- Give every tool a narrow contract and least-privilege credentials.
- Require structured output and validate it at the boundary.
- Set input limits, concurrency limits, budgets, deadlines, and cancellation checks.
- Preserve evidence and version the prompts, model settings, and policies.
- Build a review route and a safe fallback before enabling automation.
- Monitor outcome quality, completion, latency, and cost after launch.
The point is not to remove people from the loop at all costs. It is to build a system that lets people spend their attention where judgment matters, while the agent handles bounded, observable work reliably. That is when an agent stops being a feature demo and starts becoming infrastructure.