Back to Blog
AI AgentsLLMProduction SystemsBackendReliability

Agents in Production: The Systems Work That Makes Them Reliable

A production agent is not a clever prompt with tools. It is a durable workflow with boundaries, evidence, observability, and a safe route for uncertainty.

D
Dheeraj Jha
September 29, 2026
7 min read

An agent demo produces an impressive answer. An agent in production produces a dependable outcome when networks fail, data is messy, users cancel, and the model is uncertain.

The model is only one component. The dependable part is the system around it: a job that survives a restart, narrow tools, an output contract, limits on parallel work, and enough evidence that a human can understand what happened.

Production agent loop

Make the workflow durable, observable, and reviewable.

1. Durable jobPersist intent before model work begins.
2. Bounded agent poolSmall batches, explicit concurrency, time limits.
3. Validate & consolidateSchema checks, evidence links, deduplication.
4. Act or reviewAutomate safe actions; route ambiguity to people.
Heartbeats expose stalled work
Retries are bounded and idempotent
Tools run with least privilege

Start with a job, not a chat request

Long-running agent work should not live inside a single HTTP request. A user request should create a durable job record containing the input, owner, status, attempt count, timestamps, and an idempotency key. A queue then delivers that job to a worker.

That one decision changes the operating model. Work can resume after a deploy, be retried without creating duplicate side effects, be cancelled deliberately, and expose meaningful progress in the UI. The user sees collecting inputs, analysing batch 4 of 12, or preparing review instead of a spinner with no truth behind it.

request → validate → persist job → enqueue → worker → persist result → notify
                         ↑                         │
                         └──── cancellation / status ┘

State is a product feature

Use explicit states such as queued, running, awaiting_review, completed, failed, and cancelled. They make support, retries, and user communication much less ambiguous.

Idempotency is non-negotiable

Give every side-effecting action a stable key. A retry must not send a second email, create a duplicate ticket, or overwrite a newer result.

Give agents small, typed units of work

One giant prompt over an entire corpus is difficult to observe, expensive to retry, and vulnerable to context limits. Split the input into coherent batches, then give a bounded pool of agents those batches. Keep the concurrency deliberately conservative: upstream APIs, rate limits, databases, and spend are shared resources.

The agent's response should be structured data, not prose that downstream code has to guess at. Define a schema with required fields, enums, size limits, and evidence references. Validate the result before it can affect the rest of the workflow.

type Finding = {
  severity: "critical" | "medium" | "info"
  title: string
  evidence: { sourceId: string; quote: string }[]
  recommendation: string
  confidence: number
}

Design the failure path before the happy path

Every external call needs a deadline. A timeout is not necessarily a failure of the whole job; it is often a signal to preserve what is known, mark the affected unit for review, and continue. This is preferable to quietly inventing a result or throwing away a successful hour of work because one batch stalled.

Good defaults include:

  • A timeout per tool call and per agent batch.
  • A bounded retry policy with exponential backoff and jitter.
  • A cancellation check between meaningful steps and before side effects.
  • A fallback result that clearly says manual review required rather than pretending to be complete.
  • A dead-letter or review queue for work that has exhausted safe retries.

Make evidence a first-class output

An agent's conclusion is rarely enough on its own. Store the source excerpt, source identifier, tool inputs and outputs, model/version, prompt version, latency, token usage, and the decision taken. This gives people a way to verify the result—and gives engineers a way to debug regressions.

When batches overlap, the same issue will often appear more than once. Consolidate before presenting results: normalize the evidence, group equivalent findings, retain every affected source, and keep the highest applicable severity. The report gets shorter without becoming less truthful.

Why did the agent say this?
Show the exact evidence and the rule or policy used.
Can we reproduce it?
Keep versioned prompts, model settings, and tool traces.
Did it finish?
Report progress, heartbeats, duration, and terminal status.
What needs a person?
Make uncertainty and review ownership explicit.

Observability is how you operate an agent

Logs are useful, but a production workflow also needs metrics and traces. At a minimum, measure queue age, completion rate, retry count, timeout rate, model latency, cost per completed job, tool failures, and how often humans override the result.

A heartbeat is particularly valuable for long jobs. If a worker stops updating it, a reconciliation process can safely identify and recover stranded jobs. That is much more useful than discovering a stuck task from a customer email.

Operating checklist
  1. Alert on stale heartbeats, not merely on errors.
  2. Alert when retries or fallback-to-review rates change materially.
  3. Sample completed work for evidence quality, not only JSON validity.
  4. Review cost by successful outcome, not raw token totals.

Tool safety is the boundary that matters

The moment an agent can call a browser, database, filesystem, or third-party API, it becomes a security boundary. Never let a model decide its own permissions. The application should enforce allowlists, scoped credentials, input validation, rate limits, and a confirmation step for irreversible actions.

For web-facing tools, validate URLs and block private network targets to prevent server-side request forgery. For database tools, use purpose-built operations rather than arbitrary query execution. For anything that sends, deletes, pays, or publishes, prefer a human approval gate—or an exceptionally narrow policy with an audit trail.

The practical architecture

There is no universal multi-agent template. Most systems become more reliable when they begin with one worker and add specialization only when it improves a measurable constraint: throughput, domain context, tool isolation, or review quality.

Use a single agent when

The task is short, the tools are few, and one coherent context produces the best result. Add durable state and validation before adding more agents.

Use a worker pool when

Inputs can be split independently, a failed subset should be retried alone, or throughput matters. Keep an explicit cap on concurrency.

Use a reviewer when

Consequences are high, answers must cite evidence, or the first pass is intentionally broad. A reviewer should verify a clear rubric—not just rephrase the first agent.

Use human review when

The cost of a wrong action exceeds the cost of a short delay. Ambiguity is a routing condition, not an embarrassment.

A launch checklist

Operating checklist
  1. Persist the job before performing model or tool work.
  2. Define allowed state transitions, retry behavior, and idempotency keys.
  3. Give every tool a narrow contract and least-privilege credentials.
  4. Require structured output and validate it at the boundary.
  5. Set input limits, concurrency limits, budgets, deadlines, and cancellation checks.
  6. Preserve evidence and version the prompts, model settings, and policies.
  7. Build a review route and a safe fallback before enabling automation.
  8. Monitor outcome quality, completion, latency, and cost after launch.

The point is not to remove people from the loop at all costs. It is to build a system that lets people spend their attention where judgment matters, while the agent handles bounded, observable work reliably. That is when an agent stops being a feature demo and starts becoming infrastructure.