AI agents fail in production all the time. I've debugged enough of them to see the pattern, and it's almost never the model.
The usual culprits look more like this. A rate limit hits at 3 a.m. and there's no fallback. A tool call runs twice after a crash, so the customer gets two emails. A memory that was true in March is still being served in June. Someone "improves" a prompt, fixes one case and breaks eleven others that nobody checks until a user complains. An agent holds write access to a system it only ever needed to read.
In every one of those cases the model did its job. The system around it didn't.
So the useful question for agentic engineering in 2026 is less "how do I make the agent smarter?" and more "what environment makes the model I already have reliable, observable and safe?" The public work coming out of the big labs points the same way.
OpenAI's harness engineering write-up is the clearest example. Over five months, a team that grew from three to seven engineers shipped a product of roughly a million lines across about 1,500 merged pull requests. Codex wrote every line of it: application logic, tests, CI, docs, observability and internal tooling. The humans spent their time "designing environments, specifying intent, and building feedback loops". Their own summary is humans steer, agents execute.
The agent is not the system
The classic picture of an agent is an LLM in a loop with some tools:
User → LLM → Tool → Observation → LLM → AnswerThat loop is still the right core. In production, though, it sits inside several layers, and each one owns a different problem:
WORK issues, jobs, events, schedules
↓
ORCHESTRATION what runs, when, how many, retries
↓
HARNESS the agent loop: state, context, tool calls, interrupts
↓
CAPABILITIES tools (MCP), skills, memory — scoped and permissioned
↓
SANDBOX where actions actually execute
↓
EVALUATION did the work succeed? → done, retry, or replanI draw it this way to make ownership obvious. The orchestrator has no business reasoning about code, and the agent shouldn't be managing a queue of a thousand jobs. The sandbox doesn't need to know the business goal at all. When something breaks, you want to know which layer is responsible. If everything lives inside one agent.run(task), nobody is.
1. Start with the smallest loop that works
Before adding any of those layers, take the opposite lesson to heart. Most problems need less agent than you think.
mini-SWE-agent makes the case better than any argument. Its agent class is about 100 lines of Python, the model gets nothing but a bash shell, and it still scores above 74% on SWE-bench Verified, with the best runs on the official bash-only leaderboard reaching 76.8%. Extra scaffolding often buys you a harder system to debug and very little else.
My rule is to draw the state machine before writing a single prompt, then look hard at every node. A surprising number aren't reasoning steps. A lookup is a SQL query. A routing decision with three fixed outcomes is an if. Each node you turn into plain code can't hallucinate, costs nothing per call and takes minutes to test.
Agents earn their place where the input is messy: free text, documents, judgement calls. Everywhere else, write code.
2. The harness is where reliability lives
The harness is the runtime around the model. It builds the prompt and context, executes tools, holds state, handles interrupts and retries, and persists everything. Nobody finds it exciting, yet it decides whether your agent makes it through an ordinary Tuesday.
The change that has saved me the most pain is durable execution from day one. Agent runs last anywhere from seconds to hours, and in that time processes crash, deploys restart pods and APIs time out. If the agent's state lives only in memory, each of those events loses work. Sometimes it does something worse and repeats it.
In LangGraph, that means a checkpointer on every graph and side effects wrapped in @task, so a replay returns the recorded result instead of running the action again:
from langgraph.func import task
@task
def send_follow_up(deal_id: str, draft: str) -> str:
# Runs once. On replay after a crash, the cached result is returned
# instead of sending the email a second time.
return email_client.send(deal_id, draft)Without it, you get a familiar bug: the workflow sends an email, crashes before saving its state, resumes from the last checkpoint and sends the email again. Idempotent side effects are what turn that crash into a harmless retry.
There's a bigger pattern here too. Both OpenAI's Codex (a Rust runtime behind a bidirectional JSON-RPC "App Server") and the OpenHands Agent Server (REST and WebSocket, with every action and observation as a serializable event) put a service boundary around the agent runtime. The harness runs as its own process with a protocol in front of it, separate from your application. Once it's packaged like that, an orchestrator can start, pause, inspect and kill an agent the same way it would any other service.
3. Put determinism around the model
LLM APIs go down. Rate limits hit, and latency spikes without warning. A production agent needs a plan for each of these that doesn't depend on the model being available.
Every agent I ship has a deterministic fallback, with a circuit breaker so a struggling provider doesn't get hammered with retries:
@task
async def score_risk(item: Item) -> RiskScore:
if breaker.is_open(): # 3 consecutive failures → 60s cooldown
return rule_based_score(item)
try:
result = await risk_agent.ainvoke(item)
breaker.record_success()
return result
except (RateLimitError, TimeoutError):
breaker.record_failure()
return rule_based_score(item)The rule-based score is cruder than what the model produces. I'm fine with that. A slightly worse answer beats an error page, and the system keeps working while the provider recovers.
More generally, I make the system strict wherever strictness is cheap, and I save the model for places that need judgement:
| Deterministic | Model |
|---|---|
| Permissions, timeouts, budgets | Classifying messy input |
| Workflow state and retries | Planning and decomposition |
| Running tests, validation, schemas | Drafting text or code |
| Security policy, approval gates | Diagnosis and summarisation |
Structured output belongs in the left column. If the next step parses what the model returns, have the model return a schema and validate it. Regex over prose will let you down eventually.
4. Human approval is an architecture decision
Teams often bolt "human in the loop" on at the end as a UI feature. That's too late. It has to live in the state machine, because what you're really deciding is where the agent stops.
I follow one rule: the agent prepares and a person commits. That applies to anything with consequences outside the system, such as payments, messages to customers, orders or deletions. The agent does everything up to the decision, then the graph interrupts and persists its state until someone approves:
from langgraph.types import interrupt
def place_order(state: State) -> State:
decision = interrupt({"proposed_order": state["draft_order"]})
if decision["approved"]:
return {"order_id": erp.create_order(state["draft_order"])}
return {"status": "rejected", "reason": decision.get("reason")}Since the interrupt is durable, the answer can come back in five seconds or five days. The rejections are useful as well. A "no" with a reason attached is a labelled example of the agent getting something wrong, and it goes straight into the evaluation set (more on that in section 7).
5. Scope capabilities like you'd scope a junior's access
Whatever an agent can touch (tools, data, network) is its attack surface. In my experience, the common failure is rarely malicious. It's an agent with broad credentials doing something that looks reasonable, in the wrong place.
Here's what I apply to every system.
Each tool gets the least privilege it needs. A tool that reads orders gets a read-only credential, and write access lives in a separate tool behind an approval gate.
Code and shell actions run in a sandbox with an explicit network policy and an audit trail of every command. Doing this by hand stops scaling at about two agents, which is why I built nemoclaw-hub to manage several sandboxed agents from one place.
MCP is the tool boundary, so it's also where policy belongs. The Model Context Protocol standardises how agents reach tools, and it comes with a security model worth reading. Two mistakes show up again and again: passing the user's token straight through to downstream APIs, and the confused deputy, where a server with broad privileges acts for a caller who shouldn't have them.
When several customers share one agent, isolate them at the storage layer. Checkpoints, memories and stores all need hard separation, because a WHERE tenant_id = ? that someone forgets is a data leak. langgraph-tenancy handles that isolation, and per-tenant usage metering lives in the same place.
6. Context has to be engineered
The default instinct is to give the model more: longer system prompts, full chat histories, every document you have. That breaks in three ways. Too much context costs money and distracts the model. Too little and it forgets what it was doing. The worst case is wrong context, because a stale fact injected with confidence is something the agent will act on.
Three practices have held up for me.
The first is to split knowledge into modules. Agent Skills, which Anthropic introduced in October 2025 and later published as an open standard that Codex and other tools support, are folders with a SKILL.md plus scripts and references the agent loads only when they're relevant. Specialisation becomes files you can version, review and test. Repository context works the same way: an AGENTS.md and architecture docs the agent can read, with the important constraints enforced by linters and tests. If a rule matters, a machine should be able to check it.
The second is to treat memory as data with a lifecycle. Prices change, owners change, statuses change, and a memory layer that only appends will eventually contradict itself. The design I've settled on works like this. A new value supersedes the old one. A removal retracts the fact, which is softer than deleting it. History stays queryable, so you can still answer "what did we believe in March?". That's the idea behind vayl. Checking that stored facts are still true is a separate job, and MemGuard does it. The tiers, compaction and decay are covered in my post on production memory architecture.
The third is a budget for the context window. Give instructions, the task, retrieved documents and memories their own token allowance, and decide ahead of time what gets dropped when you overflow. If you don't decide, the thing that gets dropped is whatever happened to arrive last.
7. Evaluation is the definition of "done"
The loop I see most often goes like this: change the prompt, try three examples, decide it "feels better", ship. That's testing in production with extra steps.
What works is treating agent behaviour like code under test:
production traces → failures → dataset → evaluator → fix → re-run all → deployReal failures make the best test cases. On one matching agent, every recommendation a user rejected went into the evaluation dataset. I wrote a targeted evaluator for each failure type and re-ran the whole set before every change. Over three months, matching accuracy climbed from 72% to 91%. The model stayed the same the whole time. We just never shipped the same mistake twice. (The tooling for this is in my LangSmith guide.)
Prompts deserve the same discipline, since a prompt change is a code change and needs a regression suite in front of it. I built prompt-diff for exactly that: a rewritten prompt has to pass a YAML test suite before it can replace the old one. Superpowers applies the same thinking to agent skills. You watch the agent fail without the skill, write the skill, then confirm the behaviour actually changed. It's TDD for agent behaviour, and I think it's the right mental model.
Two more things I've learned here.
Check how the agent got there as well as what it produced. For a coding agent, a passing final diff tells you little on its own. Did it touch the right files? Did it preserve existing behaviour? Did it stop when approval was required? You can answer most of those from the trace.
A small model with good calibration often matches a big one. On a scoring agent, I moved from a large model to a small one with three few-shot calibration examples and got the same quality on the eval set for about 85% less. I only knew that switch was safe because the eval set existed.
8. Orchestration: when one agent becomes many
Everything up to here is about making a single agent reliable. Orchestration starts when you have many agent runs going at once against a stream of work, and the questions stop being about reasoning and start being about operations:
Which work items are eligible? How many run at once?
What happens when one fails? What if the work item changes mid-run?
Which workspace belongs to which run? How do we recover after a crash?You have two broad options.
One is a purpose-built work orchestrator. OpenAI's Symphony is the clearest public example for coding work. It polls issues from a tracker (Linear, in OpenAI's setup), gives each issue an isolated workspace and its own coding-agent run, and handles concurrency, retries and reconciliation. The spec says outright that it isn't meant to be a general-purpose workflow engine, and OpenAI released it as an engineering preview. I like its framing a lot: manage work instead of supervising coding agents.
The other is a general durable-execution engine. Temporal and Conductor record each step in an event history, so a crashed workflow picks up exactly where it stopped. Neither is an agent framework. They're the deterministic layer you run agents inside. Conductor can even call agents built with LangGraph, OpenAI Agents or Google ADK as a native AGENT task, and it records MCP tool discovery and tool calls as workflow steps you can inspect.
This also clarifies where LangGraph fits. It's excellent inside the agent layer, when one agent needs stateful, graph-shaped control. It doesn't have to be your whole platform.
Agents built on different frameworks can talk to each other over the A2A protocol, which handles discovery through Agent Cards and task updates through polling, streaming or push notifications. MCP covers the agent-to-tool side. You'll want A2A once you genuinely have different kinds of agents talking to each other, which tends to happen much later than architecture diagrams suggest.
What to adopt, and when
Adopting all of this at once is the fastest way to build something fragile. Here's what I'd actually put in place at each stage:
| Stage | Adopt | Skip for now |
|---|---|---|
| One agent, first users | Deterministic nodes where possible, durable state, idempotent side effects, an approval gate, tracing | Orchestrators, A2A, multi-agent |
| Agent in daily use | Evaluation dataset from real failures, prompt regression tests, fallbacks and circuit breakers, scoped credentials | Custom orchestration |
| Multiple agents or tenants | Tenant isolation, sandboxing with network policy, memory with supersession, cost budgets | A2A unless agents are truly heterogeneous |
| Many concurrent runs over a work queue | A durable orchestrator (Temporal, Conductor, or a Symphony-style work manager) | Nothing at this point |
The takeaway
Models matter, and they keep improving. Still, the biggest production gains I've seen rarely came from swapping one model for another. They came from turning fake reasoning steps into code, making every side effect durable and idempotent, and stopping the agent before anything consequential. They came from giving each agent exactly the access it needs, treating context and memory as data you design, and refusing to ship a change the evaluation set hasn't seen.
Building the agent is the easy part now. The harder skill is deciding where its responsibility should end and where the surrounding system takes over.
Sources
- OpenAI, Harness engineering: leveraging Codex in an agent-first world, February 2026
- OpenAI, Unlocking the Codex harness and openai/codex
- OpenAI, Symphony and announcement, April 2026
- OpenHands Software Agent SDK
- mini-SWE-agent and the SWE-bench leaderboard
- Agent Skills specification
- obra/superpowers
- Temporal, Event History
- Conductor OSS and its agent framework recipes
- A2A protocol specification