94% per-step success rate still means 88% of workflows fail
Your agent passes every demo. It handles the happy path perfectly. Then you deploy it to production and it dies within 16 steps.
This isn't a model intelligence problem. It's a reliability architecture problem.
We learned this the hard way when our coding agent had a 94% per-step success rate in testing but only completed 12% of real pull requests. The math is brutal: even 99% per-step reliability gives you just 37% success on a 100-step workflow.
Here's the deployment pattern that fixed it:
1. Separate decision-making from execution
Your agent proposes actions with structured arguments. A deterministic gateway validates, authorizes, and executes them. When the probabilistic part fails, it fails cleanly without corrupting state.
// Agent proposes
{
"action": "create_file",
"path": "src/components/Button.tsx",
"content": "..."
}
// Gateway validates & executes
if (validate_path(path) && check_permissions(user, path)) {
execute_with_rollback(action)
}2. Build state checkpoints into every workflow
Long-running tasks need persistence. When your agent crashes at step 47 of a 60-step deployment, it should resume from step 47, not restart from step 1.
We use LangGraph for state management and Postgres for durable checkpoints. The agent can crash, the API can go down, or we can kill the process for human approval — when it comes back, it knows exactly where it left off.
3. Measure end-to-end completion, not step accuracy
Stop celebrating 95% benchmark scores. Start measuring: "How many complete workflows does this agent finish without human intervention?"
Track failure cascades. One wrong assumption early in a workflow creates 20 minutes of productive-looking debugging that accomplishes nothing.
4. Design explicit handoff points
Your agent shouldn't run until it succeeds or dies. Build structured pause points where humans can review, approve, or redirect without breaking the entire workflow.
CHECKPOINT: Database migration ready - 3 tables to modify - Estimated downtime: 4 minutes - Rollback plan: automated [Approve] [Review Changes] [Cancel]
The reliability trap: Don't just add retry logic. Retries amplify bad decisions. Fix the decision boundary first, then add durability.
Production agents need to survive the messy real world: API outages, permission changes, network timeouts, and humans who go to sleep mid-approval. Build for that reality, not the demo environment.
The agents that survive deployment aren't the smartest ones. They're the ones with the best failure boundaries.