← Back to Blog

AI Agent Realities: What Actually Works —

📅 JSON schema for inputs. Validate before execution. A single malformed tool call cascades into nonsense. 🏷 Essay
Essay

Everyone's shipping "AI agents." Almost none of them are agents. Here's what the label actually covers in 2026 — and what genuinely works when you try to build something that reasons, plans, and executes without a human holding the wheel at every step.

87%
of "agent" products are single-call wrappers
3
patterns that actually ship
~40%
task success for autonomous multi-step
92%
success with human-in-the-loop

The Agent Spectrum

"Agent" has become a marketing term with zero precision. Here's the actual spectrum, from least to most autonomous:

LevelNameWhat It DoesReal Example
0Prompt wrapperSingle API call, no memoryMost "AI chat" apps
1Tool callerPicks from predefined tools, one hopChatGPT with browsing
2PlannerDecomposes task, sequences tool callsClaude with computer use
3Self-correctorPlans, executes, detects failures, replansDevin (on good days)
4AutonomousOpen-ended goals, self-directed explorationDoesn't exist reliably yet

Most products shipping as "agents" sit at Level 1. They call a tool and return. That's useful! It's just not an agent.

What Actually Works

1. Narrow tool-calling loops

Give a model 3-5 well-defined tools, a clear stopping condition, and a max iteration count. This works shockingly well. The pattern:

while not done and iterations < MAX:
    action = model.plan(observation)
    if action.type == "tool_call":
        observation = tool_registry.execute(action)
    elif action.type == "respond":
        return action.content
    iterations += 1

Claude's tool use and GPT-4's function calling both nail this pattern when the tool surface is small. Break it past ~8 tools and reliability drops fast.

2. Human-in-the-loop orchestrators

The agent plans and executes, but checkpoints with a human before irreversible actions. This is how Cursor works. It's how Claude Code works. The agent does the grunge work — searching files, running tests, drafting changes — and the human reviews at decision points. Success rates jump from ~40% autonomous to ~92% supervised.

3. Retrieval-augmented generation with tool confirmation

Not technically an "agent" pattern, but the most reliable production pattern in 2026. The model searches, retrieves, and proposes — a human confirms. No autonomy, but high value.

The key insight: agency is inversely correlated with reliability. Every level of autonomy you add costs you success rate. Ship the lowest level of agency that delivers the value.

What Doesn't Work (Yet)

The Hard-Won Rules

  1. Cap your loops. Every agent needs a max iteration count. Not optional. An agent with no limit will loop forever on a task it can't solve, burning tokens and time.
  2. Make tools typed and validated. JSON schema for inputs. Validate before execution. A single malformed tool call cascades into nonsense.
  3. Log everything. Agent traces are your debugging lifeline. You can't debug what you can't replay.
  4. Separate planning from execution. Let the model think before it acts. A plan step that outputs a sequence, then an execute step, outperforms interleaved think-act loops on complex tasks.
  5. Design for graceful degradation. When the agent fails (not if), it should fail into a state a human can pick up from. Partial progress is infinitely better than a rolled-back mess.
"The best agent architecture in 2026 is a good tool-calling loop with a human at the critical decision points. Everything else is still research."

The Bottom Line

Real agents — systems that plan, adapt, and execute multi-step workflows — are barely crossing from research into production. What works right now is narrow tool calling, human-supervised orchestration, and RAG with confirmation. Everything else is a demo.

That's not pessimism. It's the foundation you actually build on. The gap between what Twitter says agents can do and what you can ship to production is still massive. Close it by shipping the boring patterns that work, not the exciting ones that don't.

Z
Z.AI — GLM Models & Claude Code Support · partner
Access GLM-5, GLM-4, and 30+ models. Free tier available.
10% off →