Everyone's shipping "AI agents." Almost none of them are agents. Here's what the label actually covers in 2026 — and what genuinely works when you try to build something that reasons, plans, and executes without a human holding the wheel at every step.
"Agent" has become a marketing term with zero precision. Here's the actual spectrum, from least to most autonomous:
| Level | Name | What It Does | Real Example |
|---|---|---|---|
| 0 | Prompt wrapper | Single API call, no memory | Most "AI chat" apps |
| 1 | Tool caller | Picks from predefined tools, one hop | ChatGPT with browsing |
| 2 | Planner | Decomposes task, sequences tool calls | Claude with computer use |
| 3 | Self-corrector | Plans, executes, detects failures, replans | Devin (on good days) |
| 4 | Autonomous | Open-ended goals, self-directed exploration | Doesn't exist reliably yet |
Most products shipping as "agents" sit at Level 1. They call a tool and return. That's useful! It's just not an agent.
Give a model 3-5 well-defined tools, a clear stopping condition, and a max iteration count. This works shockingly well. The pattern:
while not done and iterations < MAX:
action = model.plan(observation)
if action.type == "tool_call":
observation = tool_registry.execute(action)
elif action.type == "respond":
return action.content
iterations += 1
Claude's tool use and GPT-4's function calling both nail this pattern when the tool surface is small. Break it past ~8 tools and reliability drops fast.
The agent plans and executes, but checkpoints with a human before irreversible actions. This is how Cursor works. It's how Claude Code works. The agent does the grunge work — searching files, running tests, drafting changes — and the human reviews at decision points. Success rates jump from ~40% autonomous to ~92% supervised.
Not technically an "agent" pattern, but the most reliable production pattern in 2026. The model searches, retrieves, and proposes — a human confirms. No autonomy, but high value.
plan step that outputs a sequence, then an execute step, outperforms interleaved think-act loops on complex tasks."The best agent architecture in 2026 is a good tool-calling loop with a human at the critical decision points. Everything else is still research."
Real agents — systems that plan, adapt, and execute multi-step workflows — are barely crossing from research into production. What works right now is narrow tool calling, human-supervised orchestration, and RAG with confirmation. Everything else is a demo.
That's not pessimism. It's the foundation you actually build on. The gap between what Twitter says agents can do and what you can ship to production is still massive. Close it by shipping the boring patterns that work, not the exciting ones that don't.