The clearest map of agent work is well established: five recurring workflow shapes, built on one augmented-LLM building block, topped by autonomous agents. The common advice is to implement each in a few lines of code. That's right about the patterns. It leaves out everything else — the retries, the checkpoints, the approval gates, the metrics, the audit trails.
This note uses that structure as the frame, and at each step adds the layer they skip — the part the agent runtimes leave to you. Every pattern becomes a declarative primitive that runs any CLI agent — opencode, Claude Code, or otherwise — as a single composable, checkpointable, measured step.
The common line is that the patterns are simple — and it's right. The production layer around them is what you've rebuilt three times.
Every workflow here starts from the same standard primitive: an LLM enhanced with retrieval, tools, and memory — a model that can generate its own search queries, pick the right tool, and decide what to remember. The block below assumes each call has those capabilities.
In this engine that primitive is concrete, not assumed. Retrieval is a service; tools are steps; memory is a versioned state store. They are injected, swappable, and mockable — so the same building block runs in CI without an LLM and in production against a real one.
Retrieval → RAG pipeline with citation-validation rails. Tools → any CLI agent, MCP server, or plain Python step. Memory → working / semantic / episodic stores with Redis, SQL, or vector backends. All wired through a single service registry (dependency injection), so every capability is replaceable per environment.
Below, each of the five patterns is framed the standard way — what it is, when to use it, where it shines — and then shown as a first-class primitive in the engine, with the production behavior you would otherwise write from hand.
Decompose a task into a sequence of steps, where each call processes the output of the previous one. Programmatic gates between steps keep the chain on track.
When to use it: the task breaks cleanly into fixed subtasks, and you'll trade latency for higher accuracy by making each call easier.
Examples: generate marketing copy, then translate it; write an outline, validate it against criteria, then write the document from the validated outline.
Primitive: sequential pipeline. You get for free: per-step checkpoint/resume (a 5-call chain that fails on step 4 restarts at step 4), retry with exponential backoff, a typed decision trail of every intermediate output, and automatic cost tracking per link.
Classify an input and direct it to a specialized follow-up. Separation of concerns lets you tune one branch without degrading the others — optimizing for one input type no longer hurts the rest.
When to use it: there are distinct, separable categories and classification can be made accurately.
Examples: split customer queries into general questions, refunds, and technical support; route easy questions to a cheap model (Haiku) and hard ones to a capable one (Sonnet/Opus).
Primitive: expression-evaluator conditional routing on the graph. You get for free: the branch taken is logged for audit, routes can target different models per branch (cost optimization), and a misroute is recoverable because state is checkpointed before the split.
Run calls simultaneously and aggregate programmatically. Sectioning breaks a task into independent subtasks; voting runs the same task many times for diverse outputs. For multi-faceted tasks, separate focused calls beat one overloaded call.
When to use it: subtasks parallelize for speed, or you need multiple perspectives for confidence.
Examples: one instance answers while another screens for policy violations; several prompts review code for different vulnerability classes; multiple voters with a threshold balance false positives and negatives.
Primitive: automatic graph fan-out via topological analysis — independent branches run with asyncio.gather, a ParallelStep merges results. You get for free: one branch failing doesn't crash the others (partial-failure isolation), and every parallel result is folded back into a single checkpointed state.
A central LLM dynamically breaks the task down, delegates the pieces to worker LLMs, and synthesizes their results. Unlike parallelization, the subtasks aren't predefined — the orchestrator decides them from the specific input.
When to use it: you can't predict the subtasks — e.g. the number of files to change and the nature of each change depend on the request.
Examples: coding products that edit many files per task; search that gathers and synthesizes across many sources.
Primitive: hierarchical multi-agent pattern with a manager + specialists. You get for free: each worker can be a real CLI agent (opencode or Claude Code) spawned as an AgentStep, delegation is logged, and the synthesis step is checkpointed as one atomic unit.
One call generates a response while another provides evaluation and feedback, in a loop. It mirrors the iterative revision a human writer goes through for a polished document.
When to use it: there are clear evaluation criteria, and iterative refinement measurably helps — especially when a human articulating feedback would improve the result.
Examples: literary translation where a critic LLM catches nuance; complex search that needs multiple rounds before the evaluator agrees more searching is unwarranted.
Primitive: iterative refinement with loop-with-exit-conditions on the graph (cycles are first-class, with a LoopDetector safety guard). You get for free: bounded iteration counts, each round checkpointed so you can resume or replay from any pass, and convergence metrics tracked automatically.
Past workflows, an agent operates in the open-ended regime: it plans and acts over many turns, using tool results as ground truth at each step, pausing for human input at checkpoints or blockers, and terminating on completion or a max-iteration cap. The implementation is often straightforward — "just LLMs using tools based on environmental feedback in a loop."
Here that loop is a real CLI agent — opencode or Claude Code — running as an AgentStep. The agent supplies the intelligence and the tools; the engine supplies the ground-truth plumbing underneath: state at every step, a human-approval hook at any checkpoint, and a hard iteration cap so autonomy can't run away.
Every workflow above is a shape. The list below is everything you reinvent by hand when you ship that shape — each one a first-class, documented primitive in the engine rather than a side project. This is the showcase; it is deliberately long.
Seven metrics auto-tracked on every execution — the layer regulators and SREs actually need.
SR · success rate CR · compensation/recovery rate
HIR · human-intervention rate MA · model accuracy
PC · pass-confidence TCL · tool-call latency
WCT · workflow completion time
State saved after every step; a failed multi-call pipeline resumes from the failure, not from zero.
backends → SQL · Redis · in-memory
resume · replay · fork a run
save frequency → per-step
Pause the graph, surface a typed approval UI, fire a webhook, resume on the human's decision.
UI schemas · webhooks · timeouts
conditional trigger (e.g. severity ≥ 7)
pending-queue + submit-approval API
Input/output safety as a configurable pipeline, not a post-hoc patch. Runs before and after the model.
input rails → PII · jailbreak · toxicity
output rails · dialog rails · execution rails
retrieval rails → citation validation
Five first-class patterns over a shared, versioned state manager — not manual asyncio wiring.
sequential · parallel · consensus
debate · hierarchical
state rollback + versioning
Retrieval that verifies its own sources — citations checked, hallucinations flagged at the retrieval boundary.
hybrid search · vector + keyword
embedding management · re-ranking
require-citations enforcement
Swap OpenAI, Anthropic, Google, or a local model without touching step code. Cost tracked per call.
100+ providers via LiteLLM
fallback model chains
per-execution cost accounting
Every decision leaves a trail. Dashboards come standard, not as a weekend project.
Prometheus exporter · Grafana dashboards
NDJSON lifecycle telemetry
decision trail · alert rules
Independent branches detected from the graph and run concurrently — no manual concurrency code.
topological fan-out
partial-failure isolation
automatic merge at join nodes
Database, LLM, and HTTP clients shared across steps without globals — and swapped for mocks in tests.
central registration · auto-injection
mockable in CI (no LLM calls)
per-environment overrides
Workflows aren't forced to be DAGs. Bounded cycles power refinement, debate, and retry-until-converged.
loop detector · executor
max-iteration caps
feedback edges with conditions
Extend without forking, drive workflows from the terminal, and give agents working/episodic/semantic memory.
plugin registry · decorators · builtins
CLI: init · run · validate · visualize
memory: working · episodic · semantic
Define the graph in JSON; write only business logic in Python. The framework handles orchestration — so workflows are reviewable, diffable, and runnable in CI without an LLM.
One pipeline can fan out to an opencode voter, a Claude Code executor, and a local model — mixing open and closed, cheap and capable. Vendor lock-in isn't a production feature.
Checkpointing, HITL gates, guardrails, audit trails, and seven reliability metrics — the layer every CLI agent assumes someone else builds. This is that layer.
A coding agent at the terminal is, almost without exception, a single loop with tools against a local repo. They differ in which model they'll talk to and how open the tool layer is. None ship the durability, approvals, metrics, or multi-agent coordination above.
User ─► CLI agent loop ─► one model
│
└─ tools (read/edit/bash/web) ─► local repo
one model · one loop · one process · no durability
User ─► CLI agent loop ─► ANY provider
│
├─ tools ─► local repo
└─ MCP / ACP / skills
any model · one loop · still no multi-step durability
trigger ─► PIPELINE (durable graph)
├─ Step (LLM)
├─ AgentStep ─► opencode ┐ CLI runtimes as leaves
├─ AgentStep ─► claude code ┘
├─ HITL gate (pause / resume)
└─ A2AStep ─► remote agent
│
▼
checkpoint · 7 EARF metrics · guardrails · audit trail
many models · many loops · multi-process · durable · auditable
(A) and (B) are the same shape — opencode only opens the model layer. (C) is a different layer: the CLI runtimes become leaves, and the engine owns everything above them.
| Agent | Model support | License | Role here |
|---|---|---|---|
| Claude Code | Claude only | Commercial | AgentStep leaf |
| opencode | Any provider | Open source | AgentStep leaf |
| Codex CLI | OpenAI + others | Open source | AgentStep leaf |
| Gemini CLI | Gemini | Open source | AgentStep leaf |
| Aider / Goose | Any (LiteLLM) | Open source | AgentStep leaf |
opencode specifically — open source, model-agnostic, with MCP and ACP support, subagents, and skills. The natural open counterpart to Claude Code, and a first-class target, not an afterthought.
| Capability | Claude Code | opencode | This engine |
|---|---|---|---|
| Agent runtime (tools + loop) | yes | yes | via AgentStep |
| Model-agnostic | no (Claude) | yes | yes |
| Durable checkpoint / resume | sessions | sessions | SQL / Redis / memory |
| Human approval gates | inline Q | inline Q | native, UI schemas |
| Reliability metrics | none | none | 7, auto-tracked |
| Guardrails (PII / toxicity) | none | none | built-in |
| Multi-agent collaboration | subagents | subagents | consensus / debate / hierarchy |
| Observability | none | none | Prometheus + Grafana |
| RAG w/ citation rails | via MCP | via MCP | native |
| Cyclic graphs / bounded loops | no | no | first-class |
Net: as a runtime, Claude Code and opencode are more polished. As a production system, this is the layer they both assume someone else will build.
One language. Python only today; no TypeScript surface yet.
No managed tier. Self-host only. A hosted option is on the roadmap.
Indirect MCP. It reaches the Model Context Protocol through the CLI agents it spawns, rather than speaking MCP as a first-party server today.
Adoption is early. Treat the reliability numbers as promises to validate, not benchmarks to cite — yet.
The model is the expensive part. Make the rest disappear.