Harness, loop, graph: three agent problems that need three different fixes
Harness engineering, loop engineering and graph engineering keep getting used as synonyms. They are three separate decisions: what the model can touch, what happens after the call, and where it is allowed to go next. Fixing the wrong one is how a week disappears.
Three phrases have been doing the rounds: harness engineering, loop engineering, graph engineering. They turn up in the same sentences, usually as synonyms, usually by people who are each describing something real and slightly different. A framing posted by beamnxw on X finally pulled them apart for me, and the separation is worth having, because the three terms name three different fixes for three different failures.
Shortest version I can give you: the harness is the environment, the loop is the feedback cycle, the graph is the flow. All three sit around the same model. All three can contain something that looks like a loop. All three decide whether the thing works on a Tuesday. They are still not the same decision, and the whole value of the vocabulary is being able to say which one you are currently getting wrong.
| Layer | The question it answers | What you actually build |
|---|---|---|
| Harness | What can the model touch? | House rules, tool definitions, memory, files, sandboxes, permissions, hooks, traces |
| Loop | What happens after the call? | Goal, state, evidence, compact feedback, retry limits, stop rules |
| Graph | What is allowed to happen next? | Nodes, edges, conditional branches, joins, human gates, checkpoints |
Harness: what the model can touch
A model on its own cannot read your repo, run your test suite, look at a browser, remember yesterday, or ask a human for approval. Every one of those is a capability the environment lends it. The harness is that environment: the system prompt and the house rules, the tool definitions, the memory and the filesystem, the sandbox, the permission model, the hooks that fire whether or not the model remembers them, and the traces that let you work out afterwards what actually happened.
I have written a whole post on building one, so I will not relitigate it here. The one-line test is this: if the agent failed because it could not see something or could not do something, that is a harness problem. No amount of prompt wrangling conjures a missing tool.
The classic harness mistake is the junk drawer β forty tools bolted on because each looked useful in isolation. Selection errors go up, reliability goes down, and the model burns its attention choosing rather than working. A harness is a set of deliberate affordances, not an inventory.
Loop: what happens after the call
A prompt tells the model what to do during a call. A loop decides what the system does after it.
That is the distinction most people skate past. Prompting is a single-turn concern. Looping is about the cycle: the agent acts, the world answers back, and something has to decide whether to continue, change approach, roll back to an earlier state, call for help or stop. In practice that decision is where most of an agent's apparent intelligence lives, and it is yours to design, not the model's.
A loop that survives contact with production has a small number of named parts: a trigger, a goal specific enough to be checked by something other than the model, the state the next pass needs, a policy for what the agent is allowed to do, evidence that the goal was met, compact feedback when it was not, and a stop rule. The stop rule is the part everyone skips, and it is the part that turns up on the invoice.
const result = await runUntil({
// Checkable by something that is not the agent.
goal: () => suite.passing && contract.validates(output),
act: (feedback) => agent.run(task, feedback),
// Every exit is real. "The agent said it was done" is not one of them.
limits: { attempts: 4, seconds: 600, usd: 2 },
onExhausted: (trace) => escalateToHuman(trace),
});State the goal as something checkable
"Tests green and the endpoint returns 200", not "improve the code".
Agent acts
Collect evidence
Test output, a validated schema, a citation that resolves. Not an opinion.
Goal met?
β³ no? compact the failure into feedback, increment the counter, go round again
Attempts, budget or clock exhausted?
β³ yes? stop and escalate to a human, with the trace attached
Done, with receipts
βΊ and that counter is the only reason this ever terminates
The shape is ordinary. What makes it something you can leave running unattended is that both exits are real: one on evidence, one on exhaustion.
Graph: what is allowed to happen next
A harness and a loop still leave the agent free to pick its own route. Sometimes that freedom is the entire point. Sometimes it is a liability, and you want the route written down: these steps, in this order, with this branch, and a human signs here. That is graph engineering β modelling the workflow as an explicit directed graph or state machine, where nodes are steps and edges are the transitions you are prepared to permit.
A node does not have to be a model call. The good graphs mix them: a deterministic function here, a specialist agent there, a human review gate before anything leaves the building, a retry policy per node, and checkpoints so a failure at step six does not mean redoing steps one through five. LangGraph is the best known tool for this, but the discipline matters more than the library β a state machine and a switch statement get you most of the way.
Intake: validate and scope the request
Deterministic node. No model involved.
Route by request type
β³ out of scope? reject here, cheaply, before a single token is spent
Research node β agentic, with its own loop inside
Free to iterate, but only within this node.
Screening node: does the evidence hold up?
β³ thin? back to research, two passes maximum
Drafting node
Human gate
The one step nobody gets to delegate.
β³ changes requested? back to drafting, with the notes carried as state
Publish
Every arrow is a decision you made once, at design time, instead of one the model re-makes on every run and gets right most of the time.
They all contain loops. That is why they get confused.
Here is the knot. A harness has retries. A loop is a cycle by definition. A graph can have a cycle edge. Three different things, all drawn with an arrow that points backwards.
The way out is to ask what the arrow is made of. A harness retry is infrastructure: the call timed out, try it again, nothing about the task has changed. A loop iteration is learning: the attempt produced evidence, and the evidence changes the next attempt. A graph cycle is topology: this node is permitted to hand control back to that one, and the edge exists in the diagram whether or not anybody travels it today.
Diagnose first, build second
This is the practical payoff of keeping the three apart. When an agent misbehaves, the symptom usually tells you which layer is at fault β and working on the wrong layer is exactly how a week disappears.
| What you are seeing | Layer at fault | The fix |
|---|---|---|
| It had no idea about a rule, a file or a convention | Harness | Put the truth in the environment: house rules, a doc it genuinely reads, a tool |
| It tried to do something and simply could not | Harness | A missing tool or permission. Not a prompt problem. |
| First attempt is close, but it is a coin flip run to run | Loop | Add an evidence check and let it iterate against that |
| It retries forever, or declares victory on nothing | Loop | Stop rules: attempt limit, budget, timeout, escalation path |
| Five specialists, an approval and three branches, all improvised | Graph | Draw the state machine and make the transitions explicit |
| Step six fails and the whole run restarts from scratch | Graph | Checkpoints, and a retry policy per node |
The agent produced something wrong
Could it even see and do what the task required?
β³ no? harness. Stop here β the other two layers cannot rescue you.
Was it told it was wrong, soon enough to act on it?
β³ no? loop. You have a long prompt, not a feedback cycle.
Was it free to take a route it should never have been offered?
β³ yes? graph. Constrain the transitions.
Only now is it fair to blame the model
You will not often get this far.
Models get blamed for orchestration bugs constantly. A broken API, stale state, a missing exit condition and a vague schema all look like stupidity from the outside.
The order to build them in
- 1Harness first, always. Until the agent can see the truth and act on it, nothing downstream matters.
- 2Then the loop, the moment the first attempt stops being reliable. One goal, one evidence check, one stop rule. That is an afternoon's work.
- 3Graph last, and only once the shape of the work has stopped moving. Branches, specialists, approvals and genuinely parallel paths are the signal. One agent with four tools is not.
- 4Keep evaluation outside all three. Trace replay, version comparison and a success rate you can plot. Otherwise you are tuning three layers by vibes and calling it architecture.
The punchline
None of this is about the model, which is the uncomfortable bit if you were hoping the next release would sort it out. The differentiator in a production agent is the system around the model, and that system has three distinct parts with three distinct failure modes.
Harness, loop, graph. Environment, feedback, flow. Get into the habit of naming which one you are fixing before you open the editor, and you will stop losing Tuesdays to rewriting a prompt that was never the problem.
Building something like this?
I design and ship these systems for clients: retrieval over private data, agents that complete real tasks, and the Laravel platforms underneath them.