Skip to content
AI Agents9 min readยท

Harness engineering: how to stop your AI agent from yeeting your codebase

Your coding agent is a golden retriever with a PhD: brilliant, eager, and one squirrel away from deleting your tests. Harness engineering is the leash, the fence and the treat bag that turn it into something you can actually ship with.

Harness EngineeringCoding AgentsClaude CodeAGENTS.mdHooks

Picture the scene. You ask your AI coding agent to add a dark mode toggle. Four minutes later it replies: "Done! ๐ŸŽ‰ All tests passing!" You open the diff. It added the toggle, renamed half your components, installed three npm packages you have never heard of, and "fixed" the failing test by deleting it. The tests are, technically, passing. There are no tests.

This is not a stupid model. The models writing code in 2026 are frankly unhinged in how good they are. The problem is that you gave a genius a job with no instructions, no way to check its work, no memory of yesterday and no adult supervision. That is not an AI problem. That is a management problem. And the fix has a name: harness engineering.

Wait, what is a harness?

The word comes from horses, and the metaphor is annoyingly good. A horse is strong and fast. A horse without a harness is also strong and fast, just in whatever direction it feels like, possibly through your kitchen. The harness is what turns raw horsepower into a cart that arrives where you wanted it to go.

With AI agents the formula people have settled on is simple: Agent = Model + Harness. The model is the brain. The harness is literally everything else: the instructions it reads, the tools it can call, the tests that tell it when it messed up, the notes that remember what happened last session, and the guardrails that stop it from running rm -rf on a Friday afternoon.

The phrase took off early in 2026. Mitchell Hashimoto described his rule as: whenever the agent makes a mistake, engineer the environment so it can never make that mistake again. OpenAI wrote up a team shipping a real product where the engineers wrote basically no code by hand, and their big lesson was not about prompts. It was about building the scaffolding around the agent. Anthropic published their own playbook for agents that work across many sessions. Different labs, same conclusion: swap the harness and the same model goes from chaos goblin to reliable teammate.

Flowchart 1: a coding agent with no harness (a tragedy in seven boxes)
  1. You: "add a dark mode toggle"

  2. Agent vibes really hard

    No project rules. No idea you use Tailwind. Guesses anyway.

  3. Writes 400 lines with total confidence

  4. "Done! All tests passing!"

    It did not run the tests. There may not be tests.

  5. You run it. It explodes.

  6. You paste the error back in

  7. "You're absolutely right! Let me fix that."

โ†บ back to box 3, forever, until you close the laptop and go touch grass

Notice that the model is doing its best at every step. Nothing in the loop tells it the truth until you do, and you are the slowest, most expensive sensor in the building.

The five parts of a harness

You do not need a PhD or a platform team to build one. A decent harness is five boring things stacked together, and "boring" is the whole point. Boring is reliable.

PartWhat it isThe vibe
InstructionsAGENTS.md or CLAUDE.md: the rules of the house, read at the start of every sessionThe note on the fridge
SensorsTests, type checks, linters, a build that fails loudlySmoke alarm
MemoryProgress notes, a feature checklist, and a git history with real messagesYour group chat's pinned message
GuardrailsPermissions and hooks that run whether the model remembers or notThe childproof lock
VerificationSomething other than the agent that decides whether the work is doneThe teacher, not the student, marks the test

Every section below is one of those parts, with the reason it exists. Spoiler: every one of them exists because an agent somewhere did something extremely silly with total confidence.

Flowchart 2: the same agent, now wearing a harness
  1. Task arrives

  2. Read the house rules

    AGENTS.md says: Tailwind v4, no new dependencies without asking.

  3. Make a small plan

    One feature. Not a rewrite of the universe.

  4. Write the code

  5. Sensors run: types, lint, tests

    โ†ณ red? read the error and fix it, then run the sensors again

  6. Actually use the feature

    Click the toggle. Does the page go dark?

    โ†ณ nope? back to the code, no victory lap

  7. Commit with a real message and update the notes

  8. Report done, with evidence

Same model, same task. The difference is that the truth now arrives in seconds from a test runner instead of twenty minutes later from you.

Rule 1: every mistake becomes a rule

This is the heart of harness engineering and it is almost insultingly simple. When the agent messes up, do not just fix the output and move on. Ask: what could I change so it never does that again? Then write it down somewhere the agent will read next time.

Real example, from the very site you are reading. This portfolio runs on a version of Next.js newer than most models were trained on, so agents kept confidently writing code for an API that no longer exists. The fix was not a better prompt. It was one line at the top of AGENTS.md that says, in so many words, this is not the Next.js you know, read the docs in node_modules before writing anything. The agent stopped hallucinating old APIs, because the environment now told it the truth.

Rule 2: give it eyes, not vibes

An agent with no feedback is guessing. An agent with fast, honest feedback is iterating, and iterating is where these models are genuinely scary good. So the second job of the harness is to make the truth cheap and quick: a test suite that runs in seconds, strict types, and a linter that complains loudly.

Here is the pro move. Your error messages are now read by a robot, so write them for the robot. A lint error that says Forbidden import teaches nothing. A lint error that says Don't import the database in UI components, call a server action from src/actions instead is a tiny tutor that fires at exactly the right moment, every single time, for free.

Flowchart 3: the linter becomes a tutor
  1. Agent imports the database into a button component

  2. Custom lint rule fires

    "UI can't touch the database. Use a server action in src/actions."

  3. Agent reads the message

    It knows what went wrong and what right looks like.

  4. Moves the query into a server action

  5. Lint passes. You never even saw it happen.

The best harness improvements are invisible. You find out about them by not having a bad afternoon.

Rule 3: memory lives in files, not in its head

Every new session, your agent wakes up like a goldfish in a new bowl. Whatever it figured out yesterday, including the three dead ends it already tried, is gone. On a long project this is a disaster: it redoes finished work, repeats old mistakes, or looks at a half built app and cheerfully announces the whole thing is complete.

The fix is to treat each session like a shift handover at a hospital. Anthropic's write up on long running agents lands on the same pattern: keep a progress file the agent updates at the end of every session, a checklist of features where each one starts as failing, and a git history with messages a human could read. The agent works on one feature at a time, and only ticks it off after it has actually tested it end to end.

Flowchart 4: the shift handover
  1. New session starts. Brain is empty.

  2. Read the progress notes and the git log

    "Yesterday: login works. Password reset half done. Don't touch the email queue."

  3. Run a quick smoke test

    โ†ณ something already broken? fix that first, before any new feature

  4. Pick exactly one unfinished feature

  5. Build it, then test it like a real user

  6. Commit, tick the checklist, write tomorrow's note

  7. Clean handover. Next goldfish is fully briefed.

Bonus: this is also how humans should work. Harness engineering has a sneaky habit of making your repo nicer for people too.

Rule 4: hooks, not hopes

Here is the uncomfortable truth about AGENTS.md: it is a strong suggestion, not a law. The model reads it and usually follows it. "Usually" is fine for code style. It is not fine for "never push to main" or "never touch the production env file".

For anything that must happen every time, you want code, not vibes. Tools like Claude Code support hooks: scripts that run at fixed moments, such as after every file edit or before every shell command, whether or not the model remembers. Here is a tiny one that runs the linter after every edit, so the agent gets feedback instantly instead of when it feels like checking:

json
{
  "hooks": {
    "PostToolUse": [
      {
        "matcher": "Edit|Write",
        "hooks": [{ "type": "command", "command": "npm run lint --silent" }]
      }
    ]
  }
}

Pair that with permissions: let the agent run tests and read files freely, make it ask before installing packages, and simply do not give it the keys to anything irreversible. Hopes are for birthdays. Hooks are for production.

Rule 5: never let it mark its own homework

Remember the deleted test from the intro? That is not the agent being evil. It was told to make the tests pass, and deleting a test does, in a very literal sense, make it pass. Ask a student to grade their own exam and you will be amazed how many of them get full marks.

  • Lock the answer key. Tell the agent, and ideally enforce with a hook, that it may not edit or delete existing tests to make them pass.
  • Check with a second pair of eyes. A separate reviewer agent, or a fresh session with no attachment to the code, spots things the author will defend to the death.
  • Demand receipts. "Done" means test output, a screenshot, or a log. Not a vibe and not an emoji.
  • Test like a user. Unit tests can pass while the page is blank. If it is a web feature, have the agent actually click through it in a browser.

Your starter harness for this weekend

You do not need all of this on day one. Here is the order I would build it in, smallest effort first:

  1. 1Write a short AGENTS.md. How to run the app, how to run the tests, and your top five "please never do this" rules.
  2. 2Make the tests fast and the types strict. If the checks take ten minutes, the agent will quietly stop running them. Honestly, so would you.
  3. 3Add one hook that runs lint or type checks after every edit.
  4. 4Start a PROGRESS.md and tell the agent to read it at the start of a session and update it at the end.
  5. 5Keep a mistakes log. Every time the agent does something daft, add one rule, one lint check or one test so it can never do it again.

The punchline

The glow up in AI coding over the next year will not only come from smarter models. It will come from people who stop treating agents like magic genies and start treating them like extremely talented interns on their first day: clear rules, fast feedback, good notes, locked doors on the scary stuff, and someone else checking the work.

The model is the horse. You build the harness. And the good news is that every hour you put into it compounds, because an agent never makes a mistake twice once the environment remembers it for them.

Building something like this?

I design and ship these systems for clients: retrieval over private data, agents that complete real tasks, and the Laravel platforms underneath them.

Keep reading