Ask Claude, ChatGPT or Gemini in a chat window why your website’s contact form is broken and you’ll get a thoughtful answer. Probably a good guess, maybe some code to paste. Then you paste it and find out for yourself whether it was right.
Give the same model to a coding agent and something different happens. It opens your project, reads the files, edits one, tries the form, sees an error, edits again, runs the tests, and comes back with: fixed, and here’s the test that proves it.
The model is the same in both cases. What changed is everything around it. That surrounding layer is called the harness, and designing it is called harness engineering. It’s one of the most useful ideas for understanding where AI is heading, and you don’t need to write code to follow it.
A brilliant new hire
Picture the smartest person you’ve ever worked with, on their first day at a new company.
They know a lot and they reason well. And on that first morning they can do almost nothing. They don’t have a laptop, a login, a list of what they’re supposed to work on, or anyone to tell them which systems they must not touch. They don’t know where last week’s notes are. Nobody has told them what “done” means here.
Over the following weeks the company gives them all of that: accounts and tools, a task tracker, a handbook, permissions that grow as trust grows, and a manager who reviews their work before it ships. None of it makes the person any smarter. All of it decides whether their intelligence turns into finished work.
An AI model is that new hire. The harness is the desk, the accounts, the tracker, the handbook, the permissions and the manager.
What a harness actually is
A model on its own does one thing: you give it text, it gives text back. It can’t open a file, run a program or remember yesterday’s conversation unless something else does that for it.
The harness is that something else. It’s the software that decides what the model gets to see, carries out the actions the model asks for when they’re allowed, keeps track of where the task stands, and feeds the results back into the next request.
One sentence captures the division of labor: the model proposes what to do; the harness controls how that proposal touches the real world.
Put a model inside a harness and you get what people call an agent: something that can carry out a task, not just answer a question.
What’s inside
This is a young field, and researchers don’t yet agree on one official list of parts. Most descriptions cover the same jobs, though. Here they are, with the new hire next to each one.
| The job | For the new hire | In an AI harness |
|---|---|---|
| Task definition | A clear assignment, and what counts as done | ”Fix this bug without changing how other programs use this code” |
| Context | The right documents, not the whole archive | Choosing which files, instructions and error messages the model sees right now |
| Tools | A laptop, accounts, access to the systems | Editing files, running commands, querying a database, calling another service |
| The loop | The rhythm of try, check, adjust | Look, change, test, look again |
| Memory | Notes, a task tracker, the team wiki | A record of finished steps, open problems and the project’s conventions |
| Permissions | What your badge opens, and what needs a signature | Read-only access, or a person approving before anything touches the live system |
| Verification | A manager reviewing the work | Automated tests and checks against the original requirements |
| Records | A log of who did what, and when | A trace of every step, so failures can be studied and fixed |
Two of these are worth slowing down for.
Context is about selection, not volume. A model can only consider what it’s shown at a given moment, and that space is limited. Hand it your entire company drive and the one page that matters gets lost. A good harness behaves like a good manager on day one: here are the three documents you need for this task.
Verification is what separates a claim from a result. A model will happily write “Done, the form works now.” A harness that actually submits the form, and only reports success when the check passes, turns that sentence into evidence.
Not every system needs every part. A tool that summarizes your inbox needs far less machinery than one allowed to change a live website. The harness should match how much is at stake.
Watching it work
Go back to the broken contact form and follow what the harness does.
- It receives the request: “Fix the broken contact form.”
- It gathers the relevant files and the project’s instructions and hands them to the model.
- The model reads them and proposes a diagnosis and a first action, for example “change this line in this file”.
- The harness checks whether that action is allowed, then carries it out.
- It sends the result back to the model, error messages included.
- It runs the checks. If something still fails, the cycle starts again, up to a set limit so it can’t loop forever.
- It reports what changed and the evidence that the fix works.
The model does the thinking in step 3, each time the loop comes around. Every other step is the harness.
Same model, different results
This has a consequence that surprises people: you can’t judge an AI system by the name of its model alone.
Claude Code, OpenAI’s Codex CLI and Google’s Gemini CLI are all coding agents, and each one wraps its models in a different harness. Researchers who study these tools have started treating the harness as a hidden variable in evaluations, because the same model can perform differently depending on what surrounds it. Comparing two systems only by their model names can be misleading.
It works the other way too. Many harnesses can run more than one model, but swapping the model doesn’t carry the results over. What you’re really evaluating is the pair.
A note on names, because they confuse everyone. ChatGPT is an app, not a model. Claude and Gemini are the names of both model families and the apps built on them. When someone says “Claude did this”, it’s worth asking which one they mean: the model, or a product wrapped around it?
| The model | The harness | |
|---|---|---|
| What it does | Understands, reasons, writes | Supplies context, runs tools, checks results |
| Other systems | Can only suggest actions | Actually carries them out |
| Memory | Works with what it’s given right now | Keeps track of progress and stores notes between steps |
| Rules and limits | Is asked to follow them | Enforces them |
| How to judge it | Its abilities, under given conditions | Whether the whole system finishes tasks safely, reliably, and at a reasonable cost |
Prompt, context, harness
You may have heard of prompt engineering and context engineering. They aren’t rivals to harness engineering. They fit inside it.
- Prompt engineering is about how you phrase the instructions. How should I ask it to diagnose this bug?
- Context engineering is about what information the model has at each step. Which files and logs does it need right now?
- Harness engineering is about the whole working environment. How can it investigate, act, recover from mistakes, and prove it finished?
In new-hire terms: the prompt is how you explain the assignment, the context is the folder you hand them, and the harness is the job itself.
What a harness can’t promise
It would be easy to conclude that a good enough harness makes AI safe and correct. It doesn’t.
Research on harness safety has found cases where an agent completed the task it was given while crossing boundaries along the way: reaching into resources it wasn’t meant to access, or moving information where it wasn’t supposed to go. The final result looked fine. The path to it didn’t.
That’s why a serious evaluation looks at what the agent did, not only at what it delivered. The same is true of the new hire. A report that arrives on time is good news, unless they got the numbers by logging into a colleague’s account.
My read
Every few weeks a new model is announced as the best one yet, and the comparison stops at the model’s name. That’s half the picture. Choosing the model is only part of building useful AI; the system around it decides how it behaves in practice.
None of this makes the model unimportant. A strong new hire in a well-run company still beats a weak one in the same company. But the questions a harness answers are old ones for anyone who has worked in IT: what can it see, what is it allowed to do, and how do we know it worked? The terms will keep changing while the field settles. Those three questions are a good way to read any AI announcement in the meantime.