The AI setup that decides what your designs look like
The harness around an AI tool is usually treated as engineering work. I built one to find out what a designer could bring to it.
I first heard the word “harness” at an adventure camp at school, when the instructors explained how one would keep us safe on the ropes. This year I heard it again on an AI project at work, and it brought back memories.
On this project, our design team had moved from designing in Figma to prototyping in Claude Code. I was amazed that the generated screens looked just like the ones we had designed ourselves, with the same components, spacing and style.
Curious about how that was possible, I asked around and learned there was an agent harness working behind the scenes, set up by our engineering team. The harness gave the AI model behind Claude Code our design system components and base templates to build from, so every screen started from them. That got me wondering: what exactly is a harness, and could a designer build one too?
What a harness is
I assumed the harness was just a long prompt with our components and templates pasted in. It isn’t. An AI model on its own answers one request and stops. A harness is the scaffolding around it that keeps it working through a task in steps. After each step, the model decides what to run next based on what it has already found. OpenAI calls it harness engineering; Anthropic writes about harness design. Both mean the setup that turns a model into something that can finish a job.
The harnesses I’ve read about share five parts:
- a loop that runs each step and feeds the result back
- tools the model can use, like search
- memory that carries findings from one step to the next
- checkpoints where a person steps in
- limits the code enforces, whatever the model asks for
The last two are what the camp instructors meant. A harness sets how far you can go and keeps you safe while you do.
Picking a problem I could judge
Once I understood what a harness was, I decided to build one for competitive research. It’s something I do a lot at work, it takes real time, and I know what a good analysis looks like, so I’d be able to tell whether the output was any good.
It’s in no way a replacement for doing the research yourself, going through the flows and screens and building an analysis from your own experience. However, I felt a harness could take on the desk research: pulling what’s available online about each competitor and comparing them.
So I started building it in Claude chat. I began prompting with something as basic as the inputs, deciding how many fields there should be and what each one should ask for. Then I had to work out how those inputs would shape the research, and what the results would look like. After all that back and forth, I had my first harness without writing a single line of code.
How it works
Each research session, which I call a run, has four stages and touches all five parts of a harness. You start by writing a brief describing what you want to research. The model finds competitors, you review the list, and a research loop runs until it produces a report. If the brief names a product to compare against, that product becomes the baseline. In the examples here, that’s Figma.
The loop follows a plan the model writes at the start, researching each competitor in more depth with every step. After each step, the model suggests what to do next.
For tools, I connected Parallel Search inside Claude, since the tool runs in Claude chat as an artifact and can’t search the web on its own. The model writes its own queries, and whatever it finds is saved with a link back to the source.
Those findings go into memory, which keeps the model from going in circles. For each competitor, memory holds a short summary, key facts with sources, and an open question. The model reads this before every step, so each step starts from what it has already found. The final report is written only from what’s in memory.
The last part is limits, which rein in the model and stop the tool from researching endlessly. I learnt that the hard way. My first version had no step limit, and I spent easily five minutes waiting for the results to appear. Now, before a run, you choose how many competitors to study and how deep to go, and that decides how many steps the run gets. Once those steps are used up, the run ends, even if the model wants to keep going. The harness also caps the model at two web searches per step. The part that enforces these limits is called the Guardrail, and it shows up in the log whenever it steps in.
Where I stay in the loop
The Guardrail handles limits, and you handle the most important decision, which is reviewing the list of competitors the model finds. Everything after depends on that list, so before any research starts, you can remove any or add your own. For my brief, comparing pricing and features across design tools with Figma as the baseline, I swapped one competitor for another.
Once the research starts, it continues as expected until the model finds something that makes it want to change the research plan. This could mean researching a competitor out of order, returning to one it has already covered, or writing the report early.
Before doing any of these, the model checks with you first. I added Allow and Don’t allow buttons in the log that ask for your input before it proceeds. I came across this idea in a 1999 paper by Eric Horvitz, which argues that a system should weigh the cost of interrupting you against the risk of acting on its own.
If you’d rather the research run without interruptions, you can switch on “Run unattended.” It’s off by default, and I kept it that way so I could see for myself how and where the model wants to deviate from the research plan.
Designing the interface
With the behaviour settled, the next question was visual, and it got settled iteratively: what should a person actually see?
The setup screen
The setup screen is what you see first, and how you start a research run. The first version was a form, with far too many fields to fill in before you could even begin. I realized that even I wouldn’t want to fill all that out when I just wanted to type something quickly and let the model get started. So I trimmed it down to one input field, with two basic settings: number of competitors and depth of analysis, for the final competitive research report.
Steps panel and log
As the model started its run, I could not tell where it was in the research plan. It needed a visual marker, something that showed progress at a glance rather than making you infer it from the log. So I added the Steps panel on the left, with a flowchart that maps the full pipeline, from finding competitors to the final report, and highlights the current step so you always know where the model is.
On the right is the log. It pairs with the Steps panel. The flowchart shows where the model is, and the log explains what it’s doing there in plain language. It also lets you trace any finding in the final report back to the search that produced it.
I also wanted the log to make it visible whenever the Guardrail overrules the model. This happens when the model wants to deviate from the plan, and the Guardrail keeps it on track instead.
Colour
The first interface Claude produced had a purple button. Of course it did. Purple is the AI tool’s go-to colour. To own the visual direction, I switched to a neutral palette built mostly from black, white, and grey. That meant colour could be used sparingly, reserved only for states that actually need attention. I used coral for when the Guardrail overrules the model, and teal for when a run is waiting on you. Red and blue were the obvious first picks, but both already carry too many meanings in most interfaces.
The report
Once the run finishes, you get a report with an executive summary and a comparative matrix. It ends with a record of what you changed, how often the Guardrail stepped in, and how many searches ran.
The report looked neat and organized at first glance, very compelling. But I didn’t quite trust it. When I verified the pricing details, I found one was incorrect. I also discovered the model had added a competitor on its own, Affinity, which I hadn’t approved (it’s owned by Canva, which I had approved). So for now, the report is a good first draft of the desk research, and I’d still check it before relying on it.
What I learnt
Building this helped me understand how harnesses actually work, even a small one like mine. While digging deeper into the topic, I found that the limits and checkpoints I’d defined could go stale as models improve. A future model might not need a cap of two web searches per step, or it might handle plan changes well enough that the Allow/Don’t allow prompt becomes unnecessary. That means I’ll need to revisit where they sit and what they check for.
On the project where I first encountered a harness, it was set up by engineers, but building one myself showed me that those decisions are design work too.
References
- OpenAI, Harness engineering: leveraging Codex in an agent-first world (2026). https://openai.com/index/harness-engineering/
- Anthropic, Harness design for long-running application development (2026). https://www.anthropic.com/engineering/harness-design-long-running-apps
- Horvitz, Principles of mixed-initiative user interfaces (CHI 1999). https://dl.acm.org/doi/10.1145/302979.303030
- Gibbons, The Four Design Jobs AI Created (Nielsen Norman Group, 2026). https://www.nngroup.com/articles/design-jobs-ai-created/