Notes

The AI setup that decides what your designs look like

ai, agents
A model sits at the centre of a loop. On each turn it runs a web search, files what it found in memory, and waits at a checkpoint for a person to press Continue. The loop sits inside a boundary of limits with a step meter.

The harness around an AI tool is usually treated as engineering work. I built one to find out what a designer could bring to it.

I first heard the word “harness” at an adventure camp at school, when the instructors explained how one would keep us safe on the ropes. This year I heard it again on an AI project at work, and it brought back memories.

On this project, our design team had moved from designing in Figma to prototyping in Claude Code. I was amazed that the generated screens looked just like the ones we had designed ourselves, with the same components, spacing and style.

Curious about how that was possible, I asked around and learned there was an agent harness working behind the scenes, set up by our engineering team. The harness gave the AI model behind Claude Code our design system components and base templates to build from, so every screen started from them. That got me wondering: what exactly is a harness, and could a designer build one too?

What a harness is

I assumed the harness was just a long prompt with our components and templates pasted in. It isn’t. An AI model on its own answers one request and stops. A harness is the scaffolding around it that keeps it working through a task in steps. After each step, the model decides what to run next based on what it has already found. OpenAI calls it harness engineering; Anthropic writes about harness design. Both mean the setup that turns a model into something that can finish a job.

The harnesses I’ve read about share five parts:

  • a loop that runs each step and feeds the result back
  • tools the model can use, like search
  • memory that carries findings from one step to the next
  • checkpoints where a person steps in
  • limits the code enforces, whatever the model asks for

The last two are what the camp instructors meant. A harness sets how far you can go and keeps you safe while you do.

A research agent harness drawn as a climbing harness, labelled with its five parts: the loop, tools, memory, checkpoints and limits.
Anatomy of a harness, drawn as a climbing one.

Picking a problem I could judge

Once I understood what a harness was, I decided to build one for competitive research. It’s something I do a lot at work, it takes real time, and I know what a good analysis looks like, so I’d be able to tell whether the output was any good.

It’s in no way a replacement for doing the research yourself, going through the flows and screens and building an analysis from your own experience. However, I felt a harness could take on the desk research: pulling what’s available online about each competitor and comparing them.

So I started building it in Claude chat. I began prompting with something as basic as the inputs, deciding how many fields there should be and what each one should ask for. Then I had to work out how those inputs would shape the research, and what the results would look like. After all that back and forth, I had my first harness without writing a single line of code.

The tool's empty start screen with a single text field for the research brief.
The finished tool ready for a brief.

How it works

Each research session, which I call a run, has four stages and touches all five parts of a harness. You start by writing a brief describing what you want to research. The model finds competitors, you review the list, and a research loop runs until it produces a report. If the brief names a product to compare against, that product becomes the baseline. In the examples here, that’s Figma.

The loop follows a plan the model writes at the start, researching each competitor in more depth with every step. After each step, the model suggests what to do next.

For tools, I connected Parallel Search inside Claude, since the tool runs in Claude chat as an artifact and can’t search the web on its own. The model writes its own queries, and whatever it finds is saved with a link back to the source.

Those findings go into memory, which keeps the model from going in circles. For each competitor, memory holds a short summary, key facts with sources, and an open question. The model reads this before every step, so each step starts from what it has already found. The final report is written only from what’s in memory.

The last part is limits, which rein in the model and stop the tool from researching endlessly. I learnt that the hard way. My first version had no step limit, and I spent easily five minutes waiting for the results to appear. Now, before a run, you choose how many competitors to study and how deep to go, and that decides how many steps the run gets. Once those steps are used up, the run ends, even if the model wants to keep going. The harness also caps the model at two web searches per step. The part that enforces these limits is called the Guardrail, and it shows up in the log whenever it steps in.

Start screen with a brief typed in asking to compare pricing and features across design tools against Figma, with the competitor count and depth settings below.
The competitive analysis tool’s start screen with the brief I wrote.

Where I stay in the loop

The Guardrail handles limits, and you handle the most important decision, which is reviewing the list of competitors the model finds. Everything after depends on that list, so before any research starts, you can remove any or add your own. For my brief, comparing pricing and features across design tools with Figma as the baseline, I swapped one competitor for another.

The competitor review screen, with Adobe XD removed and Framer added to the list the model found.
The run pauses after finding competitors. I removed Adobe XD, which Adobe no longer actively develops, and added Framer, which the model missed.

Once the research starts, it continues as expected until the model finds something that makes it want to change the research plan. This could mean researching a competitor out of order, returning to one it has already covered, or writing the report early.

Before doing any of these, the model checks with you first. I added Allow and Don’t allow buttons in the log that ask for your input before it proceeds. I came across this idea in a 1999 paper by Eric Horvitz, which argues that a system should weigh the cost of interrupting you against the risk of acting on its own.

The log with the model asking to research Sketch ahead of plan, and Allow and Don't allow buttons.
The model asks to research Sketch ahead of plan. Doing it now changes the order.

If you’d rather the research run without interruptions, you can switch on “Run unattended.” It’s off by default, and I kept it that way so I could see for myself how and where the model wants to deviate from the research plan.

Steps panel with the Run unattended toggle switched off.
“Run unattended” toggle in the Steps panel.

Designing the interface

With the behaviour settled, the next question was visual, and it got settled iteratively: what should a person actually see?

The setup screen

The setup screen is what you see first, and how you start a research run. The first version was a form, with far too many fields to fill in before you could even begin. I realized that even I wouldn’t want to fill all that out when I just wanted to type something quickly and let the model get started. So I trimmed it down to one input field, with two basic settings: number of competitors and depth of analysis, for the final competitive research report.

The original setup form with technical field labels, beside the new setup screen with a brief field and two choices.
Before, a form with many fields. After, a brief and two choices: how many competitors and how deep.

Steps panel and log

As the model started its run, I could not tell where it was in the research plan. It needed a visual marker, something that showed progress at a glance rather than making you infer it from the log. So I added the Steps panel on the left, with a flowchart that maps the full pipeline, from finding competitors to the final report, and highlights the current step so you always know where the model is.

The run screen mid-research, showing the current action, the loop diagram, steps used against the limit and progress per competitor.
Mid-run: what’s happening now, the loop with the web search tool in use, steps used against the limit, and each competitor’s progress.

On the right is the log. It pairs with the Steps panel. The flowchart shows where the model is, and the log explains what it’s doing there in plain language. It also lets you trace any finding in the final report back to the search that produced it.

I also wanted the log to make it visible whenever the Guardrail overrules the model. This happens when the model wants to deviate from the plan, and the Guardrail keeps it on track instead.

The log with the model's request for another pass on Sketch crossed out, and the Guardrail step highlighted in coral on the loop diagram.
The Guardrail overrules the model’s request for another pass on Sketch. Shown in preview mode with example data.

Colour

The first version with a purple button, beside the neutral palette with black buttons that replaced it.
The first version’s purple button, and the neutral palette that replaced it.

The first interface Claude produced had a purple button. Of course it did. Purple is the AI tool’s go-to colour. To own the visual direction, I switched to a neutral palette built mostly from black, white, and grey. That meant colour could be used sparingly, reserved only for states that actually need attention. I used coral for when the Guardrail overrules the model, and teal for when a run is waiting on you. Red and blue were the obvious first picks, but both already carry too many meanings in most interfaces.

Run screen with a coral highlight on a Guardrail override and a teal highlight on a prompt waiting for the user.
How the two colours show up in practice in a later version of the tool, coral for an override, teal for a prompt awaiting input.

The report

The finished competitive research report, with Figma as the baseline column.
The finished report, with Figma as the baseline column.

Once the run finishes, you get a report with an executive summary and a comparative matrix. It ends with a record of what you changed, how often the Guardrail stepped in, and how many searches ran.

The record at the end of a report, listing changes made, Guardrail interventions and searches run.
The record at the end of every report.

The report looked neat and organized at first glance, very compelling. But I didn’t quite trust it. When I verified the pricing details, I found one was incorrect. I also discovered the model had added a competitor on its own, Affinity, which I hadn’t approved (it’s owned by Canva, which I had approved). So for now, the report is a good first draft of the desk research, and I’d still check it before relying on it.

What I learnt

Building this helped me understand how harnesses actually work, even a small one like mine. While digging deeper into the topic, I found that the limits and checkpoints I’d defined could go stale as models improve. A future model might not need a cap of two web searches per step, or it might handle plan changes well enough that the Allow/Don’t allow prompt becomes unnecessary. That means I’ll need to revisit where they sit and what they check for.

On the project where I first encountered a harness, it was set up by engineers, but building one myself showed me that those decisions are design work too.


References

More notes

Six off-white vertical bands of varying width on a sand field.

A deterministic cover generator

I wanted every note here to have a cover, so I tried an image model first. It gave two different notes the same laptop-and-mug setup, then added text I never asked for. That wasn't going to work, so I wrote a spec for a generator instead and had Claude build it. The first thing it does is hash the filename, which means turning it into a number. Every letter has a number behind it, and the code combines them in a fixed way, so cover-art-generator always comes out as 283832237. That number picks the colour and the layout. The filename is the only input. Nothing random, nothing that depends on when I run it. That's what deterministic means: same input, same output, every time. I can rebuild the whole site and not one cover changes. Then I put it on a page, which meant deciding what other people get to change. As a script it needed no controls, since I was the only one running it. The obvious answer is two colour pickers, one for the background and one for the marks. I didn't do that. Every background already comes with the mark colour I picked to sit on it, and two pickers break the pair. So you never type a colour. You pick one of four palettes, the filename picks a swatch inside it, and one slider shifts the hue of the whole palette. Saturation and lightness stay fixed. It's at /lab/cover-art. Type in the name of something you've written and see what it gives you.

August 31, 2026
Four table screens rendered four different ways without a shared convention file, and the same four rendered identically with a committed CLAUDE.md.

Claude Code, and where team conventions actually live

I decided a page needed tabs. The design file didn't have them yet, so when I prompted for them the tool drew the pattern from scratch, and it drew it wrong. That took a minute to fix. The interesting part came later. Several of us were generating pages against the same app. Each of us kept hitting gaps like that one, places where the design file had no answer and the tool had to invent something. Every session invented differently. Nobody was working carelessly. We were each getting a reasonable guess from a tool with no way of knowing what the other three had already decided. So I started correcting mine. Follow the pattern the rest of the app uses, and remember it for the next instance. That worked, and my later pages held the conventions without me asking again. It just never reached anyone else. Claude Code carries knowledge between sessions two ways, and the difference between them turned out to be the whole story. Auto memory is what I'd been using. You say remember this, and the tool writes itself a note, scoped to the repository but living outside it. Which means it's yours. Your corrections, your sessions, nobody else's. A CLAUDE.md is the other option: a markdown file you write and commit, read by every session that opens that repo, including sessions belonging to other people. Same instruction, much larger blast radius. I'd hit the second use case with the first mechanism, and put a team convention somewhere only I could see. So I did a reconciliation pass. I went back through pages other designers had generated and fixed the same three things by hand. Numbers right aligned in tables. Timestamps as a date and then a time. Page names matching their label in the navigation menu. Nobody finds these decisions interesting. They're the kind of thing a team agrees once and then stops thinking about, and we were paying for them one page at a time. Here's roughly what I'd commit now: ``markdown UI conventions Navigation Underlined tabs are reserved for page-level navigation. Page names must match their label in the navigation menu. Tables Numeric columns are right aligned. Text columns are left aligned. Timestamps render as date, then time. Before generating a new screen Check an existing page for the pattern before inventing one. If no existing page has it, ask rather than choosing. `` I'd copy two things from that shape before I copied any of the rules. It says what's reserved rather than what's preferred, because a preference invites a judgment call and every judgment call is another place two sessions can diverge. And the last block is the one I'd want in there on day one. Most convention documents describe the cases you thought of. The drift came from the cases nobody had thought of yet, where the tool filled a gap because filling gaps is what it does. Telling it to ask when the pattern is missing turns a silent guess into a question. None of this is an argument against generating UI this way. The pages were good and they arrived fast. So before I correct the tool again, I ask where the correction is going to land.

August 28, 2026

Designing AI-first product interfaces

The first wave of AI products treated the model as a feature. A chat box bolted onto an existing flow. A magic button next to a regular one. Users did their normal work, and off to the side, AI was there if they wanted to summon it. Most of those never made it past the demo. The products that stuck had a different shape. The model wasn't a feature you could add or remove. It was baked into how the work got done. The flow assumed AI was in the loop. The interface let people correct or override the model, not just trigger it. On one of the manufacturing projects I worked on, the question we kept circling wasn't "where do we put the AI?" It was: if an operator has to trust this output enough to sign off on it, what does that signoff actually look like? Once we framed it that way, a lot of decisions got easier. We made the model's output visible. We let operators edit it inline. We tied every action to a clear name and timestamp. None of it was glamorous, but it was the thing that made the product real. A lot of AI design still works the old way. The product keeps doing what it always did, and the AI gets added on top. That rarely holds up. When you design with AI from the start, the questions change, and the interface gets simpler, because the model is doing the right work in the right place instead of sitting in a sidebar waiting to be useful.

June 18, 2026