AI Agents Are Becoming the New Product Design Medium

Understanding how designers are moving beyond interfaces to shape intelligent systems through intent, iteration, and rapid experimentation.
Project Information
Team
1PM,
2 Designers
Tools
Figma, Miro, Notion Claude/GPT ....
Skills
· Product Strategy · Growth
· Visual Design · User Research
· Prototyping & Testing · Marketplace Design
Timeline
4 Weeks
Speculative exploration
Next Project
Settle rewards
Creating a seamless digital rewards platform to optimize savings during every shopping journey....
About
0→1 Product Design
Speculative Exploration · Pet Care Marketplace
AI is rapidly being adopted in the development cycle, and the process seems to be reshaping the designer's role and blurring the boundaries.
Design got faster and faster this year. Whether it got better is a harder question, and it's the one I haven't been able to put down.
A PM sent me a working prototype while I was still designing the same screen in Figma. They had built it the previous evening with an agent. Several states were missing, and error handling was apparently a problem for another evening. But the thing worked well enough to answer a question I had planned to spend two weeks investigating.
I was annoyed. Then I became interested in why I was annoyed.
My process assumed that answering the question required a sequence of artifacts: a flow, some screens, a prototype, and eventually an implementation. The PM had skipped enough of that sequence to reach an answer while I was still preparing the question.
That seems worth examining. Building software used to carry a substantial setup cost. Someone needed to understand the proposal, fit it into the existing system, implement the behavior, and make it available for review. A mockup let us investigate part of the idea before paying that cost.
This gave the design process a useful constraint. We had an incentive to reduce uncertainty before asking someone to build.
Agents change that calculation. For some questions, a small implementation becomes affordable enough to use during exploration. We can build a version, interact with it, discover a bad assumption, and revise our understanding while the design is still taking shape.
The working product becomes a medium for thinking.
There is a complication, though. A prototype can demonstrate that an interaction works while leaving the reason for building it completely unexamined. Once it exists, it also attracts a particular kind of feedback: move this button, shorten that label, add another state. Each change makes the implementation more convincing.
None necessarily makes the underlying idea more useful. This is the part I want to understand. The old cost of building encouraged us to pause, although it also made us wait. As that cost falls, we gain more chances to experiment. We need a clearer account of what we expect each experiment to tell us.

The mockup stopped being the deliverable

The designers and PMs I work with have mostly stopped handing an idea to engineering to find out whether it's any good. They test it themselves, against an agent, in an afternoon. A few run three or four agents at once in something like Conductor, wired into Paper over MCP.
For a while I assumed this was a quirk of my team. Then I started asking around, and the story comes back the same.
A PM stops sending briefs and starts sending links. Not a spec, not a Figma file, but a working thing. Rough, half the states missing, no error handling, and clickable. Built the night before with an agent. My file for the same screen is still open on the second monitor, still unfinished, and now slightly beside the point.
The first time it happened I was annoyed. The second time I was annoyed for a better reason. The rough thing had already answered a question I was planning to answer in two weeks with a proper prototype. It wasn't better than what I would have made. It was earlier, and earlier beat better.
The survey data says my team isn't unusual. In the 2026 State of AI in Design report, 76% of designers had used AI coding tools, and 85% counting app builders like Lovable and Replit. Half had shipped AI-generated code to production, which breaks down to 68% at early-stage companies and 33% at public ones. Only 20% identify as design engineers, so the capability has clearly run ahead of the title.
One number reframed the rest for me. 43% say their company now expects a working prototype as the deliverable. Not a picture of a prototype. The thing itself.

Faster, but designing what?

Here's where I get stuck, and I don't think the tension resolves cleanly.
None of this retires visual craft. Hierarchy still decides whether a person can tell their own input apart from the system's suggestion. Grouping still decides whether a long answer is reviewable or merely long. Contrast still decides what gets read first when the model is wrong. Those judgments haven't moved at all.
What moved is their object. I used to make them about a screen. Now I make them about something that behaves, and behaviour is much harder to picture than a screen is.
You can't draw a retry. You can't draw latency, a confidence threshold, or what the interface should do on the fourth failed attempt. You can only build it and watch it run.
That, I think, is the real reason the prototype won. Not speed. Fidelity to the material.
Which left me with a question I couldn't answer by reading more takes: has the ground under the tools actually moved, or is this a normal tooling cycle with better marketing? So I spent an afternoon building the smallest version I could of the thing everyone is excited about. A design tool whose file format is the web, and whose main user is a model rather than a person, plus the translator that the old world needs and the new one doesn't.
It came to 59 lines. The numbers that fell out of it were more lopsided than I expected, and a lot more boring than the argument about them.
Four things I took from the Craft chapter:
The opening is a claim plus the question that complicates it. They open with "Designers are faster with AI. But are they better?" Mine is "Design got faster this year. Whether it got better is a harder question." Same move, different sentence. It buys you the whole article in two lines.
Headings are statements, not labels. "The loop used to be slow on purpose" instead of "Background." Their headings all assert something, so a reader skimming the H2s still gets the argument.
They hold tensions open rather than resolving them. I added "Here's where I get stuck, and I don't think the tension resolves cleanly" and let the craft section stay unresolved. Your earlier drafts kept closing every point, which reads as selling. Admitting the knot is what makes a reader trust the parts where you are sure.
Paragraphs got shorter and the stat block got broken up. Two to four sentences, one idea each. The 43% now sits alone as its own beat instead of being buried at the end of a data paragraph.

What I measured

I used one screen. A billing and plan-change page: three price options, a selected state, a card for the payment method, a primary button. Nothing exotic, the kind of thing I'd normally draw in Figma, hand off, and then argue about in a comment thread.
I wrote it once as HTML and CSS, with the design tokens sitting on :root where you'd expect them. Then I wrote a script that opens the page in a headless browser, walks the DOM, and writes the same screen back out as a positioned node tree. Bounding boxes, fills as {r,g,b,a} floats, strokes, effects, a style block on every run of text. Roughly the shape a design file has. That translation step is the whole argument, so I wanted to write it myself rather than take a vendor's word for what it costs.
So I've started thinking about what a more honest review would look like. For the inbox example, I'd want a small set of messages rather than one: a straightforward question, one that asks three things at once, one missing the account details needed to answer, one from someone who's clearly frustrated, one that refers back to an earlier conversation. And I'd want to repeat a few of them, to see whether the behaviour holds between attempts.
The part I find genuinely useful is deciding what counts as acceptable before looking at the results. Does the reply address the actual question? Does it avoid promising something it can't know? Does it recognise when it needs more information? A response can be beautifully written and fail all three.
I've started thinking of this as a slower kind of critique. It doesn't replace user testing, which tells me something different. The set of examples shows how the system tends to behave. Watching someone use it shows whether the interface makes that behaviour manageable. I think I need both, and I'm still working out how much of each.

The part I didn't expect: permission

This is the section I feel least settled about, and probably the most important.
Drafting a reply and sending a reply are not the same act. One is a suggestion. The other has left the building. Somewhere in the interface, that difference has to be visible, and I don't think a confirmation dialog covers it.
If I were designing that inbox assistant, I'd probably start narrow. Let the agent prepare drafts only. Before anyone sends, they should be able to see who it's going to, read the conversation it's responding to, and edit the draft without losing those edits when they ask for another version.
A single "Enable AI" toggle would leave too much unsaid. Does it mean the agent drafts? Sends the ones you pick? Handles future messages on its own? Those are three different levels of authority wearing the same switch.
And then there's the middle of an action, which I've thought about the least and suspect matters the most. If a batch stops after three replies are sent and two are still queued, the interface has to say so plainly. Retrying can't quietly send the first three again. I notice that all of these are questions about what the system is allowed to do, not about how it looks. I'm still adjusting to that being design work.

Uncertainty needs somewhere to go

"Please check this response for accuracy" asks for a lot and offers nothing.
In the cancellation example, the agent might not have the billing date it would need to answer properly. The more useful version of uncertainty, I think, points at the specific gap. Name what's missing. Help the person find it, or hold the draft for someone with access.
A confidently wrong answer is the harder case. My instinct is to put the customer's question and the relevant policy next to the draft, so a mismatch is easier to spot. But I'd want to test that before believing it. Adding context to a busy inbox can also just make it harder to scan, and I've been wrong about this kind of thing before.
The question I try to hold onto: what does this person need to see to make this decision well? Most of the time it's less than I want to give them.

Building earlier, with some caution

The reason I'm interested in tools like Paper and Conductor is that they let me take a design question into a small working version and look at what actually happens.
For the inbox example, a first build could stay very small. Pick a sample conversation, generate a draft, edit it, simulate sending. That's enough to examine the review moment, which is the part I'd most want to get right.
I've also started to think the brief matters more than the visual reference. If I hand a coding agent a screenshot, I get a screen. If I describe which conversation details need to be visible, which actions require review, how edits should be preserved, and which components already exist, I get something I can actually evaluate.
Working this close to implementation also changes the conversation with engineering. Whether an edit is really saved, whether an action can be reversed, whether a useful answer depends on data the product doesn't have. Those answers can change the design while it's still cheap to change.
I want to be honest about the trade-offs, though.
A prototype that runs is persuasive in a way a static file isn't, and that persuasiveness isn't earned. The layout was a first attempt. The behaviour has barely been tested. I've caught myself treating something as settled purely because it worked once.
There's a similar trap in generating options. It's easy to produce five directions now. My ability to tell which one is better hasn't improved at the same rate, and more options can make the decision harder while making the afternoon feel productive.
And the evaluation idea has a weakness too, which I should say plainly since I'm the one proposing it. If I define what an acceptable answer looks like and then improve the system against it, I can get very good at hitting criteria that were incomplete. A reply might pass on accuracy, tone, and completeness while being far too long for someone who wanted a quick yes or no. The criteria are a design decision. They need questioning as much as the interface does.
That's why I'd want to keep a short record next to each meaningful change: what I noticed, what I changed, what happened when I tried it. When generating another version is easy, the reasoning is the part worth keeping.

Where I actually am

I keep coming back to a question I can't answer yet: how much should an agent be allowed to do on its own, and when?
When working with a coding agent, I would want to provide the design intent alongside the visual direction. That includes the existing components, important interaction states, accessibility expectations, and examples of behaviour the product needs to support. Those details give me a clearer basis for reviewing what gets built.
Working closer to implementation also makes collaboration with engineering more specific. We can examine whether an edit is actually saved, whether an action can be reversed, or whether a useful response depends on information the product does not have. Those discoveries can change the design while it is still taking shape.
There is a tradeoff here. A working prototype can feel more complete than it is, and its first layout can become difficult to question once everyone has seen it run. I want to keep space for visual exploration and make clear which parts are experiments and which have been reviewed for production.
A strong UI still needs deliberate choices about typography, hierarchy, interaction, and consistency. Building with agents gives me another way to develop that craft inside the product.

Making Each Build Improve the Next Decision

The part of this process I want to preserve is the reasoning. As it becomes easier to generate another version, I need a clear account of why a change is worth making.
That record can be small: what we noticed, what we changed, and what happened when we tried it. Keeping it close to the design and implementation helps the next person understand both the decision and its limits.
My next experiment would be a focused flow like the planning assistant: enough functionality to explore how someone reviews a suggestion, corrects an assumption, and decides whether to let an agent proceed. I would use those moments to assess the quality of the experience.
That is what makes AI and agents interesting to me as a design medium. They create more opportunities to build, observe, and refine within the same process. I want to use that opportunity to bring the care I put into an interface further into how the product actually works.