How the 2026 AI design workflow actually works
You open a coding agent before you open Figma now, and you hand it a rough PRD instead of a wireframe. What comes back this year is not a simulation you click through and describe to an engineer later. It is a running build, with your data shape, sitting at a real URL, and it happened in one sitting.
Vibe design is not vibe coding
Andrej Karpathy's February 2025 post named the coding version of this: describing software in plain language and mostly not reading the code it produces. Google Labs relaunched Stitch on March 18, 2026 with the design-side sibling built in, calling it vibe design, alongside a voice-driven canvas and a portable format for a team's visual identity.
The distinction that survived past the marketing: vibe coding stays inside the code editor, describing implementation. Vibe design starts from product intent and still has to end in real, running code, not a screenshot of one. Everything below is what happens between those two ends once the intent has been typed or spoken.
The four stages, run in order
Practitioners converged on roughly the same sequence this year, regardless of which specific tool they used.
- Interrogate before you generate. Hand the agent a PRD, a few bullet points, whatever exists, and tell it explicitly to ask you about scope, users, and edge cases before producing anything. Under-specifying a prompt does not produce a blank result. It produces the model's best guess at your gaps, silently, and you find out which guesses were wrong after the build exists.
- Load a point of view before the first screen. Left alone, generative tools default to a narrow, recognisable look, purple gradients, identical card grids, glassmorphism. A packaged set of typography, spacing, and hierarchy rules, given to the agent before generation starts, changes the starting quality more than any amount of post-hoc prompting does.
- Build the MVP, not the finished screen. Core screens, mock data, basic interactions, achievable in one short session. The goal is something the team can react to, not something ready to ship.
- Iterate with critique language, attach references, then deploy for real. "This feels too dense" gets a better result than a rewritten spec, because the agent can revise fast enough to absorb many small corrections. When the direction is right but the look is wrong, attaching reference screenshots and asking the agent to rework against them does more than either the point-of-view file or the critique alone. The last move is publishing a real, shareable build rather than a prototype link, because feedback on something that behaves like the product is categorically different from feedback on something that only looks like it.
The failure mode worth naming is skipping straight to stage three because it's the satisfying one. A model that isn't told to interrogate you first will not sit quietly waiting for a better brief. It fills every gap you left with its own default, an assumed audience, an assumed edge case, an assumed tone, and hands back something that looks finished. You only discover which assumptions were wrong once the build already exists and someone has to argue you out of it, which costs more than the fifteen minutes of questions would have.
What changed the underlying economics this year
None of the four stages above were available to run this fast twelve months earlier, and three releases inside four months of each other in 2026 are why. Anthropic launched Claude Design into research preview on April 17, scanning a team's actual codebase and design files so generated work matches real brand tokens rather than defaulting to generic output. Anthropic followed on June 18 with Claude Code Artifacts, turning a coding session's output into a live page a team can watch update, rather than the file people were already emailing each other by hand. OpenAI shipped GPT-5.6, including the Sol variant, on July 9, tuned specifically for long, difficult, agentic coding work rather than short completions.
The pattern practitioners describe is routing by strength rather than picking one vendor: a creative-judgment model for ideation and design critique, a different model for the specific implementation problem that stays stuck after several attempts. None of this was a single company's roadmap. It reads more like three labs independently arriving at the same four-stage shape from different directions, which is itself evidence the shape is real rather than a marketing artifact.
Why the deploy step carries more weight than it sounds
The old pipeline's actual cost was never the time spent in Figma. It was the gap between what a stakeholder clicked through and what shipped, a gap that only ever showed up after the handoff, when it was expensive to fix. A build sitting at a real URL collapses that gap before it can open. Someone types garbage into a form, resizes a browser to a width nobody designed for, or clicks the thing that was never wired up, and all three responses are more concrete than "this feels a bit cluttered" said about a static frame.
That is also why stage four's reference-image step matters more than it looks. A model with the right typography rules can still hand back a well-spaced, badly organised screen. Layout, hierarchy, and whether a screen should exist at all remain a human call. The four stages compress the distance between deciding something and seeing it working. They do not remove the deciding.
Figma's own response this year is a useful confirmation of where the pressure actually sits. Rather than ceding early exploration to chat-based tools entirely, it folded generation directly onto its canvas with Dev Mode and an MCP-based pipeline into external coding agents, which only makes sense as a strategy if the canvas is still where refinement happens once a direction is validated, even as the earliest divergent thinking moves elsewhere first.
Run one real feature through all four stages this week
Pick something small enough to finish in a single sitting, not your next big feature. Spend 15 minutes answering the agent's own questions instead of skipping straight to a prompt. Spend 10 minutes writing three sections of a point-of-view file, colours, type, spacing, in real values, and load it before you generate anything. Build the MVP in one session, then run two rounds of critique-language feedback instead of a rewritten prompt. Deploy it to a real URL before you show anyone.
The version of this that fails is skipping straight to the deploy step because it's the exciting one. The interrogation stage is the one doing the actual work, and it is also the one every source that tried this a year ago admits they were tempted to skip first.