Why adoption outran trust in AI design tools

Harsh Chhajer
3m read

You open Claude or Cursor before you open Figma now. You use it every day. And you still read every line of what it hands back, because you have been burned by output that looked right and was not. That is not a discipline you built on purpose. It is the tool telling you something.

Everyone adopted, nobody settled

The AI in Design 2026 report, from Designer Fund and Foundation Capital, puts numbers on the first half. 91% of designers now use AI in their design work at least weekly, up from 54% a year ago. 75% use it daily. The average toolstack went from 3 tools to 7.

65% are running Claude Code, which could not appear on the 2025 survey because it had not launched yet. Anthropic shipped Claude Design into research preview on April 17, 2026.

Then the second half. Only 37% say they have settled on a clear set of go-to tools. 49% are still looking.

The complaint is the requirement

Ask what makes a tool stick and 80% say the same thing: reliable, high-quality output.

Ask what is hardest about using AI for design work and 62% say the same thing from the other side: output that is inconsistent or unreliable. It is the most-picked answer on the list.

Those are one number read from two ends. The property that earns a permanent slot is exactly the property most designers say is missing.

Why design hides its own failures

A broken build fails loudly. A bad API call throws.

A spacing value that is off by four pixels does not. It renders. It looks like a screen. It demos fine in a room where nobody is holding a ruler to it.

Most of the ways design work fails do not raise an error. They look fine until someone who knows the difference looks closely, which is usually after something has been built on top of them. That is why a year of daily use did not convert into trust: the failures are quiet, so the evidence that a tool is dependable never accumulates.

Trust is a thing you measure

Adoption happened by itself. Trust will not.

Nothing in a 7-tool stack tells you which of those tools has actually earned an unreviewed pass, because nobody is keeping score. So keep score, the way an engineer treats a flaky test.

Pick the one tool you reach for most. For the next two weeks, every time you correct its output, write one line: what it got wrong, and whether you have written that same line before.

The line is short. "Used its own grey instead of ours, third time." "Invented a second button component." "Spacing off, again." That is the whole practice, and it takes about fifteen seconds each time.

A tool that fails the same way twice has a pattern you can write a rule for. A tool that fails differently every time is one you cannot promote yet. The first kind is fixable with a file the tool reads before it starts. The second is a tool you are still auditing, whatever your habits say.

Ten minutes of logging, spread across two weeks, tells you something the survey cannot: which of your tools is actually dependable, and which one you have just gotten used to fixing.