Why your design system cannot prove it works

Harsh Chhajer
4m read

Someone asked what the design system saved last quarter. You said it improves consistency and speeds teams up. They nodded, and the headcount request went nowhere.

The number the system does not have

zeroheight asked 147 design system practitioners what they measure, for its fifth annual report this year. 41% measure adoption. 9% measure NPS. 5% measure return on investment.

Adoption is the comfortable one, and it answers the wrong question. It tells you how many teams use the thing. It says nothing about what happened because they used it.

A system that measures only adoption can describe its reach and not its effect. That is a fine position when nobody is asking. It is a bad one in a budget meeting.

What that costs, stated plainly

The same survey shows the cost. Satisfaction with the ability to get leadership buy-in fell from 42% to 32%. Dissatisfaction rose from 23% in 2025 to 40%. 61% of teams say they do not have enough people.

Read those in order and the chain is not subtle. The system cannot state its effect in the units leadership uses, so it loses the argument, so it stays understaffed, so it has even less time to go and measure anything.

That loop tightens on its own. Nothing external has to go wrong.

Meanwhile everyone else is buying

Demand is not the problem. In the State of Prototyping survey of 1,478 designers, 40.2% named design systems and tokens as somewhere they plan to invest over the next 12 months. It ranked third, behind AI-assisted coding at 64.0% and agent workflows at 46.3%.

So practitioners intend to put time into systems, and system teams cannot get funded. Those facts are not in conflict. Intent to spend your own time is free. Headcount is not, and headcount is decided by people who were never going to be moved by an adoption percentage.

Pick one number with a before

The fix is not a measurement framework. It is one number, chosen because you can state what it was before the system touched it.

Good candidates, in rough order of how easy they are to get honestly:

  • Time to first usable screen on a new project, measured twice: once on a team using the system, once on a team that is not.
  • Count of distinct values for one token type in production. Greys, spacing steps, font sizes. You can often get this in an afternoon with a script, and it moves in a direction anyone can read.
  • Review comments about consistency per pull request, before and after adoption on one team.

Each has a denominator and a date. None needs anyone to agree on what quality means.

The half hour worth spending first

Before instrumenting anything, do the cheapest version. About thirty minutes, filled in like this:

FieldFilled in
What I countedDistinct hex values in the checkout flow
BeforeForty-one, in March, counted with a grep over the repo
AfterNine, in September, same grep
Who felt itCheckout team, during the payments redesign
What it does not proveThat anything shipped faster, or that users noticed

That last row carries more weight than it looks. A claim that states its own limit survives a sceptical question from a finance lead. A claim that overreaches gets dismissed entirely, and takes the next three you make down with it.

If you cannot fill in "Before", that is the finding, and it is worth knowing this week rather than next quarter. You are not short of evidence that the system works. Nobody wrote down what things were like when it did not.

Measurement is the argument, not the paperwork

Design system teams tend to treat measurement as reporting, something done after the real work to satisfy someone else. It is the other way round. The measurement is how the work gets to continue.

One number, with a date and a before, beats a dashboard nobody requested. Somebody on your team still remembers what things were like beforehand. Ask them this week, while that is still true.