Why generative UI still needs your rules
13.0% to 71.7%. That is how often a small open model produced a working interface before and after one team tuned it for the job, on that team's own benchmark.
The model runs on a single consumer graphics card. Interfaces that assemble themselves for each user stopped being a lab demo this autumn.
Two things are true at once
Fact one. Generating a valid screen is close to solved, and cheap. OUI-1, released by Thesys on September 8, 2026, writes interfaces as a short component language instead of raw markup. Its error rate fell from 35.3 to 3.8 per 100 statements.
Fact two. Nobody has shown that a generated screen follows a particular product's design. The best evidence on that question is a study from September 18, 2026 with 12 practitioners in it.
Why each one holds
The first holds because validity is checkable. A screen either parses or it does not. Components either connect or they do not. That is a target a model can be trained against, and Thesys trained against it.
Clean runs rose from 24 of 184 to 132 of 184. On 60 requests the model had never seen, 55 came back valid.
The second holds because conformance is not checkable the same way. There is no parser for "this looks like our product". The September study built one by hand. Its system, GUIDE, lets a designer inspect a generated screen and correct it, then learns from the corrections and scores later screens against them.
The 12 practitioners preferred screens shaped that way over screens generated from example files and a design document the model wrote for itself. A written description of your design was the weaker input. A designer's edits were the stronger one.
Resist picking a winner
One reading says the hard part is done and design systems are next to be automated. The other says a benchmark score from the vendor proves nothing.
Neither survives the details. The validity numbers are real, and they come with a stated limit: state, queries and anything that changes data are not covered yet. The conformance study is real too, and it is 12 people and an abstract with no effect sizes.
So you have a machine that reliably produces a screen, and no machine that reliably produces your screen. Those are different problems. The first got solved by training. The second, so far, gets solved by a person correcting output.
Write the rules a screen must pass
If screens are going to be assembled at run time, the design system has to exist as checks and not only as components. You can start that today in about twenty-five minutes.
- Pick five screens your product already ships.
- For each, write the rules it obeys as yes or no questions. Is there one primary action. Does every list have an empty state. Is destructive action separated from the default one.
- Merge the lists and keep the rules that hold across all five.
- Hand one generated screen to a colleague with the list and ask them to mark it.
Where they hesitate, the rule is not written clearly enough for a machine either. That list is the thing a generator will need from you, whichever generator it turns out to be.
What nobody has measured yet
The open question is the one between the two facts. How many corrections does a designer have to make before generated screens stay inside a system without supervision.
The study shows the direction and not the amount. Nobody has published the curve. Until someone does, a team shipping generated interfaces is running that experiment on its own users, and should know it.