Design critique became an installable file

Harsh Chhajer
4m read

Sit in enough design reviews and you notice the good ones have a fixed shape. What is working, said first and specifically. What is broken, ranked by how much it costs. A verdict. Then what to try instead. The order barely changes between a senior designer reviewing a login screen and the same person reviewing a checkout flow, because the order is the method, not the subject.

That shape is a file now. Not a metaphor for one: an actual folder of instructions an agent loads and runs against work nobody has looked at yet.

The two changes that made it possible

Anthropic introduced Agent Skills on October 16, 2025, as folders of instructions Claude pulls in only when a task calls for them, and in December released the format as an open standard other platforms and agents could implement. Its plugin directory followed on May 22, 2026, a reviewed catalogue sitting inside the tool itself.

Format, then distribution. It is the same sequence that turned code from something you retyped from a book into something you installed, and the interesting part was never the format. It was that a decision one person made carefully could now be run by someone who had not made it.

What survives being written down

Not all of critique does. Split a review into the part that can be checked and the part that can only be judged, and the line is unusually clean:

CheckableJudgeable
Contrast ratio against a stated thresholdWhether this screen deserves to feel calm or urgent
Whether spacing uses the scale or improvisesWhether the improvisation was the right call
Whether the pattern matches the one used elsewhereWhether consistency is worth what it costs here
Whether hierarchy has a single clear first readWhether that first read is the right thing to lead with

Everything in the left column travels. It has a threshold, so a machine can hold it and apply it identically to the tenth screen and the four-hundredth.

The right column does not travel, and the evidence for that is now specific rather than vibes. A February 2026 aesthetic benchmark ran 400 tasks over 1,195 images against more than 100 independent evaluators. The best model scored 26.5%. Human experts scored 68.9%. That is not a gap the next release quietly closes.

Why narrow is the point

The temptation with a critique skill is to ask it for a verdict, because a verdict is what you actually want. That is the version that fails, and it fails quietly, by being confidently wrong in a register that sounds like a senior designer.

A skill that only reports what it can point to is worth running unattended. A skill that hands down judgement is worth arguing with, which defeats the purpose of having automated it.

So the useful version is deliberately smaller than the thing it is named after. It finds, it ranks, it stops.

Write the rule before you write the file

Here is the part worth doing by hand, and it takes about twenty minutes. Pick the note you give most often in reviews. Not the interesting one, the boring recurring one you are tired of repeating. Then write it in four parts:

Rule: <the one thing you keep saying>

Check:    <what makes it true or false, stated as a threshold
           someone else could apply without asking you>
Severity: <high / medium / low, and what makes it high>
Evidence: <what to point at on the screen, not an adjective>
Fix:      <the specific change, not "improve the hierarchy">

Fill that in four times and you have a critique rubric. Whether it ends up as a skill file, a checklist in a doc, or a paragraph you paste into a chat is a detail. The four fields are the work.

Most rules die at the Check line, and that is the useful failure. If you cannot state the threshold without saying "it depends", you have found something in the right column of that table, and the honest answer is that it stays yours.

What this actually changes about the job

The pressure this creates is not that agents are coming for critique. It is that the checkable half of your judgement is now something you can hand over completely, and once you can, the part of the review that is still yours gets much more exposed.

If most of what you contribute in a review is the left column, that contribution is now installable, and being the person who spots contrast failures fastest stops being a role. If most of what you contribute is the right column, you spend less time on the checking and more on the argument, which is the trade you wanted anyway.

Either way, the sorting happens whether or not you write the file. Writing it is just how you find out which column you have been living in.