The taste gap benchmarks keep finding

Harsh Chhajer
3m read

Ask an AI model to judge which of two logos is better and it will answer confidently. A February 2026 benchmark tested that confidence against 100-plus independent evaluators and found the best model was right less than 27% as often as a human expert.

The number, and why it's not close

The benchmark ran 400 tasks across 1,195 images, judged through 13,000-plus individual assessments.

Visual Aesthetic Benchmark, February 2026Score
Human expert baseline68.9%
Best model (Claude Sonnet 4.6)26.5%
Gap42.4 points

That gap is less than half, not a rounding error the next model release closes.

A second dataset published in May approaches the same gap from a different direction and finds the same wall. Researchers had professional designers judge AI-generated graphic design across nine separate criteria, typography, spatial layout, tone, and found that existing model judges, general vision-language models and dedicated scoring systems alike, couldn't reach majority agreement with the designer panel on any of them.

Two different benchmarks, five months apart, converging on the same finding: models are not currently able to reliably say what a trained eye says is good.

Why this gap specifically resists automation

Most tasks AI has gotten good at have a checkable answer: does the code compile, does the test pass, does the summary contain the right facts. Aesthetic judgment doesn't have that structure. Two professional designers can disagree about a layout and both be defensible, which means there's no ground truth for a model to be trained against the way there is for correctness.

There's a second reason underneath that one. Taste is not stored knowledge, it's the residue of a habit: someone looked at their own first instinct, decided it was wrong, and did that a few thousand times. What gets published is the last version. The rejected ones, and the reasons they were rejected, never leave the person who rejected them. A model trained on finished work is training on the output of that habit while the habit itself stays invisible.

That also explains why the gap held across two separate benchmarks built by different teams, months apart, instead of narrowing the way most model capabilities do release over release. A capability trained on visible outputs improves when you feed it more outputs. A capability built on invisible, repeated internal correction has no equivalent training signal sitting around to feed it.

What this actually means for how you spend your time

If taste is the one part of this job current models measurably cannot do, it is also the one part where your own repeated practice compounds instead of getting automated out from under you. Every other piece of the workflow, layout generation, copy variants, code scaffolding, is a place where the model is closing the gap month over month. This is the piece where the gap held steady across two independent benchmarks in the same year.

Spend fifteen minutes today doing the thing a benchmark can't simulate: pick two versions of something you made, decide which is better, and write down the actual reason, not the vibe. That's the repetition the models in these benchmarks have no equivalent of yet, and it's the one skill on this list that gets more valuable, not less, the longer the gap holds.