What a design system teaches you about AI skills

Harsh Chhajer
5m read

You wrote a skill for your design principles, then one for component specs, then one for tone of voice, and somewhere around the fourth or fifth one the agent started loading the wrong one, or three at once, and answering a design-review question with your content style guide. Nobody removed anything. It just stopped being three separate files you could hold in your head.

Skill sprawl is a design-system problem

That failure has a name in every design team that has ever shipped without a system: components built in isolation, styled slightly differently each time, until nobody can say which button is the real one. Skill sprawl is the same failure, one layer up the stack. Anthropic's Agent Skills, introduced on October 16, 2025 and released as an open standard that December, made it trivial to package a piece of knowledge into a folder an agent loads on demand. It did not, by itself, stop people from writing overlapping, contradictory folders the same way a component library doesn't stop someone from building a fourth grey.

The tangle is predictable enough to describe before you hit it. Skills written one at a time, each sensible on its own, start to disagree at the edges: the review skill assumes something the principles skill contradicts, and the agent has no way to know which one you meant. Nothing broke. The set stopped being coherent, which is a different problem and a harder one to notice.

Specs answer one question, principles answer another

The distinction underneath the whole tangle is worth separating out on its own, because conflating it is the most common way an agent gives confidently wrong design feedback.

Ask a specAsk a principle
Which component should this be?Is this the right component for this context?
Are the tokens applied correctly?Does this respect what the user actually needs here?
Anything with one definitive right answerAnything requiring judgement: is the hierarchy working, does this hold up on mobile

A spec executes. A principle evaluates. A skill file that mixes both ends up doing neither reliably, because the agent cannot tell which mode it's supposed to be reasoning in.

Play that out concretely. Ask an agent holding a mixed file whether a settings screen should use a modal or a full page, and it might answer from the spec half, whichever pattern your component library happens to list first, when the actual question is a principle question: does this decision need to interrupt the user's flow, or can it wait. Split the file and the agent knows which register to answer in before it starts, the same way a designer knows the difference between "which component exists" and "which component is right here" without having to think about it.

The three layers that hold once you split them

Once separated, the working shape has three parts. Reference skills hold knowledge only, design principles, component and token specs, content voice, and never act on their own. Capability skills perform a workflow, an audit, a hi-fi frame, a prototype, and load whichever reference skills the task actually needs. Tool connectors, built on the Model Context Protocol Anthropic introduced in November 2024, give a capability skill somewhere real to act, a design tool's API, a project tracker, a codebase.

The payoff is the same one a real design system gives a human team: update one principle, and every workflow that depends on it inherits the change automatically. Nothing has to be found and edited in four places. That is the entire reason the layering is worth the setup time, not a nice-to-have on top of it.

The same problem, at organisation scale

On June 23, 2026, Anthropic launched Claude Tag, replacing a per-user Slack integration with one shared, organisation-identity Claude per channel, carrying its own admin-provisioned access rather than any single person's personal connections. That is the skill-sprawl problem inverted: instead of one person's folder growing too tangled to trust, it is an entire team's standing context, spread across however many channels, that has to stay coherent for everyone tagging it at once.

A shared agent that gets design principles and content tone confused in front of one designer is an annoyance. The same confusion in a channel the whole team tags is a standing source of wrong answers nobody notices is wrong, because it looks confident every time. The layering discipline that fixes an individual setup is the same discipline an org-wide deployment cannot skip, just with a much more expensive failure mode if it does.

Nobody at the individual level asked to become an information architect for their own agent. That is a fair complaint, and it does not change the fact that the work is now unavoidable, the same way nobody chose to become a build engineer either, back when a team's first component library outgrew whoever was maintaining it from memory. The skill folder is the design system now. It gets the same discipline or it gets the same failure.

Audit your own folder in 20 minutes

Open whatever skills, rule files, or standing prompts you currently point an agent at and sort each one into exactly two piles.

  1. Does it answer with a definite right answer, or does it require judgement? A spec belongs in one file. A principle belongs in a separate one. If a single file does both, that is the file causing your confident-wrong answers, split it. Ten minutes.
  2. Does anything only make sense loaded alongside something else? If your component-review skill silently assumes your principles file is already in context, write that dependency down explicitly rather than hoping the agent infers it. Five minutes.
  3. Delete the file nobody has opened in a month. A skill nobody maintains is the one most likely to contradict whatever you wrote last week. Five minutes.
  4. Write down what depends on what, in one line per file. "Component review loads principles and specs" is the whole entry. It costs nothing to write and it's the exact information that goes missing first when a set of skills starts contradicting itself.

None of this requires the multiplayer version Anthropic shipped in June. It requires the same twenty minutes a component audit takes, applied one level up, before the folder you're not looking at becomes the one giving your team's shared agent its most confident wrong answer.