TL;DR
- "Match the design" gives an agent a target it can only approximate. "Use the token vocabulary" gives it a set of allowed moves. The second is smaller, and that is the point.
- A token is a guardrail only if it has a role, not a shade.
accentandaccent-inkconstrain.purple-500invites apurple-400next to it. - The agent has to be able to read the tokens live. A skill that restates values is a copy that drifts. Point at the theme file, or expose it as a resource.
- Pairs are the missing half. A background token without its ink token is how an agent puts muted grey text on a pastel band and calls it done.
- Enforce with a grep and a contrast sweep, not with review. A hex literal in a component is a failing check, not a conversation.
- Exceptions exist and are written down: demo samples, one-off brand moments, third-party embeds. Everything else uses a token or explains why not.
The two prompts
Same component, same agent, same file. The only difference is the sentence.
Prompt A. "Build the newsletter card so it matches the design."
The agent looks at the screenshot or the Figma context, reads a near-black background, a lavender-ish accent, 16px padding, and writes them down: bg-[#0b0a0f], text-[#c8b6ff], p-4, a border in rgba(255,255,255,0.08). Every value is within a few percent of the design. The card looks right in the screenshot comparison. It is also a dead end: the next theme change does not reach it, the dark-mode variant will be hand-picked, and the fourth card built this way has a fourth lavender.
Prompt B. "Build the newsletter card. Colour, spacing, radius and type come from the tokens in app/globals.css. If a value has no token, stop and say which one."
The agent reads the theme block, finds surface, accent, accent-ink, the spacing scale and the radius scale, and writes bg-surface text-accent-ink p-4 rounded-2xl. It then reports that the design's border has no token and asks whether border at 8% alpha is the one. That question is the guardrail working. You answer it once, in the token file, and it never comes up again.
The output of B is not prettier. It is smaller, and it is the same as the last card and the next one.
Why constraints beat targets
An agent generating UI is choosing from everything it has seen. Ask it to match a picture and it samples from that whole space, aiming at the picture. It will land close and slightly different every time, because there are ten thousand near-misses for every exact hit.
A token vocabulary collapses the space. There are eleven colours, eight spacing steps, four radii, one type scale. The agent is not aiming any more; it is picking. Picking is a task models are good at, and the result is reproducible because the options are enumerable.
This is the same reason a skill works better as a decision tree than as a description. Narrow the answer space and the reasoning becomes yours. Tokens do it for values the way a skill does it for process.
What a vocabulary needs before an agent can use it
Most token sets were built for humans who already know the design. An agent needs three things those sets often lack.
Roles, not shades. text, text-muted, surface, surface-hover, border, accent, accent-ink. A role tells the agent when to use the token. A shade like neutral-600 tells it nothing, and it will reach for neutral-500 when it wants something slightly lighter, which is exactly the drift the tokens were supposed to prevent. If a raw scale has to exist, keep it out of the utilities the agent can see.
Pairs. Every background token that text can sit on needs a named ink. On this site the pastel bands take #0a0a0a ink, always, because at lightness 0.85 the pastel and the muted grey sit too close and labels vanish. We learned this by watching an agent put text-accent-ink on a bg-accent band in dark mode, where accent-ink resolves to the accent itself, and the headline disappeared. The fix was not a review comment. It was writing the pair down.
Modes as remaps, not repaints. Dark mode is the same roles fed from the other end of the scale. If the agent can see that surface is one token with two values, it will not invent a dark hex. If dark mode is a second set of tokens with different names, it will mix them.
The rest is hygiene the agent will thank you for: a spacing scale on a 4px grid with no 52px step in it, radii where inner equals outer minus padding, a type scale with line heights attached. Anything the agent has to compute is a value it will compute differently next time.
Where the agent reads them from
The token file has to be the source the agent actually reads, at the moment it needs a value. Four ways to arrange that, from cheapest to most robust:
- A rules file that points at the theme. One line in
CLAUDE.mdor.cursor/rules: tokens live inapp/globals.cssunder@theme, and no colour, spacing, radius or type value is written as a literal. This costs nothing and works because the agent can open the file. - A skill that says how to use them. The
tokens-firstskill below. It does not list values. It lists the procedure: read the theme, map every value, ask when nothing matches. - A
DESIGN.mdfor portability. Tokens as front matter, roles and pairs as prose. Right when the agent works outside the repo, or when several repos share an identity. Covered in DESIGN.md: a design system your agent can read. - A resource on an MCP server. The tokens exposed as a readable resource, or the Figma variables through
get_variable_defs. Right when the tokens live somewhere the agent cannot open as a file. See MCP for designers for what a resource is.
The one arrangement that does not work is restating the values inside the skill or the prompt. The copy is right for a week.
Enforce mechanically
A guardrail that depends on someone noticing is a suggestion. Two checks catch nearly everything, and both run in seconds.
Grep for literals. Colour hexes, rgb(, oklch( outside the theme file, arbitrary Tailwind values like p-[13px] or text-[#. The unslop pass on this site ships a version of this as tells.sh; the token-specific subset is:
grep -rnE --include='*.tsx' --include='*.css' \
-e '#[0-9a-fA-F]{3,8}\b' \
-e '\[(#|rgb|oklch|hsl)' \
-e '\b(p|m|gap|rounded|text)-\[[0-9.]+(px|rem)\]' \
src app | grep -v globals.css | grep -v node_modules
Every hit is a question. Some are exceptions, most are drift.
Sweep for contrast. Roles and pairs can still combine badly, especially in the mode the agent did not test. The contrast audit skill walks every rendered text node in both themes and reports the pairs that fail. On this site it found about thirty failing labels after a token change, all of them places where a muted token had been put on a surface it was never paired with.
Run both in CI on pull requests that touch the theme or any component. A red check is cheaper than a review comment and does not get tired.
Exceptions, written down
Not everything is a token, and pretending otherwise produces workarounds. Keep a short list in the rules file:
- Demo samples. A contrast checker has to show a failing pair. A text-shadow tool shows "Aa" in whatever the user picked. These are the product, not the chrome.
- One-off brand moments. A hero illustration with its own palette. Named, scoped to one component, and not reused.
- Third-party embeds. Their colours are theirs.
- Ink on pastel bands.
#0a0a0a, deliberately not a token here, because it is the one value that must never be remapped by theme.
An exception that is written down is a decision. One that is not is the first of many.
Skill to copy
One phase: implementing or editing UI with an existing token set. It does not choose tokens for a project that has none; that is a design decision, and it belongs to you.
---
name: tokens-first
description: >-
Implement or edit UI using only the project's design tokens for colour,
spacing, radius and type. Load when building or changing components in a
codebase that has a token file. Not for creating a token set from scratch,
and not for reviewing contrast; use contrast-audit for that.
---
# Tokens first
A value written as a literal is a decision made twice: once by the
designer who set the token, and again, differently, by whoever typed the
hex. This skill makes sure the second time does not happen.
## Before writing
1. Find the token source. In order: a Tailwind `@theme` block, CSS custom
properties, a tokens JSON, a `DESIGN.md` front matter. Read it in full.
If there is more than one, ask which is canonical.
2. List the roles it defines: text, muted text, surfaces, borders, accent
and its ink, category or brand colours, the spacing scale, the radius
scale, the type scale. Note any pairs the file names explicitly.
3. Read the rules file for exceptions. If none are listed, assume there are
none.
## While writing
- Every colour, spacing, radius and type value maps to a token. No hex,
`rgb()`, `oklch()`, or bracketed arbitrary values outside the theme file.
- Text on a coloured surface uses that surface's paired ink token. If no
pair is named, use the project's darkest text token and flag it.
- Dark mode is not a second set of values. If the token has two modes, use
the token and let the mode switch. Never write a `.dark` override with a
literal.
- Spacing comes from the scale. A gap the scale does not have is either the
nearest step or a question; it is never a bracketed pixel value.
- A radius inside a padded container is the outer radius minus the padding,
taken from the scale.
## When nothing matches
Stop and report, in one line per value: what the design shows, the nearest
token, and the difference. Do not invent a token and do not write the
literal. The user decides whether to add a token or accept the nearest.
## After writing
1. Grep the changed files for literals. Every hit is either on the
exception list or a defect.
2. Render in both modes. A pair that reads in light and vanishes in dark is
the most common failure.
3. Report: tokens used, values that had no token, exceptions applied.
## Never
- Never copy a token's value into the skill, a comment, or a prompt. Point
at the file.
- Never add a token to the theme without being asked. Propose it in the
report.
- Never "match the design" with a literal because the token is slightly
off. Report the difference instead.
Prompt to copy
For the case where the skill is not installed, or the task is a one-off. It is the shape of Prompt B above.
Build <component> in this codebase.
Colour, spacing, radius and type come from the token file at <path>. Read
it before writing anything. No hex, rgb, oklch or bracketed arbitrary values
outside that file.
Text on a coloured surface uses that surface's paired ink token. Dark mode
uses the same tokens, not overrides.
If the design needs a value the tokens do not have, stop and list each one
as: what the design shows, the nearest token, the difference. Do not invent
a token and do not write the literal.
End with the tokens you used and any values you could not map.
Checklists
Is the vocabulary ready for an agent?
- Every token is a role, not a shade
- Every coloured surface has a named ink
- Dark mode is the same tokens with different values, not a second set
- The spacing scale is on a grid with no odd steps
- The rules file names where the tokens live and lists the exceptions
Per change
- Literal grep on the diff is clean, or every hit is a listed exception
- Rendered in both modes
- Contrast sweep passes on the touched pages
- Values that had no token are in the report, not in the code
Related
- DESIGN.md: a design system your agent can read, the portable form of the same vocabulary
- Contrast audit, the sweep that catches bad pairs
- Unslop pass, which greps for the literals among other tells
- OKLCH converter and brand colours, for building the roles in the first place
Recommended Tools
Was this article helpful?