Takeaways
- A skill exists to make the agent walk the same decision path every time. If the path is not written down, the agent invents one per session.
- Every rule carries its reason. A bare rule gets applied blindly; a rule with a why gets extended to cases you never wrote.
- Numbers and absolute words change behaviour. "Reasonable", "tasteful" and "where appropriate" do not.
- One skill covers one aspect, and building is a different skill from reviewing.
- Test a line by running the skill with it and without it. If the output does not change, the line is decoration.
Why the first draft fails
Ask a coding agent for a dialog animation three times and you get three answers. All of them plausible, none of them yours. The agent is sampling from everything it has seen, and everything it has seen includes a lot of average work.
The first instinct is to write down your preferences: "keep motion subtle, prefer restraint, respect the user". The agent agrees completely and changes nothing. Those sentences describe an attitude, not a decision. The agent already has the attitude. What it lacks is the decision path you take without noticing.
A good skill is that path, written out.
Narrow the answer space
Instead of describing the outcome you like, describe how you choose. Then the agent chooses the same way, and the outcome varies only where the context varies.
Here is the tree we use for depth on a surface. Without it, an agent reaches for a shadow, a border, a ring and a gradient on the same card, because each one is common.
Does the surface sit on top of other content?
├── No → no depth cue. Spacing separates it.
└── Yes
├── Light mode → 1px border, rgba(0,0,0,0.08)
│ └── Floating (popover, modal)? → add a layered shadow, three stops
└── Dark mode → 1px solid border (#2a2a2a family), no shadow
└── Floating? → lighten the surface one step instead
The card still looks different on a dashboard and on a marketing page. The reasoning is the same. That is the whole point of a skill: predictable reasoning, not identical output.
Anatomy, briefly
A skill is a folder with a SKILL.md. The frontmatter has two required keys, name and description. The name matches the folder. The description is capped at about a thousand characters and has to say both what the skill does and when to use it, because it is the only part the agent reads before deciding whether to load the rest.
Anything long lives in sibling files: a curve library, a checklist, a picker spec. The body links to them and says when to open them. This is progressive disclosure, and it is what makes a skill cheaper than a pasted prompt. The description is always in context. The body arrives when the task matches. The reference files arrive when a line in the body sends the agent there.
The description is the router
Write it last, and write it for a machine choosing between forty skills.
| Weak | Strong |
|---|---|
| UI animation guidance | Interface motion: whether to animate at all, easing choice, duration, springs, enter and exit pairs. Load when adding or reviewing transitions, dialogs, drawers, toasts, hover and press feedback. |
| Typography best practices | Type as the subject: font loading, variable axes, scales, line length, wrapping, truncation. Not for a heading inside a layout review. |
The strong version lists the nouns a task would contain. It also says what the skill is not for, so a neighbouring skill wins the tasks that belong to it.
Write the why
A rule without its reason is a rule the agent applies in the wrong place.
| Rule only | Rule with reason |
|---|---|
| Use tabular numbers in tables. | Any number that changes gets font-variant-numeric: tabular-nums. Proportional digits are different widths, so a counter jiggles and a column of prices does not line up. Static numbers in prose stay proportional. |
| Light-mode borders are alpha. | Light-mode 1px borders are rgba(0,0,0,0.08), never a solid hex. Alpha takes on the surface behind it and recedes; a solid #e0e0e0 sits on top like a pasted line. In dark mode the opposite holds, because alpha white glows. |
The second column does two things the first cannot. It tells the agent when the rule does not apply, and it lets the agent extend the logic to a case you did not list. The price column and the timer were never mentioned, but they follow.
Think of the reader as a very fast junior with excellent recall and no history. They will do exactly what you say. Say why, and they will also do what you meant.
Be exact
Soft words are invisible to an agent. Go through the draft and replace every one of them with a number, a value, or an absolute.
| Invisible | Visible |
|---|---|
| Keep animations reasonably short. | UI animations stay under 300ms. Exits run 20 to 30 percent faster than enters. |
| Use a subtle press effect. | Press scales to 0.96. Never below 0.95, which reads as exaggerated. |
| Prefer readable line lengths. | Body text is capped at 65ch. |
Try to avoid transition: all. | Never transition: all. Name the properties. |
"Never" and "always" anchor a behaviour. "Try to" and "consider" do not, and the agent reads them as optional. If a rule matters, the sentence has to say so.
Strictness feels like it removes creativity. It does not. The creative decision was yours, made once, and the skill stops the agent from re-guessing it. Whatever you have not decided stays open.
Say when to break it
A skill that is only rules produces an agent that applies them where they do not belong. Every strong rule needs its exception written next to it.
- Baseline grids: an editorial tool. Overkill in dense product UI.
- Press feedback: skip it on high-frequency controls, where speed beats feedback.
- Balanced headline wrapping: on a three-line heading it can leave the first line stranded. Sometimes a manual break wins.
- Contrast ratios: 3:1 fails WCAG body text and can pass APCA for large text. The specs disagree. Judge in context.
Trick questions belong in a skill on purpose. They teach the agent that the rule is a default, not a law.
Cut what the agent already knows
The agent knows what a CSS transform is. It knows what a modal is. Explaining either does not just waste space, it dilutes the lines that matter, because attention spreads across everything in the file.
Read the draft one sentence at a time and ask: does this line change what the agent does? Definitions, background, encouragement, and restating the framework all fail the test. What survives is rules, reasons, exceptions and the decision trees.
A good skill is defined as much by what it leaves out as by what it puts in. So is a good interface.
One aspect, and build is not review
"Design" is not a skill. Animation is. Typography is. Surfaces are. The agent picks by description, and a description that covers everything matches everything, which is the always-loaded rules file with extra steps.
Split further along the phase of work. Building an animation and reviewing one are different tasks with different outputs: one produces code, the other a Before / After / Why table. Two small skills, animate and review-animations, beat one that tries to do both and does neither well.
Test it by running it
The loop that built your taste in the first place still works: make something, notice what feels off, say why, refine. Now the thing being refined is the document.
- Pick a real task the skill should cover. Not a toy.
- Run the task with the skill loaded and without it. Compare the output side by side.
- Find the difference. If there is none, the skill is a values statement. Rewrite the vaguest line as a rule with a number and a reason.
- Find the mistake the agent still makes. That is the next rule. Write it as a pair, wrong beside right.
- When unsure whether one line earns its place, delete it and run again.
Keep the tasks you test with. Three fixtures re-run after every edit catch the regression where a new rule quietly cancels an old one.
Maintain it
- Date the file. A skill with no date is a skill nobody trusts.
- Change one rule per edit and re-run the fixtures.
- When the agent argues with a rule and turns out to be right, update the rule, not the agent.
- Keep a short list of rejected rules at the bottom. The rejected ones are the ones that come back.
Checklists
Before the first run
- Name matches the folder, lowercase and hyphens
- Description says what the skill does and when to load it, in nouns a task would contain
- Description says what the skill is not for
- One aspect of the interface, one phase of work
- Reference material lives in sibling files, linked from the body
Per line
- Changes what the agent does, or delete it
- Has a number, a value, or an absolute where one is possible
- Carries its reason
- Names its exception if it has one
- Does not explain something the agent already knows
Before sharing it
- Run with and without on three real tasks, output differs in the intended way
- The agent's remaining mistakes are written as new rules or accepted on purpose
- Dated
- Rejected rules noted at the bottom
A shortcut
The fastest way to a first draft is not to write it as a skill. Dump everything you know about one topic into a plain note, unordered, with your examples and your opinions. Then hand that note to an agent with a skill-writing skill loaded and ask for a SKILL.md. It will sort your dump into rules, reasons and exceptions, cut the background, and tighten the words.
The output is still your taste. The agent only did the formatting. Run the checklists on it, then the fixtures, and you have a skill that took an afternoon instead of a month.
---
name: <one-aspect-one-phase>
description: <What it does>: <the nouns a task would contain>. Load when <situations>. Not for <neighbouring aspect>.
---
# <Aspect>
<One paragraph on what this skill decides and the stance it takes. No definitions.>
## Decide
<A decision tree for the choice the agent gets wrong most often.>
## Rules
- <Rule with a number or absolute.> <The reason.> <The exception, if any.>
- <Rule.> <Reason.> <Exception.>
## Wrong and right
| Wrong | Right |
| --- | --- |
| <the common default> | <your choice, with the value> |
## Checklist
- [ ] <one line per rule the agent can tick>
## Rejected
- <rule you tried and dropped, and why>
Updated: <YYYY-MM-DD>