TL;DR
- A generated prototype shows the state the prompt described and nothing else. Every state you did not name does not exist, and the model will not volunteer it.
- Coverage is the quality metric. Not "does the happy path look right" but "how many of the states this screen can be in are designed, built and reachable".
- The state inventory is written before generation, at the top of the ticket, in a fixed order: empty, loading, partial, error, permission, long content, offline, destructive recovery. Twenty minutes, every time.
- Each state is a separate prompt or a separate slice. Asking for "all the states" in one go produces the happy path with a spinner on it.
- States have to be reachable to be reviewed. A prototype with an error state nobody can trigger has no error state.
- The inventory is also the acceptance test. Review counts states, not screens.
What a prototype leaves out
Ask an agent for a settings page and you get a settings page: a form, filled in, with a save button. It looks finished. It is a single frame of a screen that has, at a minimum, nine other frames:
- the form before anything has been saved
- the form while it loads
- the form when the API returned half the fields
- the form after a save that failed
- the form for a user who can view but not edit
- the form with a name that is 80 characters long, in German
- the form while offline, with a save queued
- the form after "delete account" was pressed by mistake
- the form when there is nothing to configure yet
None of these were in the prompt, so none of them are in the output. This is not a model failure. A human designer given the same one-line brief would also draw the happy path first. The difference is that the human knows the other nine exist and would eventually get to them. The agent does not know they are missing, and it reports done.
The result is a prototype that answers the wrong question. It shows what the screen looks like when everything works, which is the one situation where design matters least.
Coverage as the metric
The way out is to stop grading prototypes on the screen and start grading them on the states. A prototype with a plain happy path and six honest unhappy states is worth more than a beautiful happy path alone, because the unhappy states are where the product decisions live: what do we say when the save fails, what can a viewer see, what happens to a queued change.
So the question for any AI prototype becomes: of the states this screen can be in, how many are designed, built and reachable? That number is the coverage. It is countable, it can be put in a ticket, and it makes "done" mean something.
The inventory, in a fixed order
Write it before the first prompt. The order matters because each state tends to expose the next.
| State | The question it forces | Typical miss |
|---|---|---|
| Empty | What is here before the user has done anything? | A blank list with no explanation and no action |
| Loading | What holds the shape while data arrives? | A spinner that collapses the layout |
| Partial | What if only some of the data came back? | Cards with undefined in them |
| Error | What went wrong, and what can the user do now? | "Something went wrong" and a dead end |
| Permission | What does a user who cannot act see? | Disabled controls with no reason |
| Long content | What does an 80-character name or a German label do? | Overflow, truncation with no title, wrapped buttons |
| Offline | What happens to an action with no connection? | Silent failure |
| Destructive recovery | After the irreversible action, what is the way back? | No undo, no confirmation, or both |
Two states are not on the list on purpose. "Success" is the happy path; you already have it. "Validation" belongs inside error: a field-level failure is an error with a location.
Not every state applies to every screen. A read-only dashboard has no destructive recovery. Say so in the inventory, with one word: not applicable. An unmarked row is a state nobody thought about.
Generating the unhappy paths
Once the inventory exists, the generation has a shape. Three rules.
One state per prompt, or per slice. "Build the settings form with empty, loading, error and permission states" gets you the form with a spinner. Ask for the form, then ask for the empty state of that form, then the error state, each with its own acceptance line. The agent does its best work when the target is one thing. This is the same slicing rule as implementing from a large Figma frame.
Give each state its content. Empty needs the copy that explains why and the one action that fills it. Error needs the cause and the recovery. Long content needs the actual long string, in the actual language. A state without content is a placeholder with a different background colour.
Make each state reachable. A prototype where the error state is a separate file nobody can navigate to has not designed the error state; it has drawn it. Every state needs a switch: a query parameter, a mock toggle, a button in a dev panel, a story. If reviewers cannot get to it in under ten seconds, it is not reviewable and will not be reviewed.
The prompt at the end of this guide does the first two. The third is a project convention: decide once how states are toggled, write it in the rules file, and every prototype gets it for free.
Reviewing by state
Review with the inventory open. For each row: designed, built, reachable, or one of the two honest alternatives, deferred with a reason or not applicable. Anything else is a gap.
Some things to look for per state, from what agents get wrong most often:
- Empty. The copy says what is missing and what to do, not "no items". There is one action, and it is the same action the happy path uses to create the first item.
- Loading. A skeleton that holds the loaded layout's dimensions. No spinner in the middle of a collapsed container. Loading of a single field does not reflow the form.
- Partial. Missing fields render as missing, not as
undefined,nullorNaN. A card with three of five values still lines up with its neighbours. - Error. Three signals: colour, icon, text. The text says what happened and what to try. The failed action is still available. Field errors sit next to the field, not at the top.
- Permission. A viewer sees the values and no editing affordance, or sees the affordance disabled with a reason. Never a button that looks live and fails on click.
- Long content. Nothing overflows, nothing is clipped without a tooltip, buttons do not wrap onto two lines. The layout was tested with the long string, not with "John Smith".
- Offline. The action is queued or blocked, and the interface says which. A queued change is visibly pending, not visibly saved.
- Destructive recovery. Either a confirmation that names the consequence, or an undo with a time window, and not both stacked. The confirm and the cancel are not the same size and colour.
The review output is one line per state. That is the acceptance record for the prototype.
What this changes upstream
Once teams grade by coverage, two things happen. The inventory moves into the ticket template, because writing it at review time is too late. And the design file starts to carry the unhappy states, because they were the ones getting asked for. Figma's own MCP workflow shows the same effect from the other side: their example documents screens, states and transitions before generating anything, and catches a missing failure state on the way back through code to canvas. The Figma guide covers that loop.
The inventory is also the smallest useful eval. Run the same ticket through a prompt change, a model change, or a new skill, count states designed and reachable before and after, and you have a regression test for prototype quality that takes a minute to read.
Skill to copy
One phase: turning a screen brief into a state inventory and driving generation from it. It does not build the states; it makes sure they get asked for, one at a time, and counted.
---
name: state-inventory
description: >-
Turn a screen brief into a state inventory before generating UI, then
drive generation one state at a time and report coverage. Load when asked
to prototype, design or build a screen or flow. Not for reviewing an
existing implementation's polish or accessibility; use those passes after.
---
# State inventory
A prototype contains the states the prompt named. This skill names them
first, so the happy path is one of many instead of the only one.
## Before generating
1. Restate the screen in one line: who is looking at it, and what they came
to do.
2. Fill the inventory. Every row gets one of: a one-line description of the
state, "not applicable" with a reason, or "deferred" with a reason.
| State | Description / n.a. / deferred |
| --- | --- |
| Empty | |
| Loading | |
| Partial | |
| Error | |
| Permission | |
| Long content | |
| Offline | |
| Destructive recovery | |
3. Show the table to the user and wait. They strike and add rows. Do not
generate anything before the inventory is agreed.
4. Confirm how states are toggled in this project (query param, mock
toggle, story, dev panel). If there is no convention, propose a query
parameter and use it consistently.
## Generating
- Build the happy path first, as its own step.
- Then one state per step, in inventory order. Each step names the state,
supplies its real content (the copy, the long string, the error cause),
and makes the state reachable through the agreed toggle.
- Never combine states in one request. "Add the empty and error states" is
two steps.
- Long content uses a real long value in the project's longest supported
language, not a placeholder.
## Per-state rules
- Empty: explain what is missing and offer the one action that fills it.
- Loading: a skeleton with the loaded layout's dimensions. No layout
collapse.
- Partial: missing values render as absent, never as undefined or NaN.
- Error: colour, icon and text; text says what happened and what to try;
the action stays available; field errors sit beside the field.
- Permission: values visible, affordances hidden or disabled with a reason.
- Long content: nothing overflows or wraps a control onto two lines.
- Offline: the action is queued or blocked, and the UI says which.
- Destructive recovery: confirmation naming the consequence, or undo with a
window. Not both.
## Report
The inventory again, with a status per row: designed, built, reachable,
deferred, or not applicable. Then the coverage line: "N of M applicable
states built and reachable". Then how each state is toggled.
## Never
- Never report a screen done with the happy path alone.
- Never mark a state built if the reviewer cannot reach it in one action.
- Never invent a state's copy. If the error cause is unknown, ask.
- Never fill the inventory silently. The user agrees it first.
Prompt to copy
For a one-off, without the skill.
Before you build anything, list the states <screen> can be in, one row each:
empty, loading, partial data, error, permission-limited, long content,
offline, destructive recovery. For each, write one line describing the
state in this product, or "not applicable" with a reason. Wait for me to
confirm the list.
Then build the happy path. Then build each confirmed state as a separate
step, in that order. Every state gets real content: the actual copy, a real
80-character name, the real error cause. Every state must be reachable via
<toggle convention>.
End with the list again, marking each state built and reachable, and the
line "N of M states covered".
Checklist
Before the first prompt
- Inventory written, every row filled or marked not applicable
- Toggle convention decided and in the rules file
- Real long content and real error causes available to hand over
Before calling it done
- Each applicable state is built and reachable in one action
- Each state was reviewed against its per-state rules
- Deferred states have a reason and an owner
- Coverage line recorded on the ticket
Related
- Figma MCP for design work, where the same inventory runs before implementation
- Prototype, pick, promote, for the cases where the state's design itself is undecided
- Reviewing agent-built UI, the passes to run once the states exist
- Loader generator and placeholder generator, for the loading and empty states
Recommended Tools
Was this article helpful?