TL;DR
- Design QA has two halves. One half is measurable: tokens, contrast, console, layout at a width, markup semantics. The other half is a judgement: does this copy explain, is this the right pattern, should this exist. Automate the first half completely and never the second.
- The test for which half a check belongs to: can two people disagree about the result and both be right? If not, a machine should run it. If so, a person owns it.
- Mechanical checks go in CI and fail the build. Judgement checks go in the pull request as a forced decision, one line, that a reviewer has to write rather than tick.
- An agent can run every mechanical check and draft the judgement questions. It does not answer them. The moment it does, the review has been outsourced to the thing being reviewed.
- Ambiguous findings are routed, not resolved. The harness reports "this label is 3.9:1 at 12px" and a person decides whether that label is an accepted demo sample or a defect.
The split
Every design QA pass mixes two kinds of question.
| Mechanical: measure it | Judgement: decide it |
|---|---|
| Does this text pass 4.5:1 in both themes? | Is this the right colour for this action? |
| Is every colour, spacing and radius a token? | Is the spacing rhythm right for this content? |
| Does the console stay clean on load? | Does the page answer the visitor's question? |
| Does the layout hold at 390 wide? | Is this the right layout at 390 wide? |
| Is every icon-only button labelled? | Is the label the right word? |
Are there indigo utilities, orbs, transition: all? | Does the screen have a point of view? |
| Does each state exist and is it reachable? | Is the empty-state copy honest? |
| Did the transition drop a frame? | Is this the transition the page deserves? |
The left column has one right answer. The right column has a best answer that depends on the product, the moment and taste. Teams get into trouble in two ways: by asking people to do the left column by hand, which they do badly and resentfully, and by asking machines to do the right column, which they do confidently and wrong.
The test for any new check: can two competent people disagree and both be right? No means automate it. Yes means a person owns it, and the automation's only job is to make sure the person is asked.
WCAG itself draws the line here. Its success criteria are testable, and its guidance is explicit that automated tools cover part of them and human evaluation covers the rest. A tool can measure contrast. It cannot decide whether alt text describes the image well.
Automating the left column
This is the harness, run by CI instead of by hand. On pull requests that touch UI:
| Check | Tool | Fails the build when |
|---|---|---|
| Console and network | probe.mjs on touched routes | Any page error, hydration mismatch, or 4xx/5xx for a first-party asset |
| Contrast | sweep.mjs on touched routes | Any text pair under threshold that is not on the exception list |
| Literals and tells | tells.sh plus the token grep | Any hit not on the exception list |
| Layout at width | shots.mjs at 390 and 1440, compared to the previous run | A pixel diff above a small threshold, on a route the PR did not intend to change |
| Semantics | An accessibility linter over the rendered page | Missing labels, wrong roles, duplicate ids, skipped heading levels |
| States | The state inventory from the ticket, checked against the toggle | An applicable state that is not reachable |
Two rules make this work in practice.
Exception lists are code. The contrast demo on a colour tool fails on purpose. The one-off hero palette is not a token. Those live in a file the checks read, with a reason per line, and changes to that file get reviewed like any other. An exception that lives in someone's head is a check that gets disabled.
Only touched routes. A check that runs against the whole site on every PR is slow, noisy, and gets ignored. The workflow below reads the routes from the PR description or derives them from the changed files. Full-site runs happen nightly, and their findings become tickets, not blocked merges.
Forcing the right column
You cannot automate judgement, but you can automate the demand for it. The mechanism is a line in the pull request template that a human must write, not tick.
## Design impact
<!-- Write one of these. A reviewer will read it. -->
- No visual or interaction change.
- Visual change, intended: <what changed and why, one line>
- Visual change, and the design file was updated: <link>
- New state or copy: <which, and who wrote the words>
## Judgement calls the harness could not make
<!-- Leave the questions the checks raised. Answer them here or in review. -->
The point is not the checkbox. Checkboxes get ticked. The point is that a sentence has to be written, and a sentence that says "no change" on a PR with a screenshot diff is something a reviewer can see is false.
The second section is where the agent's output goes. When the harness finds a label at 3.9:1 on a demo sample, or a state the inventory marked deferred, it does not decide. It writes the question. The reviewer answers it. That is the whole division of labour: the agent measures and asks, the person decides.
What the agent does and does not do
An agent running this loop has a clear job.
It does: run every mechanical check before reporting, fix the findings that have one right answer (a missing label, a literal that has a token, a hydration mismatch), re-run, and draft the judgement questions with the evidence attached: the screenshot, the ratio, the state.
It does not: change a colour to make contrast pass, because which colour to change is a design decision. Rewrite copy to be shorter, because what the copy says is a product decision. Remove a state because it was hard, or add a pattern because it was common. Mark its own PR as no design impact.
The last one matters most. An agent that answers the judgement questions has reviewed itself, and the reviewer, seeing the questions answered, moves on. The questions have to arrive open.
Routing ambiguous findings
Some findings are neither clearly mechanical nor clearly judgement until someone looks. A contrast failure on a demo sample. A transition: all on an SVG path where the alternative is listing six properties. A literal hex in a one-off illustration.
These are routed, with a default:
- The check reports the finding with its evidence.
- The agent checks the exception list. Listed: accepted, cited in the report. Not listed: goes in the judgement section as a question.
- The reviewer answers with one of three words. Fix sends it back. Accept adds it to the exception list with the reason, in the same PR. Later creates a ticket and cites it.
The one thing that does not happen is a finding disappearing because nobody decided. Every finding ends as a fix, a listed exception, or a ticket, and the report says which.
The CI shape
A GitHub Actions workflow that runs the mechanical half on UI pull requests. It expects the harness scripts in a ui-harness folder in the repo, or vendored into CI, and a preview URL from the deploy step. Adapt the paths and the preview step to your setup; the shape is the point.
name: design-qa
on:
pull_request:
paths:
- "app/**"
- "src/**"
- "messages/**"
jobs:
mechanical:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: 22 }
- run: npm ci
- run: npx playwright install --with-deps chromium
# Routes to check: from the PR body, else the default set.
- name: Resolve routes
id: routes
run: |
ROUTES=$(echo "${{ github.event.pull_request.body }}" | grep -oE 'Routes: [^\n]+' | sed 's/Routes: //')
echo "urls=${ROUTES:-/,/en}" >> "$GITHUB_OUTPUT"
# Replace with your preview deploy; a local server is enough for many setups.
- name: Start app
run: npm run build && (npm run start &) && npx wait-on http://localhost:3000
- name: Tells and literals
run: sh ui-harness/tells.sh src app
- name: Console and network
run: BASE=http://localhost:3000 URLS=${{ steps.routes.outputs.urls }} node ui-harness/probe.mjs | tee probe.txt
- name: Fail on probe findings
run: "! grep -qE 'pageerror|^ (error|[45][0-9]{2})' probe.txt"
- name: Contrast
run: BASE=http://localhost:3000 URLS=${{ steps.routes.outputs.urls }} node ui-harness/sweep.mjs
- name: Fail on contrast findings not in exceptions
run: node ui-harness/check-exceptions.mjs contrast-report.json ui-harness/exceptions.json
- name: Screenshots
run: BASE=http://localhost:3000 URLS=${{ steps.routes.outputs.urls }} node ui-harness/shots.mjs
- uses: actions/upload-artifact@v4
with: { name: shots, path: ui-harness/out }
check-exceptions.mjs is ten lines: read the report, drop every finding whose page, text and class match an entry in exceptions.json, exit non-zero if anything is left. Screenshots are uploaded rather than diffed here; add a visual-diff step once you have a baseline you trust, and keep its threshold small.
Skill to copy
One phase: quality assurance on a UI change, after it is made and before it is reviewed. It runs the mechanical checks, fixes what has one answer, and hands the rest to a person as questions.
---
name: design-qa
description: >-
Design QA on a UI change: run the mechanical checks, fix what has one
right answer, and draft the judgement questions for a reviewer with
evidence attached. Load before opening or updating a pull request that
touches UI. Not for making design decisions; those go to the reviewer.
---
# Design QA
Two kinds of finding. Mechanical findings have one right answer and get
fixed. Judgement findings have a best answer that depends on taste and
product, and get asked. This skill never confuses the two.
## Sort every finding
Ask: could two competent people disagree about this and both be right?
- No: mechanical. Fix it, or list it as an accepted exception with the
reason from the exceptions file.
- Yes: judgement. Write it as a question with the evidence, for the
reviewer.
## Run
1. Run the harness in order: probe, tells, shots, sweep, frames if
anything moves. Use the `ui-harness` skill if installed.
2. For each finding, sort it.
3. Fix the mechanical ones. Examples: a missing `aria-label`, a literal
that has a token, a hydration mismatch, a heading level skipped, a
state that is built but unreachable. Re-run the checks after fixing.
4. Check the remaining findings against the exceptions file. Listed ones
are accepted; cite the entry.
5. Everything left becomes a judgement question. Attach the evidence: the
screenshot path, the ratio, the element, the state.
## Never fix these
They look mechanical. They are not.
- A colour changed to make contrast pass. Which colour moves is a design
decision. Report the ratio and the two candidates.
- Copy shortened or reworded to fit. What it says is a product decision.
- A state removed because it was hard, or a pattern added because it is
common.
- An animation removed because frames showed a jump. Report the frame.
- An exception added to the exceptions file. Propose it; the reviewer adds
it.
## Report
Write the two pull-request sections:
**Design impact**: one sentence, chosen from: no visual or interaction
change; visual change, intended, with what and why; new state or copy,
with who wrote the words. Never claim "no change" when the screenshot
diff shows one.
**Judgement calls the harness could not make**: one question per finding,
with evidence. Leave them unanswered. If there are none, say so.
Then the check results, one line each, and the screenshot paths.
## Never
- Never answer your own judgement questions.
- Never mark a change as no design impact on your own authority.
- Never let a finding disappear. Every finding is fixed, an accepted
exception with a reason, or a question in the report.
- Never run the full-site checks on a one-route change.
Checklists
Setting up the split
- Each existing review step sorted into mechanical or judgement using the disagreement test
- Mechanical checks wired into CI on touched routes, failing the build
- Exceptions file exists, every line has a reason, changes to it are reviewed
- PR template carries the two written sections, not checkboxes
- Nightly full-site run creates tickets, not blocks
Per pull request
- Mechanical checks green, or every red line is a listed exception
- Design impact written as a sentence that matches the screenshot diff
- Judgement questions present with evidence, or "none" written
- Every finding ended as fix, exception or ticket
- The agent did not answer the judgement questions
Related
- A UX-engineering harness for AI-generated UI, the scripts this guide puts in CI
- AI prototyping beyond the happy path, where the state inventory comes from
- Contrast audit and Unslop pass, two of the mechanical checks
- Reviewing agent-built UI, the human passes on the judgement side
Recommended Tools
Was this article helpful?