Skip to content
AI

Automate design QA without automating judgement

Design QA has a measurable half and a judgement half. Automate the first completely in CI, force the second as a written line in every pull request, and route the ambiguous findings to a person. Ships a design-qa skill, the PR template lines, and a CI workflow.

stellae.design

11 min read

TL;DR

  • Design QA has two halves. One half is measurable: tokens, contrast, console, layout at a width, markup semantics. The other half is a judgement: does this copy explain, is this the right pattern, should this exist. Automate the first half completely and never the second.
  • The test for which half a check belongs to: can two people disagree about the result and both be right? If not, a machine should run it. If so, a person owns it.
  • Mechanical checks go in CI and fail the build. Judgement checks go in the pull request as a forced decision, one line, that a reviewer has to write rather than tick.
  • An agent can run every mechanical check and draft the judgement questions. It does not answer them. The moment it does, the review has been outsourced to the thing being reviewed.
  • Ambiguous findings are routed, not resolved. The harness reports "this label is 3.9:1 at 12px" and a person decides whether that label is an accepted demo sample or a defect.

The split

Every design QA pass mixes two kinds of question.

Mechanical: measure itJudgement: decide it
Does this text pass 4.5:1 in both themes?Is this the right colour for this action?
Is every colour, spacing and radius a token?Is the spacing rhythm right for this content?
Does the console stay clean on load?Does the page answer the visitor's question?
Does the layout hold at 390 wide?Is this the right layout at 390 wide?
Is every icon-only button labelled?Is the label the right word?
Are there indigo utilities, orbs, transition: all?Does the screen have a point of view?
Does each state exist and is it reachable?Is the empty-state copy honest?
Did the transition drop a frame?Is this the transition the page deserves?

The left column has one right answer. The right column has a best answer that depends on the product, the moment and taste. Teams get into trouble in two ways: by asking people to do the left column by hand, which they do badly and resentfully, and by asking machines to do the right column, which they do confidently and wrong.

The test for any new check: can two competent people disagree and both be right? No means automate it. Yes means a person owns it, and the automation's only job is to make sure the person is asked.

WCAG itself draws the line here. Its success criteria are testable, and its guidance is explicit that automated tools cover part of them and human evaluation covers the rest. A tool can measure contrast. It cannot decide whether alt text describes the image well.

Automating the left column

This is the harness, run by CI instead of by hand. On pull requests that touch UI:

CheckToolFails the build when
Console and networkprobe.mjs on touched routesAny page error, hydration mismatch, or 4xx/5xx for a first-party asset
Contrastsweep.mjs on touched routesAny text pair under threshold that is not on the exception list
Literals and tellstells.sh plus the token grepAny hit not on the exception list
Layout at widthshots.mjs at 390 and 1440, compared to the previous runA pixel diff above a small threshold, on a route the PR did not intend to change
SemanticsAn accessibility linter over the rendered pageMissing labels, wrong roles, duplicate ids, skipped heading levels
StatesThe state inventory from the ticket, checked against the toggleAn applicable state that is not reachable

Two rules make this work in practice.

Exception lists are code. The contrast demo on a colour tool fails on purpose. The one-off hero palette is not a token. Those live in a file the checks read, with a reason per line, and changes to that file get reviewed like any other. An exception that lives in someone's head is a check that gets disabled.

Only touched routes. A check that runs against the whole site on every PR is slow, noisy, and gets ignored. The workflow below reads the routes from the PR description or derives them from the changed files. Full-site runs happen nightly, and their findings become tickets, not blocked merges.

Forcing the right column

You cannot automate judgement, but you can automate the demand for it. The mechanism is a line in the pull request template that a human must write, not tick.

markdown
## Design impact

<!-- Write one of these. A reviewer will read it. -->
- No visual or interaction change.
- Visual change, intended: <what changed and why, one line>
- Visual change, and the design file was updated: <link>
- New state or copy: <which, and who wrote the words>

## Judgement calls the harness could not make

<!-- Leave the questions the checks raised. Answer them here or in review. -->

The point is not the checkbox. Checkboxes get ticked. The point is that a sentence has to be written, and a sentence that says "no change" on a PR with a screenshot diff is something a reviewer can see is false.

The second section is where the agent's output goes. When the harness finds a label at 3.9:1 on a demo sample, or a state the inventory marked deferred, it does not decide. It writes the question. The reviewer answers it. That is the whole division of labour: the agent measures and asks, the person decides.

What the agent does and does not do

An agent running this loop has a clear job.

It does: run every mechanical check before reporting, fix the findings that have one right answer (a missing label, a literal that has a token, a hydration mismatch), re-run, and draft the judgement questions with the evidence attached: the screenshot, the ratio, the state.

It does not: change a colour to make contrast pass, because which colour to change is a design decision. Rewrite copy to be shorter, because what the copy says is a product decision. Remove a state because it was hard, or add a pattern because it was common. Mark its own PR as no design impact.

The last one matters most. An agent that answers the judgement questions has reviewed itself, and the reviewer, seeing the questions answered, moves on. The questions have to arrive open.

Routing ambiguous findings

Some findings are neither clearly mechanical nor clearly judgement until someone looks. A contrast failure on a demo sample. A transition: all on an SVG path where the alternative is listing six properties. A literal hex in a one-off illustration.

These are routed, with a default:

  1. The check reports the finding with its evidence.
  2. The agent checks the exception list. Listed: accepted, cited in the report. Not listed: goes in the judgement section as a question.
  3. The reviewer answers with one of three words. Fix sends it back. Accept adds it to the exception list with the reason, in the same PR. Later creates a ticket and cites it.

The one thing that does not happen is a finding disappearing because nobody decided. Every finding ends as a fix, a listed exception, or a ticket, and the report says which.

The CI shape

A GitHub Actions workflow that runs the mechanical half on UI pull requests. It expects the harness scripts in a ui-harness folder in the repo, or vendored into CI, and a preview URL from the deploy step. Adapt the paths and the preview step to your setup; the shape is the point.

yaml
name: design-qa

on:
  pull_request:
    paths:
      - "app/**"
      - "src/**"
      - "messages/**"

jobs:
  mechanical:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: 22 }
      - run: npm ci
      - run: npx playwright install --with-deps chromium

      # Routes to check: from the PR body, else the default set.
      - name: Resolve routes
        id: routes
        run: |
          ROUTES=$(echo "${{ github.event.pull_request.body }}" | grep -oE 'Routes: [^\n]+' | sed 's/Routes: //')
          echo "urls=${ROUTES:-/,/en}" >> "$GITHUB_OUTPUT"

      # Replace with your preview deploy; a local server is enough for many setups.
      - name: Start app
        run: npm run build && (npm run start &) && npx wait-on http://localhost:3000

      - name: Tells and literals
        run: sh ui-harness/tells.sh src app

      - name: Console and network
        run: BASE=http://localhost:3000 URLS=${{ steps.routes.outputs.urls }} node ui-harness/probe.mjs | tee probe.txt
      - name: Fail on probe findings
        run: "! grep -qE 'pageerror|^  (error|[45][0-9]{2})' probe.txt"

      - name: Contrast
        run: BASE=http://localhost:3000 URLS=${{ steps.routes.outputs.urls }} node ui-harness/sweep.mjs
      - name: Fail on contrast findings not in exceptions
        run: node ui-harness/check-exceptions.mjs contrast-report.json ui-harness/exceptions.json

      - name: Screenshots
        run: BASE=http://localhost:3000 URLS=${{ steps.routes.outputs.urls }} node ui-harness/shots.mjs
      - uses: actions/upload-artifact@v4
        with: { name: shots, path: ui-harness/out }

check-exceptions.mjs is ten lines: read the report, drop every finding whose page, text and class match an entry in exceptions.json, exit non-zero if anything is left. Screenshots are uploaded rather than diffed here; add a visual-diff step once you have a baseline you trust, and keep its threshold small.

Skill to copy

One phase: quality assurance on a UI change, after it is made and before it is reviewed. It runs the mechanical checks, fixes what has one answer, and hands the rest to a person as questions.

markdown
---
name: design-qa
description: >-
  Design QA on a UI change: run the mechanical checks, fix what has one
  right answer, and draft the judgement questions for a reviewer with
  evidence attached. Load before opening or updating a pull request that
  touches UI. Not for making design decisions; those go to the reviewer.
---

# Design QA

Two kinds of finding. Mechanical findings have one right answer and get
fixed. Judgement findings have a best answer that depends on taste and
product, and get asked. This skill never confuses the two.

## Sort every finding

Ask: could two competent people disagree about this and both be right?

- No: mechanical. Fix it, or list it as an accepted exception with the
  reason from the exceptions file.
- Yes: judgement. Write it as a question with the evidence, for the
  reviewer.

## Run

1. Run the harness in order: probe, tells, shots, sweep, frames if
   anything moves. Use the `ui-harness` skill if installed.
2. For each finding, sort it.
3. Fix the mechanical ones. Examples: a missing `aria-label`, a literal
   that has a token, a hydration mismatch, a heading level skipped, a
   state that is built but unreachable. Re-run the checks after fixing.
4. Check the remaining findings against the exceptions file. Listed ones
   are accepted; cite the entry.
5. Everything left becomes a judgement question. Attach the evidence: the
   screenshot path, the ratio, the element, the state.

## Never fix these

They look mechanical. They are not.

- A colour changed to make contrast pass. Which colour moves is a design
  decision. Report the ratio and the two candidates.
- Copy shortened or reworded to fit. What it says is a product decision.
- A state removed because it was hard, or a pattern added because it is
  common.
- An animation removed because frames showed a jump. Report the frame.
- An exception added to the exceptions file. Propose it; the reviewer adds
  it.

## Report

Write the two pull-request sections:

**Design impact**: one sentence, chosen from: no visual or interaction
change; visual change, intended, with what and why; new state or copy,
with who wrote the words. Never claim "no change" when the screenshot
diff shows one.

**Judgement calls the harness could not make**: one question per finding,
with evidence. Leave them unanswered. If there are none, say so.

Then the check results, one line each, and the screenshot paths.

## Never

- Never answer your own judgement questions.
- Never mark a change as no design impact on your own authority.
- Never let a finding disappear. Every finding is fixed, an accepted
  exception with a reason, or a question in the report.
- Never run the full-site checks on a one-route change.

Checklists

Setting up the split

  • Each existing review step sorted into mechanical or judgement using the disagreement test
  • Mechanical checks wired into CI on touched routes, failing the build
  • Exceptions file exists, every line has a reason, changes to it are reviewed
  • PR template carries the two written sections, not checkboxes
  • Nightly full-site run creates tickets, not blocks

Per pull request

  • Mechanical checks green, or every red line is a listed exception
  • Design impact written as a sentence that matches the screenshot diff
  • Judgement questions present with evidence, or "none" written
  • Every finding ended as fix, exception or ticket
  • The agent did not answer the judgement questions
AIai

Was this article helpful?