TL;DR
- Generated UI is cheap to produce and expensive to check. A harness flips that: every revision gets the same five checks in under a minute, so the agent can iterate and you can trust the result.
- The five checks: screenshots at two widths and two themes, a console and network probe, a contrast sweep, a grep for the tells of generated UI, and frame captures for anything that moves.
- Scripts live outside the repo in a scratch folder, take a URL, and print something short. The agent runs them itself before reporting; the person reads the images, not the DOM.
- The harness catches what a screenshot comparison cannot: an error only in dark mode, a label that fails contrast at 390 wide, a transition that jumps on frame three.
- What it does not catch stays with you: whether the empty-state copy is right, whether the interaction makes sense, whether the thing should exist. The harness makes those the only questions left.
Why a harness, not a review
An agent can produce a new version of a page in thirty seconds. A person needs ten minutes to check it properly: open it, resize it, switch themes, open the console, tab through it, read the labels. At that ratio the person is the bottleneck and the checking gets skipped, which is how a page with a hydration error and an unreadable label ships looking fine in the one screenshot somebody glanced at.
The fix is not a better reviewer. It is making the check as cheap as the change. Five small scripts, each taking a URL and printing something short, that the agent runs before it says done. The person then reads five images and a few lines of output instead of driving a browser. Everything we built on this site since the page-transition work went through this loop, and it is the reason the section on reviewing agent-built UI can end with a screenshot pass instead of a manual one.
The five checks
| Check | What it catches | Output |
|---|---|---|
| Screenshots at 1440 and 390 wide, light and dark | Layout breaks at a narrow width, a component that only exists in one theme, text on the wrong surface | Four PNGs per URL |
| Console and network probe | Hydration mismatches, runtime errors, failed requests, missing message keys | Lines per URL, or "clean" |
| Contrast sweep | Text pairs that fail WCAG in either theme, including the ones a screenshot hides | A table, or "0 failures" |
| Tells grep | Indigo utilities, gradient text, orbs, transition: all, emoji icons, placeholder proof | File and line per hit |
| Frames after an interaction | A transition that jumps, an element that appears late, a layout shift mid-motion | Seven PNGs at fixed offsets |
Each one is small on purpose. A harness that takes five minutes to run does not get run.
Setting it up
One folder, outside the repository so it never ships and never shows up in a diff. On this site it lives in a scratch directory the agent already has, with playwright installed once.
ui-harness/
shots.mjs screenshots, widths × themes
probe.mjs console, page errors, failed requests
sweep.mjs contrast, from the contrast-audit skill
tells.sh the grep, from the unslop-pass skill
frames.mjs motion frames after a click
out/ everything the scripts write
mkdir ui-harness && cd ui-harness
npm init -y && npm i playwright && npx playwright install chromium
Two scripts are already published on their own pages: the contrast sweep and the tells grep. The other three are below. All of them take the base URL from BASE and default to http://localhost:3000.
Two conventions make the scripts reusable across projects. Theme is set through THEME_INIT, a snippet run before load, because every site stores its theme differently. Intro animations and cookie banners that block the first paint are skipped the same way; on this site it is one sessionStorage key.
The scripts
shots.mjs
Every URL, at every width, in every theme, optionally scrolled. Files are named so a folder listing reads as a matrix.
// BASE=http://localhost:3000 URLS=/,/pricing WIDTHS=1440,390 THEMES=light,dark SCROLL=0 node shots.mjs
import { chromium } from "playwright";
import fs from "node:fs";
const BASE = process.env.BASE ?? "http://localhost:3000";
const URLS = (process.env.URLS ?? "/").split(",");
const WIDTHS = (process.env.WIDTHS ?? "1440,390").split(",").map(Number);
const THEMES = (process.env.THEMES ?? "light,dark").split(",");
const SCROLL = Number(process.env.SCROLL ?? 0);
const THEME_INIT = process.env.THEME_INIT ?? "";
fs.mkdirSync("out", { recursive: true });
const browser = await chromium.launch();
for (const theme of THEMES) {
for (const width of WIDTHS) {
const ctx = await browser.newContext({
viewport: { width, height: width < 600 ? 844 : 900 },
colorScheme: theme,
isMobile: width < 600,
hasTouch: width < 600,
});
if (THEME_INIT) await ctx.addInitScript(`const THEME = "${theme}"; ${THEME_INIT}`);
const page = await ctx.newPage();
for (const path of URLS) {
await page.goto(BASE + path, { waitUntil: "load", timeout: 60000 });
await page.waitForTimeout(1200);
if (SCROLL) { await page.evaluate((y) => window.scrollTo(0, y), SCROLL); await page.waitForTimeout(500); }
const name = `${path.replace(/\//g, "_") || "_root"}--${width}--${theme}${SCROLL ? "--" + SCROLL : ""}.png`;
await page.screenshot({ path: `out/${name}` });
console.log("wrote", name);
}
await ctx.close();
}
}
await browser.close();
probe.mjs
What the console said, what threw, what failed to load. Prints nothing but the URL when a page is clean, which is the output you want to see.
// BASE=http://localhost:3000 URLS=/,/pricing node probe.mjs
import { chromium } from "playwright";
const BASE = process.env.BASE ?? "http://localhost:3000";
const URLS = (process.env.URLS ?? "/").split(",");
const THEME_INIT = process.env.THEME_INIT ?? "";
const browser = await chromium.launch();
let total = 0;
for (const path of URLS) {
const ctx = await browser.newContext({ viewport: { width: 1440, height: 900 } });
if (THEME_INIT) await ctx.addInitScript(`const THEME = "light"; ${THEME_INIT}`);
const page = await ctx.newPage();
const lines = [];
page.on("console", (m) => { if (m.type() === "error" || m.type() === "warning") lines.push(`${m.type()}: ${m.text().slice(0, 240)}`); });
page.on("pageerror", (e) => lines.push(`pageerror: ${e.message.slice(0, 240)}`));
page.on("response", (r) => { if (r.status() >= 400) lines.push(`${r.status()} ${r.url().slice(0, 160)}`); });
await page.goto(BASE + path, { waitUntil: "networkidle", timeout: 60000 });
await page.waitForTimeout(1000);
console.log(`== ${path}${lines.length ? "" : " clean"}`);
for (const l of lines) console.log(" " + l);
total += lines.length;
await ctx.close();
}
await browser.close();
console.log(total ? `${total} findings` : "all clean");
frames.mjs
Click something, then capture the next second at fixed offsets. The offsets are the ones that expose a bad transition: the first frame after the click, the middle, and the settle.
// BASE=http://localhost:3000 URL=/tools CLICK='a[href*="/tools/shadow"]' SCROLL=0 node frames.mjs
import { chromium } from "playwright";
import fs from "node:fs";
const BASE = process.env.BASE ?? "http://localhost:3000";
const URL_ = process.env.URL ?? "/";
const CLICK = process.env.CLICK;
const SCROLL = Number(process.env.SCROLL ?? 0);
const OFFSETS = (process.env.OFFSETS ?? "80,200,320,460,620,800,1000").split(",").map(Number);
const THEME_INIT = process.env.THEME_INIT ?? "";
if (!CLICK) throw new Error("CLICK selector required");
fs.mkdirSync("out", { recursive: true });
const browser = await chromium.launch();
const ctx = await browser.newContext({ viewport: { width: 1440, height: 900 } });
if (THEME_INIT) await ctx.addInitScript(`const THEME = "light"; ${THEME_INIT}`);
const page = await ctx.newPage();
const errs = [];
page.on("pageerror", (e) => errs.push(e.message.slice(0, 160)));
await page.goto(BASE + URL_, { waitUntil: "networkidle", timeout: 60000 });
if (SCROLL) { await page.evaluate((y) => window.scrollTo(0, y), SCROLL); await page.waitForTimeout(400); }
const el = page.locator(CLICK).first();
await el.hover();
const t0 = Date.now();
await el.click({ noWaitAfter: true });
for (const t of OFFSETS) {
while (Date.now() - t0 < t) await page.waitForTimeout(5);
await page.screenshot({ path: `out/frame-${String(t).padStart(4, "0")}.png` });
}
console.log("frames written to out/, ended at", page.url(), errs.length ? "errors: " + errs.join(" | ") : "no errors");
await browser.close();
Read the frames as a strip. A good transition has no frame where the old page is gone and the new one is not yet there, and no frame where a fixed element has moved.
Running it in the loop
The order is fixed, and the agent runs it, not you.
- Agent makes the change.
- Agent runs the probe. Anything but "clean" is fixed before going further; a hydration error invalidates every screenshot after it.
- Agent runs the tells grep on the changed files. Hits are fixed or listed as accepted.
- Agent runs shots on the touched routes. It looks at the images itself, at both widths and both themes, and fixes what it sees.
- Agent runs the contrast sweep on the touched routes. Failures are fixed, or listed with a reason.
- If anything moves, agent runs frames on the interaction.
- Agent reports: what changed, what each check said, the image paths, and what it could not decide.
You read the images and the report. The report is short because the checks are mechanical. If step two through six pass and the images look right, the only questions left are the ones a harness cannot answer.
Two rules of thumb from running this daily. Restart the dev server before judging a stylesheet change; a stale bundle will send you chasing a bug that is not there. And read the output, not the exit code: a probe that prints one 404 is a real finding, even when everything else is green.
What the harness does not do
It does not decide whether the empty-state copy explains anything, whether the transition is the right one, whether the modal should exist, or whether the page answers the visitor's question. Those are yours. What the harness does is remove everything else from your plate, so the review that remains is the one only a person can do.
It also does not replace a real device pass before a release. Headless Chromium at 390 wide is not an iPhone. It is close enough to catch the layout that breaks, and not close enough to catch the tap target that is 2px too small under a thumb.
Skill to copy
One phase: verification after a UI change. It runs the harness and reads the results. It does not make the change and it does not fix what it finds without asking, because the fix is a separate decision.
---
name: ui-harness
description: >-
Run the UI verification harness after a change: console probe, tells grep,
screenshots at two widths and two themes, contrast sweep, and motion
frames when something animates. Load before reporting any UI change as
done. Not for making the change; run this after it.
---
# UI harness
A UI change is not done when it compiles. It is done when the five checks
have run and been read. This skill runs them in order and reports what they
said, with the image paths, so the person reads pictures and a short list.
## Setup
The harness lives in `<harness path>` with `shots.mjs`, `probe.mjs`,
`sweep.mjs`, `tells.sh` and `frames.mjs`. `BASE` defaults to
`http://localhost:3000`. `THEME_INIT` is the project's theme snippet; read
it from the rules file. If the folder or the snippet is missing, say so and
stop.
## Order
1. **Probe.** `URLS=<touched routes> node probe.mjs`. Anything but clean is
fixed first; a hydration error invalidates every later check.
2. **Tells.** `sh tells.sh <changed files or folder>`. Each hit is fixed or
listed under accepted with a reason.
3. **Shots.** `URLS=<touched routes> node shots.mjs`, then again with
`SCROLL` for anything below the fold. Open every image. Look at 390 wide
and dark mode as carefully as 1440 light; that is where the failures are.
4. **Sweep.** `URLS=<touched routes> node sweep.mjs`. Failures are fixed by
moving lightness only, or listed with a reason.
5. **Frames.** Only if the change animates: `URL=<route> CLICK=<selector>
node frames.mjs`. Read the strip for a gap or a jump.
Fixing something re-runs the checks from step one.
## Reading, not trusting
- Look at the images. Do not describe a screenshot you have not opened.
- A probe line is a finding even when the page renders. A 404 for a font is
a finding.
- "0 failures" from the sweep means the pairs it checked pass. Samples that
are meant to fail (colour demos) are accepted exceptions, not passes.
- If a check could not run (server down, missing script), report it as not
run. Never as passed.
## Report
- What changed, in one line.
- Per check: clean, or the findings, or not run and why.
- Image paths, grouped by route.
- What you fixed on the way, and what you left for the user to decide.
## Never
- Never report a UI change done without the probe and the shots.
- Never judge a stylesheet change on a dev server that has not been
restarted since the edit.
- Never fix a finding that changes design intent (a colour, a layout)
without asking. Mechanical fixes (a missing key, a literal that has a
token) are fine to make and report.
- Never run the sweep or shots across the whole site for a one-route change.
Checklist
Setting up
- Harness folder outside the repo,
playwrightinstalled, chromium downloaded -
THEME_INITsnippet for this project written down in the rules file - The five scripts run once against a known page and their output read
Per change
- Probe clean on touched routes
- Tells grep clean on changed files, or hits accepted with reasons
- Shots at both widths and both themes opened and looked at
- Sweep clean on touched routes, or failures explained
- Frames read if anything moves
- Report names the images and what is left to decide
Related
- Contrast audit, the sweep script and its skill
- Unslop pass, the tells grep and its skill
- Reviewing agent-built UI, the human passes that run on top
- Responsive tester, for the device pass the harness does not replace
Was this article helpful?