Skip to content
AI

A UX-engineering harness for AI-generated UI

Five checks that run in under a minute after every UI change: screenshots at two widths and two themes, a console and network probe, a contrast sweep, a grep for generated-UI tells, and motion frames. The scripts, the loop, and a ui-harness skill.

stellae.design

12 min read

TL;DR

  • Generated UI is cheap to produce and expensive to check. A harness flips that: every revision gets the same five checks in under a minute, so the agent can iterate and you can trust the result.
  • The five checks: screenshots at two widths and two themes, a console and network probe, a contrast sweep, a grep for the tells of generated UI, and frame captures for anything that moves.
  • Scripts live outside the repo in a scratch folder, take a URL, and print something short. The agent runs them itself before reporting; the person reads the images, not the DOM.
  • The harness catches what a screenshot comparison cannot: an error only in dark mode, a label that fails contrast at 390 wide, a transition that jumps on frame three.
  • What it does not catch stays with you: whether the empty-state copy is right, whether the interaction makes sense, whether the thing should exist. The harness makes those the only questions left.

Why a harness, not a review

An agent can produce a new version of a page in thirty seconds. A person needs ten minutes to check it properly: open it, resize it, switch themes, open the console, tab through it, read the labels. At that ratio the person is the bottleneck and the checking gets skipped, which is how a page with a hydration error and an unreadable label ships looking fine in the one screenshot somebody glanced at.

The fix is not a better reviewer. It is making the check as cheap as the change. Five small scripts, each taking a URL and printing something short, that the agent runs before it says done. The person then reads five images and a few lines of output instead of driving a browser. Everything we built on this site since the page-transition work went through this loop, and it is the reason the section on reviewing agent-built UI can end with a screenshot pass instead of a manual one.

The five checks

CheckWhat it catchesOutput
Screenshots at 1440 and 390 wide, light and darkLayout breaks at a narrow width, a component that only exists in one theme, text on the wrong surfaceFour PNGs per URL
Console and network probeHydration mismatches, runtime errors, failed requests, missing message keysLines per URL, or "clean"
Contrast sweepText pairs that fail WCAG in either theme, including the ones a screenshot hidesA table, or "0 failures"
Tells grepIndigo utilities, gradient text, orbs, transition: all, emoji icons, placeholder proofFile and line per hit
Frames after an interactionA transition that jumps, an element that appears late, a layout shift mid-motionSeven PNGs at fixed offsets

Each one is small on purpose. A harness that takes five minutes to run does not get run.

Setting it up

One folder, outside the repository so it never ships and never shows up in a diff. On this site it lives in a scratch directory the agent already has, with playwright installed once.

ui-harness/ shots.mjs screenshots, widths × themes probe.mjs console, page errors, failed requests sweep.mjs contrast, from the contrast-audit skill tells.sh the grep, from the unslop-pass skill frames.mjs motion frames after a click out/ everything the scripts write
bash
mkdir ui-harness && cd ui-harness
npm init -y && npm i playwright && npx playwright install chromium

Two scripts are already published on their own pages: the contrast sweep and the tells grep. The other three are below. All of them take the base URL from BASE and default to http://localhost:3000.

Two conventions make the scripts reusable across projects. Theme is set through THEME_INIT, a snippet run before load, because every site stores its theme differently. Intro animations and cookie banners that block the first paint are skipped the same way; on this site it is one sessionStorage key.

The scripts

shots.mjs

Every URL, at every width, in every theme, optionally scrolled. Files are named so a folder listing reads as a matrix.

javascript
// BASE=http://localhost:3000 URLS=/,/pricing WIDTHS=1440,390 THEMES=light,dark SCROLL=0 node shots.mjs
import { chromium } from "playwright";
import fs from "node:fs";

const BASE = process.env.BASE ?? "http://localhost:3000";
const URLS = (process.env.URLS ?? "/").split(",");
const WIDTHS = (process.env.WIDTHS ?? "1440,390").split(",").map(Number);
const THEMES = (process.env.THEMES ?? "light,dark").split(",");
const SCROLL = Number(process.env.SCROLL ?? 0);
const THEME_INIT = process.env.THEME_INIT ?? "";
fs.mkdirSync("out", { recursive: true });

const browser = await chromium.launch();
for (const theme of THEMES) {
  for (const width of WIDTHS) {
    const ctx = await browser.newContext({
      viewport: { width, height: width < 600 ? 844 : 900 },
      colorScheme: theme,
      isMobile: width < 600,
      hasTouch: width < 600,
    });
    if (THEME_INIT) await ctx.addInitScript(`const THEME = "${theme}"; ${THEME_INIT}`);
    const page = await ctx.newPage();
    for (const path of URLS) {
      await page.goto(BASE + path, { waitUntil: "load", timeout: 60000 });
      await page.waitForTimeout(1200);
      if (SCROLL) { await page.evaluate((y) => window.scrollTo(0, y), SCROLL); await page.waitForTimeout(500); }
      const name = `${path.replace(/\//g, "_") || "_root"}--${width}--${theme}${SCROLL ? "--" + SCROLL : ""}.png`;
      await page.screenshot({ path: `out/${name}` });
      console.log("wrote", name);
    }
    await ctx.close();
  }
}
await browser.close();

probe.mjs

What the console said, what threw, what failed to load. Prints nothing but the URL when a page is clean, which is the output you want to see.

javascript
// BASE=http://localhost:3000 URLS=/,/pricing node probe.mjs
import { chromium } from "playwright";

const BASE = process.env.BASE ?? "http://localhost:3000";
const URLS = (process.env.URLS ?? "/").split(",");
const THEME_INIT = process.env.THEME_INIT ?? "";

const browser = await chromium.launch();
let total = 0;
for (const path of URLS) {
  const ctx = await browser.newContext({ viewport: { width: 1440, height: 900 } });
  if (THEME_INIT) await ctx.addInitScript(`const THEME = "light"; ${THEME_INIT}`);
  const page = await ctx.newPage();
  const lines = [];
  page.on("console", (m) => { if (m.type() === "error" || m.type() === "warning") lines.push(`${m.type()}: ${m.text().slice(0, 240)}`); });
  page.on("pageerror", (e) => lines.push(`pageerror: ${e.message.slice(0, 240)}`));
  page.on("response", (r) => { if (r.status() >= 400) lines.push(`${r.status()} ${r.url().slice(0, 160)}`); });
  await page.goto(BASE + path, { waitUntil: "networkidle", timeout: 60000 });
  await page.waitForTimeout(1000);
  console.log(`== ${path}${lines.length ? "" : "  clean"}`);
  for (const l of lines) console.log("  " + l);
  total += lines.length;
  await ctx.close();
}
await browser.close();
console.log(total ? `${total} findings` : "all clean");

frames.mjs

Click something, then capture the next second at fixed offsets. The offsets are the ones that expose a bad transition: the first frame after the click, the middle, and the settle.

javascript
// BASE=http://localhost:3000 URL=/tools CLICK='a[href*="/tools/shadow"]' SCROLL=0 node frames.mjs
import { chromium } from "playwright";
import fs from "node:fs";

const BASE = process.env.BASE ?? "http://localhost:3000";
const URL_ = process.env.URL ?? "/";
const CLICK = process.env.CLICK;
const SCROLL = Number(process.env.SCROLL ?? 0);
const OFFSETS = (process.env.OFFSETS ?? "80,200,320,460,620,800,1000").split(",").map(Number);
const THEME_INIT = process.env.THEME_INIT ?? "";
if (!CLICK) throw new Error("CLICK selector required");
fs.mkdirSync("out", { recursive: true });

const browser = await chromium.launch();
const ctx = await browser.newContext({ viewport: { width: 1440, height: 900 } });
if (THEME_INIT) await ctx.addInitScript(`const THEME = "light"; ${THEME_INIT}`);
const page = await ctx.newPage();
const errs = [];
page.on("pageerror", (e) => errs.push(e.message.slice(0, 160)));
await page.goto(BASE + URL_, { waitUntil: "networkidle", timeout: 60000 });
if (SCROLL) { await page.evaluate((y) => window.scrollTo(0, y), SCROLL); await page.waitForTimeout(400); }
const el = page.locator(CLICK).first();
await el.hover();
const t0 = Date.now();
await el.click({ noWaitAfter: true });
for (const t of OFFSETS) {
  while (Date.now() - t0 < t) await page.waitForTimeout(5);
  await page.screenshot({ path: `out/frame-${String(t).padStart(4, "0")}.png` });
}
console.log("frames written to out/, ended at", page.url(), errs.length ? "errors: " + errs.join(" | ") : "no errors");
await browser.close();

Read the frames as a strip. A good transition has no frame where the old page is gone and the new one is not yet there, and no frame where a fixed element has moved.

Running it in the loop

The order is fixed, and the agent runs it, not you.

  1. Agent makes the change.
  2. Agent runs the probe. Anything but "clean" is fixed before going further; a hydration error invalidates every screenshot after it.
  3. Agent runs the tells grep on the changed files. Hits are fixed or listed as accepted.
  4. Agent runs shots on the touched routes. It looks at the images itself, at both widths and both themes, and fixes what it sees.
  5. Agent runs the contrast sweep on the touched routes. Failures are fixed, or listed with a reason.
  6. If anything moves, agent runs frames on the interaction.
  7. Agent reports: what changed, what each check said, the image paths, and what it could not decide.

You read the images and the report. The report is short because the checks are mechanical. If step two through six pass and the images look right, the only questions left are the ones a harness cannot answer.

Two rules of thumb from running this daily. Restart the dev server before judging a stylesheet change; a stale bundle will send you chasing a bug that is not there. And read the output, not the exit code: a probe that prints one 404 is a real finding, even when everything else is green.

What the harness does not do

It does not decide whether the empty-state copy explains anything, whether the transition is the right one, whether the modal should exist, or whether the page answers the visitor's question. Those are yours. What the harness does is remove everything else from your plate, so the review that remains is the one only a person can do.

It also does not replace a real device pass before a release. Headless Chromium at 390 wide is not an iPhone. It is close enough to catch the layout that breaks, and not close enough to catch the tap target that is 2px too small under a thumb.

Skill to copy

One phase: verification after a UI change. It runs the harness and reads the results. It does not make the change and it does not fix what it finds without asking, because the fix is a separate decision.

markdown
---
name: ui-harness
description: >-
  Run the UI verification harness after a change: console probe, tells grep,
  screenshots at two widths and two themes, contrast sweep, and motion
  frames when something animates. Load before reporting any UI change as
  done. Not for making the change; run this after it.
---

# UI harness

A UI change is not done when it compiles. It is done when the five checks
have run and been read. This skill runs them in order and reports what they
said, with the image paths, so the person reads pictures and a short list.

## Setup

The harness lives in `<harness path>` with `shots.mjs`, `probe.mjs`,
`sweep.mjs`, `tells.sh` and `frames.mjs`. `BASE` defaults to
`http://localhost:3000`. `THEME_INIT` is the project's theme snippet; read
it from the rules file. If the folder or the snippet is missing, say so and
stop.

## Order

1. **Probe.** `URLS=<touched routes> node probe.mjs`. Anything but clean is
   fixed first; a hydration error invalidates every later check.
2. **Tells.** `sh tells.sh <changed files or folder>`. Each hit is fixed or
   listed under accepted with a reason.
3. **Shots.** `URLS=<touched routes> node shots.mjs`, then again with
   `SCROLL` for anything below the fold. Open every image. Look at 390 wide
   and dark mode as carefully as 1440 light; that is where the failures are.
4. **Sweep.** `URLS=<touched routes> node sweep.mjs`. Failures are fixed by
   moving lightness only, or listed with a reason.
5. **Frames.** Only if the change animates: `URL=<route> CLICK=<selector>
   node frames.mjs`. Read the strip for a gap or a jump.

Fixing something re-runs the checks from step one.

## Reading, not trusting

- Look at the images. Do not describe a screenshot you have not opened.
- A probe line is a finding even when the page renders. A 404 for a font is
  a finding.
- "0 failures" from the sweep means the pairs it checked pass. Samples that
  are meant to fail (colour demos) are accepted exceptions, not passes.
- If a check could not run (server down, missing script), report it as not
  run. Never as passed.

## Report

- What changed, in one line.
- Per check: clean, or the findings, or not run and why.
- Image paths, grouped by route.
- What you fixed on the way, and what you left for the user to decide.

## Never

- Never report a UI change done without the probe and the shots.
- Never judge a stylesheet change on a dev server that has not been
  restarted since the edit.
- Never fix a finding that changes design intent (a colour, a layout)
  without asking. Mechanical fixes (a missing key, a literal that has a
  token) are fine to make and report.
- Never run the sweep or shots across the whole site for a one-route change.

Checklist

Setting up

  • Harness folder outside the repo, playwright installed, chromium downloaded
  • THEME_INIT snippet for this project written down in the rules file
  • The five scripts run once against a known page and their output read

Per change

  • Probe clean on touched routes
  • Tells grep clean on changed files, or hits accepted with reasons
  • Shots at both widths and both themes opened and looked at
  • Sweep clean on touched routes, or failures explained
  • Frames read if anything moves
  • Report names the images and what is left to decide
AIai

Was this article helpful?