Skip to content
BlogBuild

Making accessibility a build step

Why our contrast checks are a script that can fail, not a tick that fades, and what a list of 38 colour pairings taught us about where bugs hide.

3 min readTanveer Singh
Terminal output of npm run audit:contrast: 76 studio pairings and 2,280 generated-palette pairings checked, all pass; no token misuse found.

Most accessibility checks run once. Someone opens a contrast checker, confirms the brand colours pass, and moves on. Then the site grows: a new chip, a hover state, a dark theme. Nothing checks those.

We treat colour contrast the way we treat a type error: something a script can refuse.

A list, and a script that fails

This site's colours live as design tokens in one stylesheet. Next to them is a list of every pairing the design actually uses — body text on paper, a copper label on a card, white text on the green button while it's hovered. Each has the WCAG threshold it has to meet: 4.5:1 for text, 3:1 for borders and focus rings.

// label, foreground, background, threshold
["body text", "--ink", "--paper", 4.5],
["copper chip", "--copper-on-tint", "--copper-tint", 4.5],
["primary button hover label", "#ffffff", "--graph-green", 4.5],

npm run audit:contrast rebuilds both themes from the stylesheet, measures all 38 pairings in each, and exits with an error if any falls short. It also checks every palette the site customiser can produce — 2,280 more pairings — because a visitor can repaint this site, and their colours have to pass too.

What it caught

The first audits found three real failures:

  • A copper chip whose text measured 4.07:1 against its own tint.
  • A green chip in dark mode at 2.97:1.
  • The contact form's input border at 2.21:1, against the 3:1 that interface components need.

None of them looked like mistakes. Each was fine on the screen it was designed on.

One colour, two jobs

The copper chip and a later dark-mode bug had the same cause. A "deep" green was doing two jobs: the background of a button, where white text sits on it, and the colour of text on a light tint. In light mode those two jobs happen to want the same shade, so the bug hides. In dark mode they pull in opposite directions — as text, that green measured 2.90:1.

The fix wasn't a different shade. It was a separate token for each job — --graph-green-on-tint for text — so each can be tuned in each theme without breaking the other.

A green run only proves the list

The same audit once passed while two failures were live on the site. Those pairings were built from utility classes the list didn't know about, so they were never measured.

So the script now does a second thing: it scans the components for the known wrong combinations — a "deep" colour used as text on a tint — and fails on those too. That narrows the blind spot. It doesn't close it. A check only covers what it has been told about, so every new colour combination gets a line in the list, in the same change that introduces it.

Why bother

For a care provider or a public service, readable text isn't a style preference; it decides who can use the site. A contrast check costs a few seconds each time. A regression that reaches a user costs far more.

From the portfolioNobility Care AustraliaCompany Website · See the project →

Have a project in mind?

Tell us what you're building. We'll tell you if we're the right fit — and if not, we'll point you somewhere that is.