Making accessibility a build step
Why our contrast checks are a script that can fail, not a tick that fades, and what a list of 38 colour pairings taught us about where bugs hide.
3 min readTanveer Singh
Most accessibility checks run once. Someone opens a contrast checker, confirms the brand colours pass, and moves on. Then the site grows: a new chip, a hover state, a dark theme. Nothing checks those.
We treat colour contrast the way we treat a type error: something a script can refuse.
A list, and a script that fails
This site's colours live as design tokens in one stylesheet. Next to them is a list of every pairing the design actually uses — body text on paper, a copper label on a card, white text on the green button while it's hovered. Each has the WCAG threshold it has to meet: 4.5:1 for text, 3:1 for borders and focus rings.
// label, foreground, background, threshold
["body text", "--ink", "--paper", 4.5],
["copper chip", "--copper-on-tint", "--copper-tint", 4.5],
["primary button hover label", "#ffffff", "--graph-green", 4.5],npm run audit:contrast rebuilds both themes from the stylesheet, measures all 38 pairings in each, and exits with an error if any falls short. It also checks every palette the site customiser can produce — 2,280 more pairings — because a visitor can repaint this site, and their colours have to pass too.
What it caught
The first audits found three real failures:
- A copper chip whose text measured 4.07:1 against its own tint.
- A green chip in dark mode at 2.97:1.
- The contact form's input border at 2.21:1, against the 3:1 that interface components need.
None of them looked like mistakes. Each was fine on the screen it was designed on.
One colour, two jobs
The copper chip and a later dark-mode bug had the same cause. A "deep" green was doing two jobs: the background of a button, where white text sits on it, and the colour of text on a light tint. In light mode those two jobs happen to want the same shade, so the bug hides. In dark mode they pull in opposite directions — as text, that green measured 2.90:1.
The fix wasn't a different shade. It was a separate token for each job — --graph-green-on-tint for text — so each can be tuned in each theme without breaking the other.
A green run only proves the list
The same audit once passed while two failures were live on the site. Those pairings were built from utility classes the list didn't know about, so they were never measured.
So the script now does a second thing: it scans the components for the known wrong combinations — a "deep" colour used as text on a tint — and fails on those too. That narrows the blind spot. It doesn't close it. A check only covers what it has been told about, so every new colour combination gets a line in the list, in the same change that introduces it.
Why bother
For a care provider or a public service, readable text isn't a style preference; it decides who can use the site. A contrast check costs a few seconds each time. A regression that reaches a user costs far more.
From the portfolioNobility Care AustraliaCompany Website · See the project →