How a project actually runs.
Two tracks. Every deliverable written down before we start so there's nothing to argue about at the end.
Build track
Websites, content systems, learning platforms
AI Studio
Local and hosted models, wired into your product
Discovery
Understand before we build.
We start by asking awkward questions. What does success actually look like? Who's the real audience? What's been tried before and why did it fail? For NDIS sector clients we map accessibility here rather than at the end, because it changes how a thing is built and not only how it looks. None of this locks you in — projects change halfway through, that is normal, and we work with it. What discovery buys you is knowing what a change costs while there is still time to decide.
- Scope document with agreed success criteria
- Feature list with priorities (must/should/could)
- Timeline and milestones agreed in writing
- Cost estimate — fixed, not a range
- Accessibility requirements doc (NDIS clients)
Scope
Define the real problem, not the hype version.
Most AI projects fail at this stage — not because the model is wrong, but because the problem wasn't defined precisely enough to know whether it's working. We spend as long as needed here. What data goes in, what comes out, what failure looks like, and whether the thing is worth building at all. Part of that is deciding where the model should run: on our hardware, where your material never leaves the building, or on a hosted API where the best model for the job happens to live. We use both and we will tell you which one this needs.
- AI problem definition doc (input/output spec)
- Success criteria and explicit failure modes
- A shortlist of models — local and hosted — and what each one trades away
- What it actually costs to run, whichever way it goes
- Go/no-go decision before any build starts
Design
Decisions made visible.
We design in the browser rather than in a mockup tool — static prototypes of the real pages, at desktop and phone width, with your actual copy in them. Structure first: what goes on the page and in what order, before anything gets a colour. Tokens, type scale and spacing are written down before a single page is built, and nothing moves to build without your sign-off.
- Page structure and content hierarchy agreed before any layout
- Static mockups of the real pages, at desktop and phone width
- Design system written down: tokens, typography, spacing
- A written specification the build works from
- Your sign-off before build starts
Prototype
Validate before you commit to architecture.
A working proof of concept with real data before any production architecture is committed. We test model choice, stress-test the prompt surface, and build an evaluation dataset — a set of concrete examples where we know what correct output looks like. Without this, you have no way to know if the build improved or broke anything.
- Working PoC connected to real or representative data
- A set of examples where we already know what correct output looks like
- The model chosen, with the reason written down
- Prompt architecture draft with version history
- Speed measured on the hardware it will actually run on
Build
Short cycles, and a link you can check any time.
Development runs on Next.js, TypeScript and Tailwind — no mystery stack. Work comes back to you in short cycles: usually two or three days, a week at the outside, depending on how much is in front of us. You get a live link early on and it updates as work lands, so you can look whenever you want rather than waiting for a scheduled demo. You can change your mind within reason without the project unravelling.
- A live preview link, updated as work lands
- Nothing lands without a passing production build and contrast check
- Firestore security rules and auth wired from the start
- Accessibility handled during the build, not bolted on at QA
- Server-rendered pages, so crawlers see real content
Build & Evaluate
Assume the model will be wrong sometimes.
Integration into your product with an output validation pipeline and explicit handling for when the model gets it wrong. We work through the failures deliberately — adversarial inputs, edge cases, the ways a model invents things, and how it behaves under load. What happens when it fails matters more than what happens when it works, because the second one takes care of itself.
- Production integration with schema validation on every output
- Defined behaviour for low-confidence and refused responses
- Running locally, so capacity is planned rather than metered
- Rate limiting on anything publicly reachable
- A written record of the failure modes we found
QA
Break it before launch.
We test across devices, browsers, screen readers and slow connections. For NDIS sector clients WCAG 2.1 AA is not a checkbox at the end — it runs through the whole build, and this is where it gets verified rather than assumed. Contrast is checked by a script that fails the build, so a regression cannot quietly reach you.
- Cross-browser and device testing
- Screen reader and keyboard navigation pass
- Contrast audited by script, not by eye
- A written list of what was checked and what was fixed
No separate stage
On the AI track, checking is not something that happens once at the end. Every output is validated against a schema as it is produced, and the failure modes are worked through during Make.
Ship
Launch without holding your breath.
We handle DNS, deployment to Vercel, and smoke-testing live. We stay reachable in the days right after launch. Then we hand over: documented codebase, credentials in a password manager, and a written guide for whoever maintains it — so you're never dependent on us for day-to-day changes.
- Live deployment, smoke-tested against the real site
- Monitoring active from the first visitor
- Credentials handed over through a password manager
- Written maintenance guide in plain English
- Someone reachable in the days immediately after launch
Ship
Deploy with observability, not optimism.
Live deployment with monitoring active from minute one. Documentation covers AI behaviour, known model limits, and what to do when it fails — not just how to use it when it works. Handoff includes prompt files, evaluation datasets, and the exact models used, so whoever maintains this after us can reproduce the same result rather than flying blind.
- Live deployment with monitoring active
- AI behaviour documentation and known limits
- Prompt files and workflows versioned and handed over
- The models named, so the same result can be reproduced later
- What to check first when output starts looking wrong
Ready to start?
Discovery and Scope are the same first step: a conversation about whether the project is worth building, and which of the two tracks it belongs on.