Limewater Labs
AI adoption

Make AI actually productive for your team — safely, and with proof.

Not a licence rollout and not a workshop nobody applies. A spec-driven playbook, agents built for the jobs that repeat, and review gates that keep humans accountable — measured on one team before it goes anywhere near your core systems.

Where this usually starts

The tools are already in the building. The results aren't.

Everyone's promising AI miracles. We can't tell what's real.

We bought the licences. Usage is low and nobody can show a result.

Our seniors won't let AI near the code that matters — and they're not wrong.

One team is flying with it. Nothing they learned is written down.

How we do it

Don't roll out a tool. Change how the work gets decided.

Spec-driven development is table stakes vocabulary by now — every major tool ships a flavour of it. What decides whether adoption sticks is who applies it, to which systems, and what happens when the AI is confidently wrong.

Specifications first, code second

With AI the code is the cheap part; deciding what to build is not. Teams that get little out of AI are almost always feeding it vague intent — so we start with how a specification gets written and reviewed, not with which tool to buy.

A playbook, not prompting

Ad-hoc prompting gives a different result every time, which is why pilots don't survive contact with a real backlog. A playbook is the deterministic version: the same steps, reviewable, repeatable by someone who wasn't in the room.

Specialized agents for the jobs that repeat

Generic assistants plateau fast. The leverage is in narrow agents built for one recurring task — cataloguing features in legacy code, extracting specifications, drafting technical documentation from tickets.

Human review gates where mistakes are costly

On core business and clinical systems the plan gets reviewed, the spec gets validated, and drift between spec and code gets caught deliberately. AI stays accountable because a named person signs off at each gate — never vibe-coding your core systems.

Clear the blockers before the training

AI cannot help a team that can't run its own stack locally. Locked-down desktops, twenty-minute builds, no usable test data — we fix the environment first, because that is what quietly kills most adoption programmes.

Measured on one team, then scaled

One pilot team, instrumented, with usage guidelines and capacity planning written down. If the numbers move, the playbook rolls out; if they don't, you have spent a month instead of a budget cycle.

Under the hood: spec-driven dev: Spec Kit · BMAD · OpenSpec · layered specs: raw → business → technical · custom agents for legacy analysis · CLI-first, fresh-thread agent workflows · human review gates + spec-drift checks
What the first three months look like

Adoption you can measure, not a licence count.

  1. Week 1

    Workshop with the people who have to use it

    Hands-on, with your codebase and your backlog: specification strategies, iterative AI workflows, test automation. Engineers leave having done the work once, not having watched a demo.

  2. ~1 month

    One pilot team, instrumented

    A real slice of the roadmap delivered the new way, with baseline and after numbers — throughput, cycle time, review load. The pilot is chosen so the result is arguable in front of a sceptic.

  3. ~3 months

    A playbook the rest of the org can run

    The steps, the review gates, the agents and the guidelines, written down and owned by your team. Adoption stops depending on whoever was enthusiastic first.

  4. From then on

    Yours, not ours

    We would rather leave than be needed. The measure of the engagement is whether your team keeps improving the playbook after we stop attending the stand-up.

Proof, not adjectives

We have done this where mistakes are expensive.

Enterprise ERP modernization

AI put to work where nobody remembered how the system worked

Constraint
A ~30-year codebase across multiple legacy layers, a hard deadline, and no one left who could say what the system actually did.
Approach
Specialized agents for legacy analysis, feature cataloguing and specification extraction. Specs generated with AI, reviewed by humans, and prioritised using complexity metrics from the legacy code.
Result
Validated specifications out of code no one fully remembered — and a day-and-a-half proof that the migration was real.

And as an adoption programme rather than a delivery one:

  • Hands-on AI workshops for several development teams at a pharmaceutical company — specification strategies, iterative workflows, test automation.
  • AI usage guidelines, coding-tool evaluation and capacity planning for a healthcare software organization adopting AI on brownfield code.
  • Custom agents for decades-old COBOL at a public-safety software vendor: reverse-engineering what the system does before anyone proposes a migration.
  • The playbook methodology presented to 100+ engineers across an industry AI summit, and a keynote on AI-first modernization at another.

On one of these programmes the pilot measured a ~20% productivity gain and the highest sprint throughput the team had recorded. That is the shape of evidence we aim for, because “the developers like it” does not survive a budget conversation.

Client work is confidential; described here in general terms. See all three engagements — or the two people who would do the work.

Start a conversation

Tell us where AI is supposed to be helping and isn't.

Start with a low-commitment AI readiness assessment: we look at what your team already tried, what the environment and the code allow, and give you a straight answer about where AI pays off first — and where it should stay out.

Email us to get startedhello@limewaterlabs.com