Applied AI

AI that earns its place inside the organisation — by being governable, observable, and tied to outcomes worth measuring.

On this page
  1. Top of page
  2. Patterns we design our practice to avoid
  3. How we approach AI work
  4. What you walk away with
  5. Common questions from clients
  6. Evidence & references
  7. What to read next

§ 01 ·

It is rarely the right answer on its own. Most of the work is the surrounding plumbing: the data, the integration, the governance, the change management, and the model of how a decision moves from an AI recommendation to a trusted action.

The four phases the firm uses to do this work are on the How we work page.

§ 02 ·

Patterns we design our practice to avoid

Six failure modes we have seen repeatedly, across AI engagements of every size. Each one motivates a specific choice in how we work. The page opens with the failure modes because that is where the work is hardest to do well.

A model that does a hard thing badly is still a hard thing. Use cases should be selected for the value of the decision, not the novelty of the technology.

Most AI failures are data failures. If the data pipeline cannot produce a reliable, fresh, labelled signal, the model will not work in production regardless of its quality on the test set.

Most production AI systems include a person somewhere in the loop, even when the architecture does not show one. The handoff is the difference between a system people use and a system people route around.

A model that improves accuracy on a metric that does not matter is a model that has been optimised in the wrong direction. Outcome metrics are the ones that justify the investment.

When governance reviews happen after the system is built, the only available lever is approval or rejection. Effective governance sets the paved roads up front so teams move quickly within them.

An LLM is an untrusted component that produces plausible output. Treating it as authoritative, without an explicit boundary between its output and the action it triggers, is the most common path to a production incident.

§ 03 ·

How we approach AI work

The common approach treats the model as the deliverable. Our approach treats the system around the model: the data, the integration, the governance, the operating model, as the deliverable, and the model as one component of it.

The common approach

A model is trained on historical data, wrapped in a thin API, handed to delivery, and "released". Success is measured on accuracy against a holdout set. Failures are investigated by the data team alone.

  • Model accuracy is the primary success metric.
  • Governance is a checklist at the end.
  • Production behaviour is observed only when something breaks.
  • The model is the artefact; the system is an implementation detail.

What we do

A model is one component of a system designed for a specific decision or operation. Readiness across data, workflow, governance, and operations is assessed before the model is built. The evaluation harness and the operating model are part of the deliverable.

  • Outcome and trust metrics, not just model metrics.
  • Governance paved roads up front; the review board sees exceptions.
  • Continuous monitoring of model behaviour and business outcome.
  • The system around the model is the artefact; the model is a component of it.

What you walk away with

A flat list of the artefacts an AI engagement produces. The full shape of the engagement is on the [How we work](/services/methodology/) page.

A use-case shortlist

A short list of use cases the organisation is willing to pursue, scored on value, feasibility, and risk. The output is not a roadmap; it is a small set of bets.

A readiness assessment and integration design

A readiness assessment across the four axes (data, workflow, governance, operations) for each selected use case, and a design for where the model sits in the decision flow, what its inputs and outputs are, and how its output is reviewed and overridden.

A working system with its evaluation harness

The model in production, the surrounding integration built, and the evaluation harness in place. The harness is part of the deliverable, not a side artefact.

Operating-model documents and monitoring

An operating model: who monitors the model, who escalates, who retrains, who reviews. Monitoring in production for drift, output sampling, and outcome metrics. Governance review for significant changes.

The LLM safety pattern (if the use case involves an LLM)

A defined contract for how the LLM is invoked, what its inputs are validated against, what its outputs are checked against before they reach the user or the system, and what actions it never has authority to take without human review.

§ 05 ·

What each phase produces in an AI context is on the How we work page.

§ 06 ·

Common questions from clients

The questions we hear most often, in the order they tend to come up.

We start from the decision or operation, not from the technology. An AI investment is justified when the underlying task is one a machine can credibly do, where the value of doing it better is large, and where the cost of being wrong is bounded and recoverable.

AI is rarely the right answer when the value of being right exceeds the cost of being wrong by an order of magnitude, when the data does not exist, or when the simpler alternative: process change, a rules engine, a better interface, would close most of the gap.

Most AI engagements run between three and nine months. The framing phase is six to twelve weeks; the readiness assessment and integration design is the bulk of the early work; the build phase is sized to the use case; the operation phase is ongoing. The total horizon is set in the framing session and revised only by mutual agreement.

Engagements are priced against the use-case shortlist produced in the framing phase. The cost depends on the use case, the readiness of the data, the depth of the integration, and the operating model the work needs. A typical range for a multi-phase engagement sits in the low-to-mid six figures; the final figure is set in the scoping call after the framing phase.

The first 30 days are the framing phase: an enumeration of candidate use cases with the leadership team, structured conversations with the people who will own the engagement, and the production of the use-case shortlist. By the end of the first 30 days, the client has a small set of bets it is willing to make.

An LLM in production is treated as an untrusted component with a defined contract. Inputs are validated and bounded. Outputs are checked against the contract before they reach the user or the system. The LLM never has authority it has not been explicitly given, and every action it influences is logged.

This is unglamorous work and most of it. The model is the visible part; the safety, evaluation, and observability layers around it are what make the system production-grade.

§ 07 ·

Evidence & references

Public frameworks and writing that inform the firm's practice.

Designing Machine Learning Systems
Chip Huyen, O'Reilly, 2022

The end-to-end treatment of the operational side of ML — data, feature stores, monitoring, continual learning.

MLOps — Engineering Machine Learning Models
Treivel, Gombar & Moreau, O'Reilly, 2020

The pattern language for operationalising models: deployment, monitoring, drift detection, retraining cadence.

Designing Data-Intensive Applications
Martin Kleppmann, O'Reilly, 2017

The data-pipeline framing: data flow, data lineage, the contracts between systems. Most of an AI engagement's hard work lives at this layer.

Power and Progress
Daron Acemoglu & Simon Johnson, PublicAffairs, 2023

The technology-and-society framing. The choice of which AI work to do is a choice the firm and the client are making about what the technology is for, not a technical question with one right answer.

§ 08 ·

What to read next

The shape of the engagement, and the services the AI work most often sits next to.

Want to talk AI?

If you are evaluating a use case, building a model into a production system, or strengthening an existing AI practice, the firm is useful at that boundary. A short conversation is the right next step.

Start a conversation