Most organisations do the steps in the wrong sequence: they pick the actors, buy the tools, and think about risk and ownership afterwards. Reverse the order and the same work produces something defensible.

I did not arrive at the sequence below in a workshop. It is what came out of conversations with leaders at Tech Week Singapore 2026 about where their AI programmes had gone wrong, and it is the order I would run.

What enterprise AI adoption usually looks like

Most of what gets called enterprise AI adoption is a procurement decision. A vendor is signed, seats are issued, and the organisation is told it now has AI. Nothing about the work has changed.

The models are rented. The capability is not. What the purchase bought is access, and access is worth nothing until someone changes how a process runs — which is a change-management problem wearing a technology costume.

That reframing matters because it moves the hard part. It is not picking the model. It is getting a few hundred people to run their work differently and to trust a system they did not build.

Why AI literacy comes before the workflow

Literacy comes first, across the whole organisation, not just the technical teams. Not prompt engineering. A working grasp of what these systems are good at, where they fail quietly, what they cost per call, and which decisions should never be handed to one.

Skip it, and the sequence fails in three predictable ways. Leaders cannot tell a real use case from a demo, so they approve the wrong ones. The people doing the work treat the new system as a threat or an audit, so they route around it. IT gets asked to govern something nobody can describe.

Literacy is also the cheapest thing on this list. It is a cost in time and attention, not in licences, and it is the only step that makes every later step possible. If it is everyone's responsibility, it is nobody's — so name the person who owns it and the standard they are teaching to.

Choosing the first workflow

Do not start with the most interesting problem. Start with a workflow that touches expenses or revenue, and pick one where the current number is already being measured.

Then write the objective and key results before anything else. One I have used as an example: reduce the volume of support tickets handled by a human by 20%. It is measurable, it has a baseline, and it matters to somebody who signs off budget.

The trap is in the wording. If the key result becomes "automate 20% of the work", you will get automation at the cost of the outcome — tickets closed twice, customers bounced, the number hit and the service worse. Write the human outcome as the key result and put the agent's throughput underneath it as a signal, never as the goal.

Two more things belong in this step: the number today, measured rather than remembered, and the do-nothing option. If spending nothing changes nothing, the case is not there yet.

The feasibility study, and why it comes first

Before decomposing work, establish whether the workflow can actually be improved. Four questions, and the honest output is often "not yet".

Desired impact. What changes outside the team, in one sentence, with no technology words in it. If that sentence is hard to write, the rest is a science project.

Return. What the improvement is worth per year, against what it costs to build and to run. Be specific about the unit: cost per completed task at a stated volume, including review and rework time. Agent cost is variable and human cost is fixed, so a workflow can be automated end to end and still save nothing if a human sits at the bottleneck.

Data readiness. Does the data exist, is it reachable, is it clean enough, and can it legally be used for this. Most AI initiatives are blocked here and nowhere else, and the block is discovered late because nobody asked early.

Infrastructure readiness. Where does it run, who has access, how does it reach the systems it needs, and can it be logged. A workflow that needs three integrations nobody owns is not a two-week build.

Report readiness per item: ready, needs work, or not feasible yet. What survives goes forward. What does not becomes a named prerequisite with an owner — which is how one AI initiative turns into the data programme nobody had to argue for.

Task decomposition, with the who left empty

Now, and not before, list the tasks in order. Every task gets a what and a when, and each one names the output it produces. A task with no named output is a task nobody can verify or hand over.

Leave the who column empty. This is the single most important ordering decision in the whole sequence. The moment you write "agent" next to a step because that is what you came here to do, everything downstream becomes a justification of a choice you already made.

Mark the tasks honestly while you are here. Which are judgment, where two competent people could disagree on the output. Which are rules, lookups, calculations and checks — those do not need a model, and putting one there buys variance for nothing.

Two things to write before starting. A process owner for the workflow as a whole, because per-task owners alone leave the handoffs unowned and the handoffs are where these fail. And a kill criterion: if every row can only be done by an agent with no fallback, the workflow should not start.

Risk triage, before any actor is named

For each task, ask what it can break. Three questions, in this order.

Blast radius. If this task is wrong, who notices and how many people are affected. One record, one customer, or the whole enterprise.

Reversibility. Can it be undone, in what time, at what cost, by whom. Cheap to undo is a lower tier.

Data sensitivity. What data does it touch, and what is the worst realistic disclosure.

Those answers produce a tier, and the tier sets every decision after this — how deep the controls go, how much review is proportionate, what evidence gets kept.

Tiering by risk is what stops both failure modes at once: over-controlling trivial steps until nobody uses the tool, and under-controlling the steps that matter. Two published frameworks carry most of the weight here, and each applies under its own condition. IMDA's Model AI Governance Framework for Agentic AI is the one to reach for when an agent may act with real autonomy: it grades a task on the scope of its actions, whether they can be reversed, and how much autonomy it runs with, which is the same three-part question as the one above. MAS's Guidelines on AI Risk Management apply to financial institutions, so use them as the reference when a supervisor will want to see the register: they ask for a risk materiality assessment covering impact, complexity and reliance, and for controls proportionate to the tier that comes out. Where neither condition holds, tier the tasks anyway. A tier nobody is inspecting is still better than an actor nobody derived.

Allocating the who from the risk tier

With a tier in hand, the actor reads off — the tier constrains the choice rather than decorating it. Deterministic rule where there is one right answer. An agent where the work is bounded judgment with a fixed set of labels. A human in the loop where the action is costly to undo or touches a customer. A human, with a model assisting, where the call is contested, novel, or has no right answer.

The failure mode is choosing the actor and then writing the risk assessment to fit. It is easy to spot in a plan: a task list where every step is an agent and the risk section mentions governance generically. That plan is a preference with a document attached.

Two rules make the allocation real. A human-in-the-loop step is only a control if the person can stop the line — one approver clearing four hundred items in a row without reading them is an agent with a person's name on it. And name a human, not a function: "ops approves" is not accountability. Every task needs one accountable name, and the organisation needs one person with the authority to stop a deployment.

Controls as their own task list

The tier says how much control; this step says which controls, and it is the same who/what/when exercise run over the controls themselves — with an agent as one of the options, and deliberately not the default.

For each control: where it is enforced, what it blocks, who reviews it, and what evidence it leaves. Enforcement belongs in the system, not the prompt, and that holds whether or not a model is anywhere in the step. A prompt cannot refuse; a permission check can. A guardrail that asks the model politely to respect human approval is a suggestion, and suggestions get bypassed or forgotten.

Then check that the control still works, because a control with no number beside it has already decayed. Two tests catch most of it: an override rate near zero means nobody is watching, and an approval that completes in a second means nobody read it. Human oversight is a mechanism that has to be audited like any other, or it quietly becomes decorative.

The one control an agent should never own is the authority to stop itself. Detection can be automated. The decision to halt belongs to a person, and at the top tier it has to be a person with the authority to use it.

Implementation, which is shorter than everything around it

The build is often the shortest part, days rather than months, because the thinking was done in the steps above. Three things belong here.

A contract per task. Schema in, schema out, validated at the boundary by code. When an input is missing or stale, fail closed and stop the flow rather than carrying on with a guess.

Cost, measured before it runs. Cost per completed task, then divided by the success rate. A cheap model that fails a third of the time is the more expensive one. Set a ceiling per workflow and one documented way to stop the spend today — a budget without a switch is a forecast.

Determinism where it can be had. Only the judgment steps get a model, and whenever an output has to be traceable to a source, they return identifiers rather than words. A model that selects which sentence matters and returns its id cannot invent a sentence it never typed. Most of the engineering value in these systems is here, and it does not demo well, which is why it gets skipped.

The operating concept, decided before the build

Define the handoff and the operating concept up front. Not because process demands it, but because a build that takes days leaves no time to invent a maintenance model afterwards — and a system nobody owns starts decaying the day it ships.

The operating concept covers the whole lifecycle: who monitors it, what triggers a rebuild, how it is retired or replaced. Retirement is the part that gets forgotten and it is the part that hurts. Retired agents keep live credentials, and a forgotten agent with a valid key is how the quiet incidents happen. Give every agent its own identity so revocation is clean, rehearse the manual fallback, and schedule a review of the fallback rather than trusting a document that says one exists.

Then measure, or it was never an objective. Review the key result against the baseline on a stated cadence, and watch the two numbers that tell you whether the humans are still in the loop: override rate and approval latency. A workflow that automated the work but lost the oversight is not a success with a caveat. It is a failure that has not surfaced yet.

A worked example: verifiable or it does not publish

I built a system that reads every sitting of Singapore's Parliament and publishes a summary of each one with the summarised sentences marked inside the transcript. A quotation attributed to the wrong speaker would have sunk the whole thing, so the requirement came before any design existed: verifiable, or it does not publish.

The two designs I tried first both failed the same way — by asking a model to be accurate and then trying to catch it when it was not. The first had the model write and select the quotes; it repeated itself and produced sentences that were not in the transcript at all, and the checker waved them through because it only ever saw what the model produced. The second narrowed the input but kept surfacing procedure instead of substance.

The fix was not a better prompt or a stricter checker. It was moving the selection into a database and reducing the model to one job: decide which sentences matter and return their identifiers. The system drops the real words in afterwards. A fabricated quotation is not a risk to catch — it is a state the design cannot reach.

And when a check fails, nothing is published. The failure I care about is not a missing page. It is a wrong one going out while nobody is looking. That is the whole argument of this article in one rule, and it is the rule I have to defend hardest, because a caveat is cheap and a missing page is visible.

Why the order is the control

Enterprise AI adoption is not a technology purchase, and the sequence is what keeps it honest. Literacy comes before the workflow, the feasibility study before the plan, the risk tier before the actor, and the operating concept before the build.

Each step is cheap next to the one it gates, and each one removes work from the step after it. That is what makes the build days rather than months, and what makes the result something an organisation can keep running after the people who built it have moved on.

Get the order wrong and no amount of model capability recovers it. Get it right, and the technology decision — which model, which vendor, which framework — becomes the smallest decision in the programme, which is where it belongs.

The pieces behind this are what a system does when it cannot be sure and the Parsnips case study, which is the honest version, failures included. Both, and anything after them, sit on the Writing page.