A colleague has produced an impressive AI demonstration using material from the organisation. The output is fluent, the team can see several possible uses and a supplier is ready to arrange a pilot.

What has not been established is which operational problem the pilot will solve, what information the tool may receive, who remains accountable for the result or how success will be judged after human checking.

Without those decisions, the organisation is not running a pilot. It is allowing a tool to search for a role.

A safe and useful pilot starts with one bounded workflow, controlled information, clear human accountability and measures agreed before the first result. This does not require a data-science team. It requires the same operational discipline used for any change that affects information, decisions and service quality.

The pilot should answer one operational question

“Can we use AI?” is too broad to test. “Can an approved tool help the policy team produce a first summary of published guidance without increasing factual errors or total review time?” is a pilot question.

The second version identifies the workflow, owner, information boundary, intended benefit and important failure. It can produce a decision.

In early AI meetings, the tool can be described in detail while the current workflow survives only in the head of the person who performs it.

Suitable early workflows happen often enough for the result to matter, but remain narrow enough to describe from input to outcome. They use information the organisation is permitted to provide, already belong to an accountable team, can be reviewed by somebody with relevant knowledge and can be reversed if the trial stops.

Examples may include preparing a first summary of approved internal material, categorising non-sensitive enquiries, suggesting tags for published content or helping staff retrieve passages from a controlled policy collection.

Avoid beginning with autonomous action, a member-facing decision or a workflow where an error could materially affect eligibility, assessment, safeguarding or payment. The issue is not that these uses are impossible. They are poor places to learn the basics of ownership, quality and control.

Write down the current process before adding AI. Record volume, elapsed time, review effort, common errors and the person accountable today. Otherwise the pilot can generate visible activity without showing improvement.

A pilot should reduce uncertainty, not merely produce output.

The Bounded Workflow Test

Before choosing a tool, score the proposed workflow against six questions:

Test A credible early pilot has
Purpose One defined operational outcome, not a general invitation to experiment
Boundary Named inputs, outputs, users and actions that are outside scope
Consequence Errors that can be found and corrected before material harm
Review A competent person who can verify each output efficiently
Measure A baseline and pre-agreed value, quality, cost and failure measures
Reversibility A practical way to stop without losing an essential service or important information

Classify each line ready, needs control or unsuitable for this pilot. A weak line does not always end the idea. It may indicate that the workflow needs to be narrowed or that a prerequisite must be resolved first.

The test also prevents technology choice from dictating scope. Teams are easily drawn towards the most impressive feature in a demonstration. The better question is which bounded change will generate useful organisational evidence.

In pilot selection meetings, visible output drives the choice while the effort needed to verify it is left for later design.

Draw the information boundary

State which information may enter the workflow, where it may be processed and stored, who may access it and when it must be deleted. Also state what is prohibited.

“Do not use sensitive data” is not an operating control. Staff need an approved information set and the tool should have no wider access than the workflow requires.

If personal information is involved, bring in the organisation’s data protection owner to review the actual workflow, information and supplier terms. A generic statement that the pilot will comply with policy does not establish what staff may enter, what the service retains or what evidence the organisation will keep.

Before approving a supplier service, follow one real pilot transaction. Trace the prompts, files, outputs and logs through processing and storage, including who can access them and whether they are used to train or improve any service. Then test how retention, deletion, security, supplier changes and exit would work in practice.

Supplier review meetings can spend more time on model quality than on the quieter question of who can see prompts, outputs and logs after the meeting ends.

Supplier assurances are useful evidence, but they are not the organisation’s information decision. In early reviews, the contract can be scrutinised closely while nobody has written down the prompts, files and outputs used in the actual workflow.

Not every early pilot needs integration. Manual transfer of approved information may be less efficient, but it can test output quality before the organisation grants broader access or creates technical dependency.

Make human accountability operable

“A human is in the loop” sounds reassuring but says little. Name the person or role, the decision they retain and the checks they must complete.

Most organisations do not struggle because they chose the wrong AI tool. They struggle because nobody agreed where human judgement stopped and automation started.

The reviewer may need to verify facts against source material, check that important qualifications are preserved, remove confidential content, assess tone and decide whether the output can be used. They also need a route for rejecting an output and recording why.

Human review is a control only when the reviewer has the competence, source material, time and authority to say no.

One common pilot result is that drafting becomes faster while checking becomes slower. If only generation time is measured, the tool appears valuable even when total effort has increased. Record review time separately and include rework.

Another warning sign is that reviewers gradually trust fluent output and reduce their checks. Sample quality assurance by the workflow owner can reveal whether the intended control is still operating.

Make the retained decision visible:

A bounded AI workflow in which approved information is AI-assisted, human-verified, recorded when it fails and used only after a human decision.

Measure value, failure and dependency

Agree measures before the pilot begins. They should cover:

  • total time per item, including review and correction;
  • accuracy or quality against an agreed sample and standard;
  • types and frequency of material errors;
  • service effect, such as response time or consistency;
  • staff confidence and usability after the novelty period;
  • cost at a realistic volume;
  • governance, support and supplier effort; and
  • new dependency on a provider, model, integration or specialist prompt.

Keep examples of failures as well as successes. A controlled pilot is valuable precisely because limitations can be found before the workflow becomes important.

Pilot readouts accumulate the best outputs; the failures are scattered across chat messages, reviewer notes and examples that nobody retained.

Set stop conditions in advance. These might include prohibited information entering the service, a material error pattern, a supplier control that cannot be verified, review effort removing the benefit or staff using outputs outside the approved workflow.

Stopping is not a failed pilot when it answers the question. Continuing without a decision is.

The one-page Pilot Brief

The brief should be short enough to review in one meeting:

Brief section Decision to record
Operational question The specific change being tested and why it matters
Current baseline Volume, time, quality, failure and current owner
Workflow boundary Approved inputs, AI-assisted step, human decision and permitted output
Information control Classification, access, processing, retention and prohibited uses
Supplier evidence Contract, security, data use, change, exit and incident answers
Measures Value, quality, review effort, cost and dependency
Stop conditions Events that pause or end the pilot
Final authority Named owner and the date for the decision

Where no pilot brief exists, purpose and controls can change in separate conversations while everybody continues to refer to the work as the same pilot.

Use existing governance proportionately. Data protection, information security, procurement and operational ownership still apply when AI is involved. A new AI committee is not a substitute for the people already accountable for those areas.

A short internal trial using approved public information needs less assurance than a member-facing service using personal data. Both need a purpose, owner, boundary and recorded decision.

End with a decision, not another demonstration

At the agreed end date, choose one of four outcomes:

  1. Stop: the value is insufficient or the risk and dependency cannot be justified.
  2. Revise: narrow the workflow, strengthen a control or test a material uncertainty again.
  3. Adopt within limits: operate the workflow with named ownership, monitoring and review.
  4. Expand deliberately: invest in integration or wider use through a new, governed decision.

Do not let a time-limited pilot quietly become a business process because people enjoy using it. That is a common route from a controlled experiment to an unowned service.

Bring one proposed workflow to a 60-minute review with its operational owner, information or data protection owner, technology lead and the person who would verify the outputs. Complete the Bounded Workflow Test and one-page Pilot Brief before approving a tool.

If the group cannot define the human decision, information boundary, measures and stop conditions, the workflow is not ready to pilot. The uncertainty has been found early, when it is still inexpensive to resolve.