AI at Work

How to Run an AI Pilot at Work

Run an AI pilot on one workflow, with a baseline, a data class, a review rule, and a kill date. Demos are not pilots.

Keywordshow to run an ai pilot at workBlogai for workai at workworkplace aiai for professionalshow to use ai at workai productivity

Related searcheshow to run an ai pilotworkplace ai pilotai proof of concept at workai pilot scopeai pilot kill datehow to run an ai pilot at work

A pilot has a baseline and a kill date. A demo has a round of applause

Run an AI pilot on one workflow, with a baseline, a data class, a review rule, and a kill date. Demos are not pilots. A demo proves a model can write a fluent page in a meeting. A pilot answers whether a named team, on real files, under your policy, can reach a trusted artefact faster or with fewer errors than they do today. If you cannot say what you will measure, what you will not put in the tool, who reviews, and when you will stop, you are hosting a showcase. Showcases create licence requests. They do not create an operating method.

Most failed AI programmes begin as unbounded trials. Everyone may try anything. Confidential data leaks into consumer chats. Success is a slide of screenshots. Nobody recorded how long the old pack took, so nobody can say whether the new path is better. The team then argues from vibe. Vibe is not a governance artefact. A bounded pilot is slightly boring on purpose: one job, one tool, one class of data, a start and an end, and a written rule for what happens to the work if the numbers are bad.

This essay is that bound. You will choose a workflow that already has an owner, capture a baseline without theatre, freeze the risk controls, run for a short fixed window, and decide keep, change, or kill with evidence. You will not add a second use case mid-flight because someone had a clever weekend. Scope creep is how a pilot becomes an unofficial production system with no owner. Write the fence before the first prompt. If a new idea cannot wait, it waits on a list, not in the live path.

If there is no kill date, it is not a pilot. It is production without the honesty of calling it production. Write the date before the first prompt.

The four fences of a workplace AI pilot

The workflow fence is one artefact: a weekly commentary, a ticket summary, a first-pass vendor table. The data fence is a class and a tool that is allowed to see it. The review fence is who checks names, numbers, and promises, and how long that takes. The time fence is a start date, a sample plan, and a kill date that does not move because a senior person is keen. Write the four fences on one page. If a proposed extra use case cannot fit those fences, it waits. Waiting is the point of a fence.

Baseline before tools. Time the current path on a few real instances: minutes to a draft you would sign, error types you already know, rework after review. Do not baseline a heroic week. Baseline a normal one. Then run the new path on comparable instances, not on a polished demo pack. Keep people and policy constant so you are not attributing a good manager to a model. If the pilot needs a new legal opinion, get it before day one. A pause for counsel during week two is a sign you started as a demo.

FenceWrite this downPilot is broken if
WorkflowOne artefact and its ownerA second job is added mid-run
Data classWhat may go in which toolSomeone pastes a higher class
ReviewNamed checker and the five checksSend without the check to hit a date
TimeStart, sample plan, kill dateThe date moves because of enthusiasm

Run the window, sample, then decide in writing

Staff the pilot with people who already do the job, not with a digital team performing the job for them. Give them the brief, the tool, and the review rule. Sample twice a week: time-to-trusted-output, errors found in review, and any policy near-miss. Log those in a table, not in chat. Do not change five variables at once. If the brief is wrong, change the brief and say so. If the tool is wrong for the data class, stop. Curiosity is not a reason to keep a leaking path open until the kill date.

On the kill date, compare the table to the baseline. Decide keep, change one variable and re-run, or stop. Write the decision where the original fence lived. Stopping is a good outcome when the evidence is poor. Extending without a new fence is how shadow production starts. If you keep the method, you then do the change work: default, training, surplus time. A successful pilot that nobody operationalises is a report. Plan that step before you celebrate, including who will coach the default in week one of production. Name that coach in the decision note.

Role: You are a programme lead who refuses to treat demos as pilots.
Task: Draft a one-page AI pilot brief from the notes I paste.
Context: I will name the workflow, owner, tool, data class, and dates.
Constraints:
- Allow only one workflow.
- Require a baseline, a review rule, a sample plan, and a kill date.
- If any of those is missing in my notes, write missing instead of inventing it.
- Do not add success metrics based on word count or licence use.
Output: Fences, baseline plan, sample table headings, decision rule for the kill date.
Quality checks: What would turn this brief back into a demo?

Pin the one-page brief next to the work, not in a steering pack that meets monthly. When someone asks to add a use case, point at the fence. When someone asks how it is going, open the sample table. At the kill date, write keep, change, or stop in a sentence a sceptic could read. Then either operationalise or shut the path. A third state called still seeing how it goes is unofficial production.

Demo theatre, moving dates, and unpaid production

Demo theatre uses a clean file, a prepared prompt, and a friendly audience. Real files are messier and will beat that memory. Moving the kill date because leaders liked the demo converts a test into an entitlement. Unpaid production is the quietest failure: the pilot tool becomes how the pack gets made, with no owner, no logging, and no decision to accept the risk. Security will find it later. So will a client, if the wrong class of data went in. Write production criteria before you need them, including identity, retention, and who signs.

Watch success metrics that cannot fail. Number of prompts, number of users, number of words. Those numbers rise when the work is worse. Hold to time-to-trusted-output, error rate in sample, and near-misses. If you cannot staff the sample, you cannot staff a pilot. Run a smaller window. A two-week honest test beats a quarter of unmeasured enthusiasm. Enthusiasm is cheap. A table that can show failure is the artefact that makes the kill date real rather than social. If the table is empty on the kill date, you do not yet have a pilot.

MistakeWhat it looks likeWhat to do instead
Demo as proofOne clean file in a meetingReal instances and a baseline
No kill dateWe will know when it feels rightA date that does not move
Scope creepJust this extra use caseThe extra waits
Vanity countPrompts and licencesTrusted output and errors

A pilot without a kill date is production. Call it that, or give it a date you will honour when the evidence is dull. Honour the date even if leaders liked the demo.

Related reading on StudyGrid

Read next: Adoption and Value How to Measure AI Productivity at Work Change Management for Workplace AI. Those essays sit beside this one. Use them when you need the neighbouring skill, not as a substitute for the check you still have to make.

What to do this week

Choose one workflow you already own. Write the four fences and a baseline from three recent instances. Set a kill date within four weeks. Run only that path. Sample twice a week. On the date, write keep, change, or stop. Do not start a second pilot until that sentence exists. The sentence is the difference between a test and a habit you never chose. Put it where the original fence lived, not only in a slide.

FAQ: How to Run an AI Pilot at Work

Common questions about this page.

How do I run an AI pilot at work?

Pick one workflow, record a baseline, lock the data class and the review rule, set a kill date, and compare time-to-trusted-output with the old path.

What is the difference between an AI demo and a pilot?

A demo shows that the model can produce text. A pilot tests whether a named team can produce trusted work, under real constraints, for a fixed period.

When should I stop an AI pilot?

Stop on the kill date if quality, risk, or time-to-trusted-output did not beat the baseline. Stopping is a successful pilot, not a social failure.

Is this StudyGrid essay free?

Yes. The full blog on StudyGrid (studygrid.in) is free. Open Blog in the header, or follow Previous and Next at the bottom of each essay.

Where should I start the StudyGrid blog?

Start at The AI Opportunity if you want the series in order. Open a single essay if you searched for a specific workplace task such as email, Excel, policy, or prompting.