Keywordshow to run an ai pilot at workBlogai for workai at workworkplace aiai for professionalshow to use ai at workai productivity
Related searcheshow to run an ai pilotworkplace ai pilotai proof of concept at workai pilot scopeai pilot kill datehow to run an ai pilot at work
A pilot has a baseline and a kill date. A demo has a round of applause
Run an AI pilot on one workflow, with a baseline, a data class, a review rule, and a kill date. Demos are not pilots. A demo proves a model can write a fluent page in a meeting. A pilot answers whether a named team, on real files, under your policy, can reach a trusted artefact faster or with fewer errors than they do today. If you cannot say what you will measure, what you will not put in the tool, who reviews, and when you will stop, you are hosting a showcase. Showcases create licence requests. They do not create an operating method.
Most failed AI programmes begin as unbounded trials. Everyone may try anything. Confidential data leaks into consumer chats. Success is a slide of screenshots. Nobody recorded how long the old pack took, so nobody can say whether the new path is better. The team then argues from vibe. Vibe is not a governance artefact. A bounded pilot is slightly boring on purpose: one job, one tool, one class of data, a start and an end, and a written rule for what happens to the work if the numbers are bad.
This essay is that bound. You will choose a workflow that already has an owner, capture a baseline without theatre, freeze the risk controls, run for a short fixed window, and decide keep, change, or kill with evidence. You will not add a second use case mid-flight because someone had a clever weekend. Scope creep is how a pilot becomes an unofficial production system with no owner. Write the fence before the first prompt. If a new idea cannot wait, it waits on a list, not in the live path.
If there is no kill date, it is not a pilot. It is production without the honesty of calling it production. Write the date before the first prompt.
The four fences of a workplace AI pilot
The workflow fence is one artefact: a weekly commentary, a ticket summary, a first-pass vendor table. The data fence is a class and a tool that is allowed to see it. The review fence is who checks names, numbers, and promises, and how long that takes. The time fence is a start date, a sample plan, and a kill date that does not move because a senior person is keen. Write the four fences on one page. If a proposed extra use case cannot fit those fences, it waits. Waiting is the point of a fence.
Baseline before tools. Time the current path on a few real instances: minutes to a draft you would sign, error types you already know, rework after review. Do not baseline a heroic week. Baseline a normal one. Then run the new path on comparable instances, not on a polished demo pack. Keep people and policy constant so you are not attributing a good manager to a model. If the pilot needs a new legal opinion, get it before day one. A pause for counsel during week two is a sign you started as a demo.
| Fence | Write this down | Pilot is broken if |
|---|---|---|
| Workflow | One artefact and its owner | A second job is added mid-run |
| Data class | What may go in which tool | Someone pastes a higher class |
| Review | Named checker and the five checks | Send without the check to hit a date |
| Time | Start, sample plan, kill date | The date moves because of enthusiasm |
Run the window, sample, then decide in writing
Staff the pilot with people who already do the job, not with a digital team performing the job for them. Give them the brief, the tool, and the review rule. Sample twice a week: time-to-trusted-output, errors found in review, and any policy near-miss. Log those in a table, not in chat. Do not change five variables at once. If the brief is wrong, change the brief and say so. If the tool is wrong for the data class, stop. Curiosity is not a reason to keep a leaking path open until the kill date.
On the kill date, compare the table to the baseline. Decide keep, change one variable and re-run, or stop. Write the decision where the original fence lived. Stopping is a good outcome when the evidence is poor. Extending without a new fence is how shadow production starts. If you keep the method, you then do the change work: default, training, surplus time. A successful pilot that nobody operationalises is a report. Plan that step before you celebrate, including who will coach the default in week one of production. Name that coach in the decision note.
Role: You are a programme lead who refuses to treat demos as pilots.
Task: Draft a one-page AI pilot brief from the notes I paste.
Context: I will name the workflow, owner, tool, data class, and dates.
Constraints:
- Allow only one workflow.
- Require a baseline, a review rule, a sample plan, and a kill date.
- If any of those is missing in my notes, write missing instead of inventing it.
- Do not add success metrics based on word count or licence use.
Output: Fences, baseline plan, sample table headings, decision rule for the kill date.
Quality checks: What would turn this brief back into a demo?Pin the one-page brief next to the work, not in a steering pack that meets monthly. When someone asks to add a use case, point at the fence. When someone asks how it is going, open the sample table. At the kill date, write keep, change, or stop in a sentence a sceptic could read. Then either operationalise or shut the path. A third state called still seeing how it goes is unofficial production.
Demo theatre, moving dates, and unpaid production
Demo theatre uses a clean file, a prepared prompt, and a friendly audience. Real files are messier and will beat that memory. Moving the kill date because leaders liked the demo converts a test into an entitlement. Unpaid production is the quietest failure: the pilot tool becomes how the pack gets made, with no owner, no logging, and no decision to accept the risk. Security will find it later. So will a client, if the wrong class of data went in. Write production criteria before you need them, including identity, retention, and who signs.
Watch success metrics that cannot fail. Number of prompts, number of users, number of words. Those numbers rise when the work is worse. Hold to time-to-trusted-output, error rate in sample, and near-misses. If you cannot staff the sample, you cannot staff a pilot. Run a smaller window. A two-week honest test beats a quarter of unmeasured enthusiasm. Enthusiasm is cheap. A table that can show failure is the artefact that makes the kill date real rather than social. If the table is empty on the kill date, you do not yet have a pilot.
| Mistake | What it looks like | What to do instead |
|---|---|---|
| Demo as proof | One clean file in a meeting | Real instances and a baseline |
| No kill date | We will know when it feels right | A date that does not move |
| Scope creep | Just this extra use case | The extra waits |
| Vanity count | Prompts and licences | Trusted output and errors |
A pilot without a kill date is production. Call it that, or give it a date you will honour when the evidence is dull. Honour the date even if leaders liked the demo.
Related reading on StudyGrid
Read next: Adoption and Value How to Measure AI Productivity at Work Change Management for Workplace AI. Those essays sit beside this one. Use them when you need the neighbouring skill, not as a substitute for the check you still have to make.
What to do this week
Choose one workflow you already own. Write the four fences and a baseline from three recent instances. Set a kill date within four weeks. Run only that path. Sample twice a week. On the date, write keep, change, or stop. Do not start a second pilot until that sentence exists. The sentence is the difference between a test and a habit you never chose. Put it where the original fence lived, not only in a slide.