AssociationAI / AI Literacy
Trihelix AI team Published

Tutorial

Beginner

Run Your First AI Pilot on One Workflow in Two Weeks

A two-week plan for testing AI on one real workflow: pick the task, record a baseline, set a goal and a stop rule, run it with a human check, then measure and decide.

Time needed: About 30 minutes of setup, then 15 minutes a day for two weeks

Before you start:

  • One repeatable workflow you own, such as drafting recaps, summarizing notes, or sorting inquiries
  • An AI tool your organization has approved, used under your AI policy
  • A notebook or shared document to use as the pilot log

AI Pilot Operations Getting Started Association staff

By the end of this tutorial you will have run a two-week AI pilot on one real workflow, with a baseline, a goal, a stop rule, and a written decision about what to do next.

The hesitation is well documented. University of York researchers surveying fundraisers found ethics, data privacy, and reliability concerns were the barriers to bringing AI into daily work. A pilot answers those concerns with evidence from your own workflow. We built this plan around two weeks: long enough to see a pattern, short enough to finish. No source prescribes a length; two weeks is our design choice.

Plan your fortnight: the pilot planner turns your workflow, baseline, goal, and stop rule into a day-by-day two-week plan with a measurement log and a go/no-go decision. The planning worksheets are free with your email.

Week one: set up and run

  1. Pick one workflow that is narrow, internal, and repeats at least weekly. Narrow means one task with a clear input and output. Internal means members never see the raw AI output. Start internal and narrow: OWASP’s Top 10 for LLM Applications lists prompt injection and sensitive information disclosure among the top risks, which a narrow pilot keeps small by design.

  2. Write down how the workflow goes today, in your real numbers. List the steps, the owner, and the minutes. The UK government’s generative AI framework says to use the technology alongside your organisation’s policies, with the right assurance in place. Your baseline is the assurance part: without it, week two has nothing to compare.

  3. Write one goal sentence and one stop rule. The goal says what success looks like (“recap drafts need one editing pass instead of three”). The stop rule says when you quit (“if a draft invents a quote or date, we stop”). Write the stop rule now, while still neutral: the UK government’s generative AI framework opens with knowing what the technology is and what its limitations are. A stop rule is that knowledge written down before enthusiasm sets in.

  4. Run the workflow with AI help on real tasks, with a person reading every output before it goes anywhere. Use real work, not a demo; real tasks surface the problems controlled tests miss. The UK government’s framework calls for meaningful human control at the right stage, and for a pilot that stage is every output, every time.

  5. Log each run: what you asked, what came back, what you fixed. Keep it to five lines. OWASP’s Top 10 for LLM Applications lists misinformation, confident-sounding wrong content, among the top risks. The log is where you catch it.

Diagram of a two-week pilot plan: week one covers picking the workflow, recording the baseline, setting the goal and stop rule, running the task with a human check, and logging each run; week two covers comparing against the baseline, counting the costs, deciding keep change or stop, reporting what was learned, and filing the log.

Week two: measure and decide

  1. Compare the two weeks against your baseline: minutes and editing passes, before and after, side by side. Write down what improved, what declined, and what stayed flat. Flat numbers are a finding, not a failure.

  2. Count the costs, including the annoying ones: fixing outputs, reviewing, writing better prompts. A workflow that saves twenty minutes and creates thirty of cleanup is not a win.

  3. Make the call: keep it, change it, or stop it. Check the result against the goal and stop rule from step 3. “Change it” counts: narrower input, a better prompt, a different human-check step. If the answer is stop, write one sentence saying why.

  4. Tell the team what you learned, in one page or less. Share the baseline, result, and decision, and say where AI helped. University of Arizona researchers found across 13 experiments with more than 5,000 participants that disclosing AI use can reduce trust, but trust falls furthest when use is discovered by someone else. An honest one-pager makes the next pilot easier to approve.

  5. File the log and the one-pager where the next pilot can find them. Your filed pilot makes the second experiment cheaper than the first, and each one builds staff confidence in using the tools responsibly.

A worked example: the webinar recap pilot

This is a teaching example, not a case study, showing the plan filled in for one workflow.

The workflow. Monthly, a program coordinator turns a webinar transcript into a draft recap email, a half-day job today.

The goal and the stop rule. Goal: “Recap drafts need at most one editing pass instead of three.” Stop rule: “If the draft invents a speaker quote or a date even once, we stop the pilot and report.”

The prompt used for every run. She pastes the transcript after this exact prompt:

Draft a 200-word recap email for members who missed our webinar. Use plain language a busy member can skim. List the three main takeaways as bullets. Do not invent speaker names, dates, or numbers. If the transcript does not state something, leave it out rather than guessing.

Example output (shortened). “Subject: Three takeaways from Tuesday’s volunteer webinar. Our panel covered (1) pairing new volunteers with a buddy for the first 90 days, (2) a one-page role description before recruitment, and (3) thank-you calls within a week.” Each run needed one fix: the draft buried the recording link.

The week-two comparison. Editing passes fell from three to one; time per recap fell from four hours to ninety minutes, including the human check. One run merged two speakers’ points but nothing was invented, so the stop rule never triggered. Decision: keep, and add a speaker-attribution check.

Check your result

Before you call the pilot done, confirm each of these in your log:

  • The baseline numbers are written down from before the pilot started.
  • Every run has a log entry with what was fixed.
  • The week-two numbers are compared against the baseline.
  • The costs you counted include fixing and review time.
  • The decision is written in one sentence with the reason.
  • The one-page report exists and is filed where others can find it.

Mistakes that sink a two-week pilot

Picking a workflow that is really five workflows. “Handle member email” is not one workflow; “draft first replies to dues-renewal questions” is. If step 1 took more than ten minutes, narrow it further.

Skipping the baseline. A baseline written from memory in week two flatters the pilot. Document your starting numbers up front, with the same discipline you will apply to the results.

Letting the pilot sprawl past its dates. Set a hard end date: open-ended pilots quietly become permanent habits. Step 8 happens on schedule even if the data feels thin.

Treating “stop” as failure. The UK framework’s principles include using the right tool for the job, and sometimes the right tool is no AI at all. A pilot that proves a workflow should stay human is a successful pilot.

Sources

Sources