A pilot results presentation should answer whether to expand, adjust, extend, or stop a limited trial. Show the original success criteria beside the observed results, then explain the trade-off. A promising headline metric does not cancel a missed quality or workload threshold.
Use a data storytelling structure that connects the finding to a decision: what the pilot tested, what happened, what the evidence cannot settle, and what should happen next. The example below turns a mixed result into a seven-slide review without calling the trial a success by default.

Start with the decision the pilot was meant to inform
“Pilot update” describes a meeting, not its purpose. A useful question is more specific: “Should we introduce the new request form to another department, repeat the trial after a change, or keep the current process?”
Keep the scope beside the question: who took part, which workflow changed, how long the pilot ran, and what was excluded. A result from one department may support a decision about the next test without supporting a company-wide rollout.
The GOV.UK Service Manual’s guidance on measuring benefits recommends agreeing metrics at the start of beta and comparing observed benefits with forecasts when deciding whether to continue. For a pilot presentation, that means bringing the original criteria into the room instead of choosing attractive measures after seeing the outcome.
If the criteria were not agreed in advance, say so. You can still describe what happened and propose how to judge the next trial. Do not retrospectively label a convenient cutoff as the original goal.
A worked example: more complete requests, with a support burden
Suppose an internal service team tests a revised request form. The baseline covers four weeks under the old process; the pilot covers two weeks in one department. The following figures and thresholds are fictional teaching inputs, not a customer case or a result achieved with Presenti.
| Measure | Old process: four weeks | New form: two weeks |
|---|---|---|
| Distinct requests submitted | 100 | 50 |
| Requests with complete information at first submission, requiring no clarification | 76 | 46 |
| Requests resolved by the observation cutoff | 80 | 45 |
| Resolved requests requiring correction of a recorded entry error | 4 | 3 |
| Additional support labor logged for the new form | Not measured on this basis | 120 minutes |
“Complete information at first submission” is determined from the submitted fields and the clarification record. It is not the same as resolved: a complete request can still wait for review, and an incomplete one can be resolved after clarification. A request is counted once, even if it produces several messages.
The pilot’s five unresolved requests remain in the submission count. Three are waiting for the requester and two for a reviewer. Excluding them would change the population and make the presentation harder to interpret.
Keep each denominator beside its result
First-submission completeness rises from 76 ÷ 100 = 76% to 46 ÷ 50 = 92%, an increase of 16 percentage points. “Up 16%” is not the same statement. Show the counts as well as the percentages so readers can see the size of each group.
The recorded-entry correction rate answers a different question: of the requests resolved by the cutoff, how many required a correction? It is 4 ÷ 80 = 5% under the old process and 3 ÷ 45 ≈ 6.7% during the pilot. Do not use 3 ÷ 50 for a measure explicitly defined among resolved requests.
There are only three recorded corrections in the pilot. Their nature and severity matter, and the five unresolved requests could still change the later picture. The current figures describe the observation cutoff; they do not establish a stable long-term defect rate.
The support figure is 120 ÷ 50 = 2.4 additional staff minutes per submitted request. It is labor devoted to helping people use the new form, not customer waiting time, processing time, or time saved. Because the baseline did not measure the same support activity, this table cannot establish a net labor saving or a before-and-after workload change.
Show the mixed result instead of averaging it away
In this example, the team agreed three criteria before the pilot. These thresholds belong to the fictional decision; they are not industry benchmarks.
| Agreed criterion | Pilot observation | Result at the cutoff |
|---|---|---|
| At least 90% complete at first submission | 46/50 = 92% | Met |
| No more than 5% of resolved requests need an entry-error correction | 3/45 ≈ 6.7% | Not met |
| No more than 2 additional support minutes per submitted request | 120/50 = 2.4 minutes | Not met |
A headline such as “New form passes pilot with 92% success” would conceal two missed criteria and turn information completeness into an overall success score. Better: “More requests arrive complete; correction and support targets remain unmet.”
Do not average the three measures into one traffic light. They describe different things, have different units, and were selected to protect against different problems. If leaders choose to proceed despite a missed criterion, record that decision and its conditions rather than rewriting the result as a pass.
Separate the observed improvement from its explanation
The form changed, but so did the observation window and possibly the people or types of requests. This was a convenience pilot without random assignment or a parallel comparison group. The 16-point difference does not by itself prove that the form caused the improvement.
Look for explanations that would change the next action. Were simpler requests overrepresented? Did staff provide more help? Did the trial miss a busy period? Were the same rules used to mark information complete in both groups? Those questions belong beside the result, not in a generic disclaimer at the end.
For the fictional trial, the team suspects that an ambiguous field contributes to correction work. That is a hypothesis. Inspect the three corrected records and the support notes before claiming a common cause. If the records show different problems, a single rewritten field may not address all of them.
Keep qualitative observations distinct from the counts. A participant’s comment can explain confusion, but it does not show how common the issue is. Use approved, attributable wording if you have it; do not invent a quote to make the slide feel human.
Make a recommendation with a bounded next step
Here, a reasonable proposal is to review the correction cases, clarify the disputed field where the evidence supports that change, and repeat a trial of similar scope before expanding. It is a proposal, not an approval or a forecast that the targets will be met.
The next brief should state the revised form version, participating department, observation period, support provision, metric definitions, and decision criteria. Keep the same measures if they still answer the question. If a definition changes, show why and avoid comparing the revised number with the old one as though nothing changed.
Choose an observation period and participation plan appropriate to the decision with the responsible analyst or service owner. “Fifty requests were enough last time” is not a justification for a statistical claim. This example supplies no sample-size calculation or significance test.
If the sponsor wants to expand immediately, specify what would be accepted: which departments, what support, which unresolved risks, and what observation would trigger a pause. A full rollout is a different commitment from inviting one additional team.
A seven-slide pilot results outline
- Recommendation. Repeat a bounded trial after reviewing the correction and support issues; approval is still requested.
- Pilot scope. Show the old and new process, participants, windows, cutoff, and comparison limitations.
- Original criteria. Present the three agreed thresholds before the results.
- Observed outcomes. Show 76/100 versus 46/50, 4/80 versus 3/45, and 120/50 staff minutes. Keep the measures separate.
- What remains unexplained. Show the unresolved requests, correction cases, support notes, and the suspected field problem without declaring a cause.
- Next trial or alternative. Specify what would change, what would stay comparable, and what decision the next observation should support.
- Decision and conditions. Record what the sponsor actually agrees, with an owner and follow-up date once confirmed.
Keep request-level data and calculations in supporting material. The main deck should show enough detail to explain the recommendation without requiring the audience to reconstruct every denominator during the meeting.
Prepare a drafting brief that preserves the mixed result
Once the service owner has reviewed the figures, use Presenti’s Paste Text input to organize the notes into a slide draft. Enter the definitions and limits with the numbers. Do not ask the tool to decide whether the pilot was successful.

Create a seven-slide pilot results presentation from this fictional internal-form trial. Baseline: four weeks, 100 distinct submissions, 76 complete at first submission, 80 resolved, 4 resolved requests requiring entry-error correction. Pilot: two weeks in one department, 50 submissions, 46 complete at first submission, 45 resolved, 3 resolved requests requiring correction, and 120 additional staff support minutes. Five requests remain unresolved: three await the requester and two a reviewer. Show 76% versus 92%, a 16-percentage-point difference; 5% versus approximately 6.7% corrections among resolved requests; and 2.4 support minutes per pilot submission. The example’s prior criteria are at least 90% completeness, no more than 5% corrections, and no more than 2 support minutes. Only completeness meets its threshold. This was not a randomized trial. Recommend reviewing the correction and support issues before a similar-scope repeat; keep the decision unapproved. Do not invent causality, savings, significance, or a successful rollout.
In the draft, check that each percentage retains its count, that unresolved requests have not disappeared, and that the support time has not become “time saved.” A useful pilot review can recommend continuing while saying clearly that the current trial has not met every condition for expansion.