Documentation-based proposal; not a hands-on test
Start with the decision, not the tool
A useful pilot answers a bounded decision: keep the current process, revise the proposed workflow, or stop. Write that decision above the scorecard. Then choose work that resembles normal use, including at least a few awkward inputs. NIST's AI Risk Management Framework says performance should be evaluated in conditions similar to deployment and that methods, metrics, uncertainty, and limitations should be documented. The UK Government Service Manual makes the companion point: establish a baseline before claiming improvement.
This is not a request for a laboratory experiment. A freelancer or small team may have only ten suitable jobs in a month. That can still reveal broken handoffs, review burden, and obvious regressions. It cannot establish a general productivity percentage. Label the result a small operational pilot and keep the denominator visible.
Use three columns that resist demo logic
| Measure | Record | Why it matters |
|---|---|---|
| Baseline effort | Elapsed working time and waiting time for the old process | Prevents a fast-looking step from hiding setup or queue time |
| Quality | Whether the output met a written acceptance check | Keeps completion separate from correctness |
| Rework | Minutes spent checking, correcting, rerunning, or explaining | Captures labor moved downstream |
Use the same task boundary before and during the pilot. If baseline timing ends when a draft exists, the pilot timing must not end only after approval. Record obvious context such as task type, unusual source quality, and who performed the work. Do not combine unlike jobs into one precise-looking average.
A deliberately small example
Suppose, hypothetically, a two-person studio pilots a new brief-to-draft workflow on twelve ordinary briefs. The scorecard records old-process observations from six recent comparable briefs and the twelve pilot runs. A draft passes only if required sections are present, source links open, and the reviewer finds no unsupported factual claim. The team records active minutes, waiting time, pass or fail, and correction minutes.
If eight pilot drafts pass and four need substantial repair, the honest result is not “33% failure” as a universal rate. It is: four of twelve pilot items in this task mix crossed the team's repair threshold, with the failure notes attached. That may be enough to revise the workflow before wider use. It is not evidence about another team, a future version, or all writing tasks.
Close with a decision note
- State what changed, what stayed constant, and which jobs were included.
- Show baseline and pilot rows, including missing observations.
- Summarize recurring failure modes rather than only the average.
- Choose keep, revise, or stop, and name the next review date.
Keep perceived ease as a separate note. People may prefer a workflow that takes longer, or dislike one that reduces rework. Both observations matter, but they are different. For a related caution about study boundaries, see what the developer slowdown study actually measured. A pilot scorecard earns trust by staying small enough to inspect and modest enough to describe accurately.
Before you hand it over
Use this as a working check, not certification. Checks stay in this page only and reset on reload.
0 of 5 checked
Sources & verification
Product details are based on the linked documentation. The proposed workflow and worked examples are editorial guidance, not measured test results.
- AI RMF CoreSource date: 2023-01-26 · Retrieved: 2026-09-19T19:39:00Z
Use of documented metrics, comparisons, uncertainty, deployment-like conditions, and explicit limits on generalization.
- Measuring the benefits of your serviceSource date: 2018-06-20 · Retrieved: 2026-09-19T19:39:00Z
Establishing a baseline, testing assumptions, tracking costs and benefits, and deciding whether to continue.
Continue the workflow
- What the developer slowdown study actually measured
A surprising result matters most when its task and participants stay attached.
- A revised experiment is not a retraction
METR's follow-up changed the study design while preserving the earlier finding.
- A 1993 paper named the computer productivity paradox
Brynjolfsson's paper defines the productivity paradox and tests four explanations for it, not a resolved verdict.