Desk Trials

GUIDE / Evidence

Run a workflow pilot that can answer a real question

A small scorecard can compare baseline effort, output quality, and rework without pretending a handful of jobs is a controlled study.

Documentation-based practical guide

Documentation-based proposal; not a hands-on test

Start with the decision, not the tool

A useful pilot answers a bounded decision: keep the current process, revise the proposed workflow, or stop. Write that decision above the scorecard. Then choose work that resembles normal use, including at least a few awkward inputs. NIST's AI Risk Management Framework says performance should be evaluated in conditions similar to deployment and that methods, metrics, uncertainty, and limitations should be documented. The UK Government Service Manual makes the companion point: establish a baseline before claiming improvement.

This is not a request for a laboratory experiment. A freelancer or small team may have only ten suitable jobs in a month. That can still reveal broken handoffs, review burden, and obvious regressions. It cannot establish a general productivity percentage. Label the result a small operational pilot and keep the denominator visible.

Use three columns that resist demo logic

MeasureRecordWhy it matters
Baseline effortElapsed working time and waiting time for the old processPrevents a fast-looking step from hiding setup or queue time
QualityWhether the output met a written acceptance checkKeeps completion separate from correctness
ReworkMinutes spent checking, correcting, rerunning, or explainingCaptures labor moved downstream

Use the same task boundary before and during the pilot. If baseline timing ends when a draft exists, the pilot timing must not end only after approval. Record obvious context such as task type, unusual source quality, and who performed the work. Do not combine unlike jobs into one precise-looking average.

A deliberately small example

Suppose, hypothetically, a two-person studio pilots a new brief-to-draft workflow on twelve ordinary briefs. The scorecard records old-process observations from six recent comparable briefs and the twelve pilot runs. A draft passes only if required sections are present, source links open, and the reviewer finds no unsupported factual claim. The team records active minutes, waiting time, pass or fail, and correction minutes.

If eight pilot drafts pass and four need substantial repair, the honest result is not “33% failure” as a universal rate. It is: four of twelve pilot items in this task mix crossed the team's repair threshold, with the failure notes attached. That may be enough to revise the workflow before wider use. It is not evidence about another team, a future version, or all writing tasks.

Close with a decision note

  1. State what changed, what stayed constant, and which jobs were included.
  2. Show baseline and pilot rows, including missing observations.
  3. Summarize recurring failure modes rather than only the average.
  4. Choose keep, revise, or stop, and name the next review date.

Keep perceived ease as a separate note. People may prefer a workflow that takes longer, or dislike one that reduces rework. Both observations matter, but they are different. For a related caution about study boundaries, see what the developer slowdown study actually measured. A pilot scorecard earns trust by staying small enough to inspect and modest enough to describe accurately.

Before you hand it over

Use this as a working check, not certification. Checks stay in this page only and reset on reload.

0 of 5 checked

Sources & verification

Product details are based on the linked documentation. The proposed workflow and worked examples are editorial guidance, not measured test results.

  1. AI RMF CoreSource date: 2023-01-26 · Retrieved: 2026-09-19T19:39:00Z

    Use of documented metrics, comparisons, uncertainty, deployment-like conditions, and explicit limits on generalization.

  2. Measuring the benefits of your serviceSource date: 2018-06-20 · Retrieved: 2026-09-19T19:39:00Z

    Establishing a baseline, testing assumptions, tracking costs and benefits, and deciding whether to continue.

Continue the workflow

  1. What the developer slowdown study actually measured

    A surprising result matters most when its task and participants stay attached.

  2. A revised experiment is not a retraction

    METR's follow-up changed the study design while preserving the earlier finding.

  3. A 1993 paper named the computer productivity paradox

    Brynjolfsson's paper defines the productivity paradox and tests four explanations for it, not a resolved verdict.