Practice guide

Measure what improved

Compare accepted delivery and quality using the same definitions before and during the program.

Agree on the comparison before the pilot

Use this guide in Prepare to establish a baseline, then revisit it in Scale and Outcomes. Bring ticket acceptance records, delivery timestamps, review history, and defect records. Name a measurement owner and use sources the team can inspect.

Choose comparable work and equal observation intervals. Record team size, available working days, ticket mix, and changes in scope or staffing. If those differ materially, explain the difference or narrow the comparison. When historical evidence is missing, start collecting it before claiming improvement.

Use four consistent definitions

  1. Accepted work: count work that meets the same acceptance standard during each interval. Keep work categories visible; splitting tickets more finely can increase the count without increasing delivered value.
  2. Cycle time: choose an explicit start and end event, such as work started to accepted. Compare the same statistic, such as the median, in both periods and include the number of completed items.
  3. Rework: define what counts, such as review rounds required to correct behavior. Use the same collection method and report the total and sample size.
  4. Escaped defects: define defects found after acceptance, their severity, and the follow-up window. Give each period the same time to reveal defects; a newly completed interval has had less exposure.

Review these together. Faster completion with rising rework or defects needs investigation before expansion. Keep links to the underlying records so the pod and sponsor can challenge the interpretation.

Worked comparison

Consider two equal, four-week periods for the same pod and comparable work. The baseline contains 20 accepted items; the current period contains 24. The count increase is (24 − 20) ÷ 20 = 20%. A median cycle time moving from five days to four is a 20% reduction.

If behavior-correction review rounds fall from 12 across 20 items to 10 across 24, the average changes from 0.60 to about 0.42 per item. Check whether review practice or recording changed. Compare escaped defects and severity after the same follow-up window before concluding quality held.

These figures show how to calculate a comparison. Use your own records to assess results and examine other explanations, including ticket mix, staffing, or an unusually difficult baseline.

Make an expansion decision

Bring the before-and-after scorecard, source records, limitations, and receiving-pod feedback to the sponsor and coach. Expand when comparable delivery evidence supports improvement, quality holds, and another pod can use the playbook. Otherwise assign a repair or further measurement step.

The 50-engineer planning example assumes 30% improvement per trained engineer and cumulative training of 5, 30, then 50 engineers. Those inputs are planning assumptions. Record your observed results separately and use them to revise your rollout plan.

All practice guides →