Milestone map
Milestone map
3 milestones
Baseline Sprint Metrics and Diagnose Root Cause
2–4 weeks for baseline data collection and diagnosis (or immediate if historical data exists)
You cannot improve what you have not measured. Before any intervention, you need a reliable 4–6 sprint baseline: sprint completion rate (stories or points committed vs. delivered), cycle time, and the distribution of incomplete work by cause. The baseline must draw from real historical data, and the root cause diagnosis must be specific — not 'we need to improve planning' but '60% of missed commitments over 6 sprints were caused by scope additions in the last 3 days from a single stakeholder.'
Proof required
Submit: (1) a baseline metrics report covering 4–6 sprints: sprint completion rate per sprint and the average, cycle time or equivalent flow metric, and the distribution of incomplete work by cause (blocked, scope-added, underestimated, deprioritised); (2) a root cause diagnosis (200 words minimum) identifying the 1–2 specific causes that account for the majority of missed commitments, with supporting data; (3) pre-committed improvement targets for each planned intervention — written before the intervention begins, not after seeing results.
What gets checked
- Baseline draws from actual sprint data — '60% completion on average' requires the per-sprint data showing how the average was calculated; a single 'typical sprint' estimate is not a baseline
- Root cause diagnosis names specific causes with proportional data — '80% of missed commitments came from these three causes' with each quantified is a diagnosis; 'we need better planning' is not
- Pre-committed targets are specific and measurable — 'improve completion rate from 60% to 75% within 8 sprints of implementing the change' is a target; 'improve delivery' is not
Common mistakes
- Diagnosing root cause without data — a list of likely causes that matches prior beliefs, without sprint-level data, is a hypothesis not a diagnosis
- Choosing the easiest metric to show improvement rather than the most meaningful — if the real problem is missed commitments, cycle time improvement that leaves completion rate unchanged does not address the stated problem
- Setting targets after seeing post-intervention results — the target must be documented before the intervention starts
Resources
Foundationstart here
Depthgo deeper
Masteryfor the dedicated
What a verifier looks for
- Baseline data quality: ask the submitter to show the per-sprint data — a single average figure has no variance information and is not a real baseline
- Root cause specificity: ask the submitter to name the top cause of missed commitments and the percentage it accounts for — a diagnosis that cannot answer this is too vague
- Target pre-commitment: ask when the improvement target was written down — if it was after the first post-intervention sprint results, it was not pre-committed
Implement Specific Intervention and Collect Post-Intervention Data
6–10 weeks for 4–6 post-intervention sprints
The intervention is the specific change made to address the M1 root cause. It must be specific enough that you can explain exactly what was different in the sprints after the intervention versus before. Vague interventions produce uninterpretable results: 'we improved our planning process' does not tell you whether the improvement was due to a shorter planning session, a new estimation technique, a change in how backlog items are written, or a stakeholder communication boundary. Four to six post-intervention sprints are required to distinguish signal from noise.
Proof required
Submit: (1) a pre-intervention plan written before implementation: the specific change (process, cadence, tooling, or stakeholder boundary), why it addresses the M1 root cause, and the sprint it went live; (2) post-intervention sprint data for 4–6 sprints using the same metrics from M1; (3) a brief notes log for each post-intervention sprint capturing any confounding events — team changes, holidays, major incidents — that might explain metric changes independently of the intervention.
What gets checked
- Intervention is specific enough to replicate — a reader not present should be able to implement it from the description; 'we improved estimation' is not replicable; 'we introduced three-point estimation for all stories above 3 points in the last two days before sprint planning' is
- Post-intervention data uses the same metric definitions as M1 baseline — if the definition changed, document why
- Confounding events log is honest — a 6–10 week period with zero confounding events is unusual; team changes, holidays, and incidents should be noted
Common mistakes
- Implementing multiple changes simultaneously and attributing improvement to one cause — if estimation and stakeholder boundaries both changed in the same sprint, neither can be isolated
- Switching to a different metric after the intervention produces disappointing results — the M2 data must use the M1 metric
- Running only 2–3 post-intervention sprints — too much variance to distinguish real effect from random fluctuation; 4–6 sprints is the minimum
Resources
Foundationstart here
Depthgo deeper
What a verifier looks for
- Intervention specificity: ask what an engineer experienced differently in sprints after the intervention — if the answer is vague, the intervention was not specific enough to produce interpretable results
- Metric consistency: compare metric definitions in M1 and M2 — if they differ without explanation, ask why
- Confounding events: ask whether anything unusual happened in the post-intervention sprints; a 6–10 week period with no team changes or holidays is unusual
Report Improvement Against Pre-Committed Targets with Expert Review
1–2 weeks for analysis, lessons-learned, and expert review session
The final milestone compares M1 pre-committed targets with M2 actual results and requires a real-time review with a named, qualified engineering leader or delivery coach who can challenge the analysis, probe the intervention logic, and stress-test the attribution. The ADVERSARIAL VERIFICATION Level 2 standard applies: the reviewer must ask undisclosed questions about specific decisions in the baseline, intervention design, and attribution, and the submitter must reason in real time.
Proof required
Submit: (1) a before/after comparison table: each M1 pre-committed target, the baseline value, the post-intervention value, and the delta; plus a 200-word attribution analysis — what you attribute to the specific intervention vs. other factors; (2) a lessons-learned document (300 words minimum): what you would do differently in the intervention design, what the root cause analysis got right or wrong, and whether the improvement is likely to persist; (3) documentation of a real-time Q&A session with a named reviewer who has ≥3 years of engineering delivery leadership experience and a verifiable professional profile — the documentation must include specific challenge questions and your responses; the reviewer must not be your current manager or a direct report.
What gets checked
- Before/after table uses M1 pre-committed targets — the table must include the target value; if the target was not met, explain why rather than celebrating a different metric that improved
- Attribution analysis acknowledges plausible alternative explanations — a delivery improvement coinciding with two new team members cannot be fully attributed to the intervention without addressing that confound
- Q&A documentation shows the reviewer asked attribution questions — look for questions like 'why do you attribute this to your intervention rather than the team composition change in sprint 3?'
Common mistakes
- Reporting improvement against a target added after the intervention results — check that the M3 comparison table target matches the one documented at the end of M1
- Attribution analysis that only credits the intervention — every engineering team improvement has confounding variables; the analysis must name the two or three most plausible alternatives and explain the attribution
- A reviewer who validates the result without challenging methodology — 'yes this is a good improvement' without challenging attribution, target-setting, or intervention isolation does not meet the adversarial standard
Resources
Foundationstart here
Depthgo deeper
What a verifier looks for
- Target fidelity: compare M3 targets with M1 documented targets — if they differ without explanation, ask when the targets were set
- Attribution honesty: ask the submitter to name the most plausible alternative explanation for the improvement that is NOT their intervention — inability to name one means the attribution analysis is not credible
- Reviewer adversarial quality: look for methodology challenge questions about attribution and intervention isolation, not only process questions