All outcomes
Teams

Lead Incidents From Alert to Resolution

12 weeks · 4 milestones

Milestone map

Milestone map

3 milestones

Audit existing incident response process and baseline metrics

3 weeks

Document the current incident response process — how incidents are detected, triaged, escalated, resolved, and reviewed — and establish baseline metrics covering: mean time to detect (MTTD), mean time to respond (MTTR), and the proportion of incidents that produce a completed post-mortem or blameless retrospective. Review at least five historical incidents to understand where the process breaks down in practice. Produce an audit report identifying at least three specific failure modes in the current process.

Proof required

Submit an incident response audit report of at least 500 words covering: current process documentation (detection, triage, escalation, resolution, review steps); baseline metrics for MTTD and MTTR (measured or estimated from historical incident records); post-mortem completion rate; and at least three specific named failure modes identified from at least five historical incidents. ADVERSARIAL VERIFICATION: a named Site Reliability Engineer (SRE), DevOps engineer, or infrastructure engineering lead (5+ years experience) must read the audit and provide written confirmation that the failure modes are substantive process breakdowns, not surface-level observations.

What gets checked

  • 500-word audit with documented process, baseline MTTD/MTTR, post-mortem rate, and ≥3 failure modes from ≥5 historical incidents
  • Failure modes are specific to your process — not generic statements that apply to any organisation
  • Written confirmation from a named SRE/DevOps/infrastructure lead (5+ years) that the failure modes are substantive

Common mistakes

  • Using only documentation to understand the process rather than reviewing real incident records — documented incident response processes routinely diverge from what actually happens during incidents; the audit must review real incidents
  • Identifying generic industry failure modes ('we lack observability') without connecting them to specific gaps in your incidents — each failure mode must be evidenced by at least one named incident type or pattern
  • Not measuring MTTR because exact data is unavailable — approximate values derived from incident records or engineering team memory are acceptable; absence of any measurement leaves the baseline undefined

Resources

Foundationstart here

Depthgo deeper

What a verifier looks for

  • The audit must reference real incidents — ask the submitter to name at least two of the five incidents reviewed (approximate dates and incident type are sufficient, not PII)
  • Failure modes must be mechanistic — each failure mode should identify the specific step in the process that breaks and why; 'our on-call response is slow' is a symptom, 'on-call engineers are not paged until a human notices the alert, averaging 22 minutes post-incident start' is a failure mode
  • The confirming reviewer must have direct SRE or incident response experience — a product manager or general engineering manager without infrastructure responsibility does not satisfy the 5+ years SRE/DevOps requirement

Implement improved incident response components and measure results

6 weeks

Design and implement improvements to at least two of the three incident response failure modes identified in M1. Improvements may include: adding or improving alerting and detection (monitoring, SLO/SLA alert thresholds, PagerDuty configuration), implementing a structured incident commander role, creating or improving the post-mortem template and completion process, running incident response training or a game day, or implementing an on-call rotation improvement. Measure the impact after at least 4 weeks of operation using the same MTTD/MTTR metrics from the baseline.

Proof required

Submit an incident response improvement report covering: (1) two improvements implemented with their target failure modes; (2) evidence of implementation (alerting configuration screenshots, updated runbook, post-mortem template, game day report, or equivalent); (3) updated MTTD/MTTR measurements from at least 4 weeks of operation after the improvements, compared to the M1 baseline. ADVERSARIAL VERIFICATION: your named SRE/DevOps reviewer reads the improvement design before implementation and provides written feedback on whether the improvements are likely to move the MTTD/MTTR metrics, or whether they primarily improve documentation without addressing the detection/response mechanism.

What gets checked

  • Two improvements implemented with implementation evidence — not just designed on paper
  • MTTD/MTTR measurements from ≥4 weeks post-improvement compared to M1 baseline
  • Reviewer feedback on the improvement design before implementation confirming the changes target the detection/response mechanism

Common mistakes

  • Improving documentation (runbooks, playbooks) without improving the detection or escalation mechanism — better documentation helps responders once paged; it does not reduce MTTD, which is the detection problem
  • Not measuring after 4 weeks because no major incidents occurred — MTTD and MTTR can be measured on any incident type; document your measurements even if incident count is low
  • Measuring MTTR only on resolved incidents and excluding incidents that required multiple responders or cross-team escalation — complex incidents are where MTTR failures are most visible and most important to measure

Resources

Foundationstart here

Depthgo deeper

What a verifier looks for

  • Implementation evidence must be specific — a screenshot of a new PagerDuty alert rule, a link to the updated runbook in a wiki, or the completed post-mortem template are all acceptable; 'we updated our process' is not
  • Post-improvement MTTD/MTTR must be compared to the M1 baseline — if no change is observable, that is a valid result that should prompt discussion about why
  • The reviewer's pre-implementation feedback is the core check — the question is whether the improvements target MTTD (detection) or MTTR (response), not both simultaneously

Present incident response improvements with SRE practitioner challenge

2 weeks

Compile an incident response improvement report covering the M1 audit, M2 results, and a reliability roadmap recommendation covering the next improvement cycle. Present the report in a live session of at least 30 minutes to a panel that includes your named SRE/DevOps reviewer plus at least one additional SRE, infrastructure lead, or engineering manager. The panel poses adversarial challenges on your metric interpretation, causal claims, and the reliability of the measurement approach.

Proof required

Submit your incident response improvement report and a record of the challenge session: names and professional roles of both panellists, date and duration (≥30 minutes), and notes on at least three challenges posed and your responses. A written statement from the primary SRE reviewer (50+ words) confirming that the improvements represent genuine changes to detection or response capability. ADVERSARIAL VERIFICATION RULE Level 2: live session required.

What gets checked

  • Incident response improvement report: M1 audit + M2 results + reliability roadmap recommendations
  • Challenge session record with ≥2 named panellists (both with SRE/infrastructure experience), duration ≥30 minutes, ≥3 challenges documented
  • Written statement from primary SRE reviewer (50+ words) confirming genuine capability improvements

Common mistakes

  • Presenting MTTD/MTTR improvements without the underlying incident data — good SRE panellists will ask to see sample incident records to validate the metric calculations
  • Attributing MTTD/MTTR changes to your improvements when other confounding changes occurred in the same period (infrastructure migration, team growth, technology change) — identify and acknowledge confounders explicitly
  • Choosing a panel without direct SRE or reliability engineering experience — a product manager and a general engineering manager cannot provide the adversarial challenge that an SRE or infrastructure lead can

Resources

Foundationstart here

Depthgo deeper

What a verifier looks for

  • Both panellists must have direct SRE or infrastructure engineering experience — check professional roles before accepting the session record
  • ADVERSARIAL VERIFICATION RULE Level 2 requires a live synchronous session — async written questions do not qualify
  • The reviewer statement must confirm capability changes — 'the team now detects incidents faster because…' or 'the post-mortem process now produces…' are the types of claims that satisfy the requirement

We use analytics to improve Powstik. No ads, ever.