Improve Incident Response
12 weeks · 4 milestones
Milestone map
Milestone map
3 milestones
Document the current incident response process — how incidents are detected, triaged, escalated, resolved, and reviewed — and establish baseline metrics covering: mean time to detect (MTTD), mean time to respond (MTTR), and the proportion of incidents that produce a completed post-mortem or blameless retrospective. Review at least five historical incidents to understand where the process breaks down in practice. Produce an audit report identifying at least three specific failure modes in the current process.
Proof required
Submit an incident response audit report of at least 500 words covering: current process documentation (detection, triage, escalation, resolution, review steps); baseline metrics for MTTD and MTTR (measured or estimated from historical incident records); post-mortem completion rate; and at least three specific named failure modes identified from at least five historical incidents. ADVERSARIAL VERIFICATION: a named Site Reliability Engineer (SRE), DevOps engineer, or infrastructure engineering lead (5+ years experience) must read the audit and provide written confirmation that the failure modes are substantive process breakdowns, not surface-level observations.
What gets checked
- 500-word audit with documented process, baseline MTTD/MTTR, post-mortem rate, and ≥3 failure modes from ≥5 historical incidents
- Failure modes are specific to your process — not generic statements that apply to any organisation
- Written confirmation from a named SRE/DevOps/infrastructure lead (5+ years) that the failure modes are substantive