# Incident Rollout Playbook ## Scope This playbook describes staged rollout for: - `/api/agg/incidents` - `/api/live/events` (SSE snapshots) - UI pages `/incidents` and `/live` - Live polling controls and stale indicators on key pages ## Stage 0 - Baseline Capture (2-3 days) - Record current MTTD and MTTR from on-call logs. - Record manual refresh usage on main pages. - Save top recurring failure patterns (degraded nodes, read-only modes, bad connections spikes). Outputs: - baseline MTTD / MTTR - top 5 incident categories by frequency ## Stage 1 - Shadow Mode (3-5 days) - Enable incidents and live pages for operators. - Do not change paging/escalation yet. - Compare incident feed against existing monitoring and mark false positives. Targets: - false positive ratio < 20% - no increase in upstream load beyond acceptable budget ## Stage 2 - Assisted Triage (1 week) - Use `/incidents` as primary triage board. - Require owner + ack for active critical incidents. - Use runbook links from incident items. Targets: - ack coverage for critical incidents >= 90% - owner coverage for critical incidents >= 90% ## Stage 3 - Policy Tuning (ongoing) - Adjust thresholds: - `bad_connections_warn` (default 1000) - `bad_connections_high` (default 10000) - Review alert fatigue weekly. - Promote stable thresholds into documented policy. ## KPI Tracking - MTTD (minutes): incident first observed -> first ack - MTTR (minutes): incident first observed -> resolved - Stale time share: percentage of time live views are stale - Manual refresh share: manual refresh / total data update actions ## Fast Rollback If noise or load is excessive: 1. disable auto-refresh by setting `refresh=0` in shared ops links 2. switch operators back to dashboard summary only 3. keep `/api/agg/incidents` for diagnostics while disabling SSE consumers ## Weekly Review Template - KPI delta (MTTD, MTTR) vs baseline - top noisy rules - incidents with missing owner/ack - policy changes applied this week