- Added a new endpoint `/api/agg/incidents` to provide a normalized snapshot of incidents for fleet triage, including severity and recommended actions. - Implemented live event streaming via `/api/live/events` for real-time updates on fleet status and incidents, enhancing observability. - Updated the Web UI to include dedicated sections for incidents and live updates, improving user navigation and access to critical information. - Enhanced API documentation to reflect new endpoints and their functionalities, ensuring clarity for developers and users.
2.0 KiB
2.0 KiB
Incident Rollout Playbook
Scope
This playbook describes staged rollout for:
/api/agg/incidents/api/live/events(SSE snapshots)- UI pages
/incidentsand/live - Live polling controls and stale indicators on key pages
Stage 0 - Baseline Capture (2-3 days)
- Record current MTTD and MTTR from on-call logs.
- Record manual refresh usage on main pages.
- Save top recurring failure patterns (degraded nodes, read-only modes, bad connections spikes).
Outputs:
- baseline MTTD / MTTR
- top 5 incident categories by frequency
Stage 1 - Shadow Mode (3-5 days)
- Enable incidents and live pages for operators.
- Do not change paging/escalation yet.
- Compare incident feed against existing monitoring and mark false positives.
Targets:
- false positive ratio < 20%
- no increase in upstream load beyond acceptable budget
Stage 2 - Assisted Triage (1 week)
- Use
/incidentsas primary triage board. - Require owner + ack for active critical incidents.
- Use runbook links from incident items.
Targets:
- ack coverage for critical incidents >= 90%
- owner coverage for critical incidents >= 90%
Stage 3 - Policy Tuning (ongoing)
- Adjust thresholds:
bad_connections_warn(default 1000)bad_connections_high(default 10000)
- Review alert fatigue weekly.
- Promote stable thresholds into documented policy.
KPI Tracking
- MTTD (minutes): incident first observed -> first ack
- MTTR (minutes): incident first observed -> resolved
- Stale time share: percentage of time live views are stale
- Manual refresh share: manual refresh / total data update actions
Fast Rollback
If noise or load is excessive:
- disable auto-refresh by setting
refresh=0in shared ops links - switch operators back to dashboard summary only
- keep
/api/agg/incidentsfor diagnostics while disabling SSE consumers
Weekly Review Template
- KPI delta (MTTD, MTTR) vs baseline
- top noisy rules
- incidents with missing owner/ack
- policy changes applied this week