Skills Plugins MCP Prompt Model 博客 我的中心
生活与工具 #data #database #ai

incident-runbook-templates

Create structured incident response runbooks with step-by-step procedures, escalation paths, and recovery actions. Use this skill when building a service outage runbook for a payment processing system; creating database incident procedures covering connection pool exhaustion, replication lag, and disk space alerts; onboarding new on-call engineers who need step-by-step recovery guides written for a 3 AM brain; or standardizing escalation matrices across multiple engineering teams.

DeepseekModel 官方收录技能 质量 优秀 · 90 v1.0.0

获取

https://deepseekmodel.com/api/download.php?id=wshobson-agents-plugins-incident-response-skills-incident-runbook-templates-skill-md&format=skill
下载 .skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用
.skill 文件中 system_prompt 字段的实际内容。
name incident-runbook-templates description Create structured incident response runbooks with step-by-step procedures, escalation paths, and recovery actions. Use this skill when building a service outage runbook for a payment processing system; creating database incident procedures covering connection pool exhaustion, replication lag, and disk space alerts; onboarding new on-call engineers who need step-by-step recovery guides written for a 3 AM brain; or standardizing escalation matrices across multiple engineering teams. Incident Runbook Templates Production-ready templates for incident response runbooks covering detection, triage, mitigation, resolution, and communication. When to Use This Skill Creating incident response procedures Building service-specific runbooks Establishing escalation paths Documenting recovery procedures Responding to active incidents Onboarding on-call engineers Core Concepts 1. Incident Severity Levels Severity Impact Response Time Example SEV1 Complete outage, data loss 15 min Production down SEV2 Major degradation 30 min Critical feature broken SEV3 Minor impact 2 hours Non-critical bug SEV4 Minimal impact Next business day Cosmetic issue 2. Runbook Structure 1. Overview & Impact 2. Detection & Alerts 3. Initial Triage 4. Mitigation Steps 5. Root Cause Investigation 6. Resolution Procedures 7. Verification & Rollback 8. Communication Templates 9. Escalation Matrix Detailed patterns and worked examples Detailed pattern documentation lives in references/details.md . Read that file when the navigation tier above is insufficient. Best Practices Do's Keep runbooks updated - Review after every incident Test runbooks regularly - Game days, chaos engineering Include rollback steps - Always have an escape hatch Document assumptions - What must be true for steps to work Link to dashboards - Quick access during stress Don'ts Don't assume knowledge - Write for 3 AM brain Don't skip verification - Confirm each step worked Don't forget communication - Keep stakeholders informed Don't work alone - Escalate early Don't skip postmortems - Learn from every incident Troubleshooting Runbook steps work in staging but fail during a real incident Steps often assume preconditions that are true in a healthy environment but not during an outage. For each command in your runbook, add a prerequisite check and a "what to do if this command fails" note: # Step: Check pod status kubectl get pods -n payments # Prerequisites: kubectl configured, kubeconfig points to correct cluster # If this fails: run `aws eks update-kubeconfig --name prod-cluster --region us-east-1` # Expected output: pods in Running state On-call engineer panics and skips steps out of order Add a numbered checklist at the top of the runbook that mirrors the section numbers, so responders can track progress under stress without reading the full document: ## Quick Checklist - [ ] 1. Declare incident severity and open war room - [ ] 2. Check service health (Section 4.1) - [ ] 3. Check recent deployments (Section 4.1) - [ ] 4. Roll back if deploy is suspect (Section 4.1) - [ ] 5. Post initial notification to #payments-incidents - [ ] 6. Escalate if > 15 min unresolved Runbook is outdated — commands reference old cluster names or endpoints Runbooks rot because they're updated manually. Include a "Last Verified" date and owner at the top, and add a CI check that validates all curl endpoints and kubectl context names are still valid: ## Runbook Metadata | Field | Value | |---|---| | Last verified | 2024-11-15 | | Owner | @platform-team | | Review cadence | After every SEV1/SEV2 | Stakeholder communication is delayed while engineers are heads-down Assign a dedicated incident communicator role (separate from the incident commander) whose only job is to post status updates. Add a standing agenda in the communication template: Update every 15 minutes (even if no new information): - Current status (Investigating / Mitigating / Monitoring) - Impact (what is broken, who is affected, % of traffic) - What we are doing right now - Next update in: 15 minutes Database runbook commands cause additional downtime when run incorrectly Add explicit warnings before destructive SQL commands and require a dry-run output check before executing: -- WARNING: This terminates active connections. Verify count first. -- DRY RUN (check count before terminating): SELECT count ( * ) FROM pg_stat_activity WHERE state = 'idle' AND query_start < now() - interval '10 minutes' ; -- EXECUTE only after verifying count is reasonable (< 50): SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE state = 'idle' AND query_start < now() - interval '10 minutes' ; Related Skills postmortem-writing - After resolving an incident, use postmortem templates to capture root cause and preventive actions on-call-handoff-patterns - Structure shift handoffs so the incoming responder has full context on active incidents
Agent 识别该技能的关键词,点击任意一个即可复制。

该技能未提供触发词。

下载的 .skill 包内含以下字段。
字段 说明
format格式标识(skill/v1)
skill_id技能唯一 ID
name技能名称
version版本号
description技能描述
category所属分类(数组)
trigger_words触发词列表
tags标签列表
source来源标识
source_url来源链接(本页地址)
exported_at导出时间(每次下载生成)
system_prompt系统提示词正文
model_config模型参数:provider / model / temperature / max_tokens / top_p
examples示例
install_guide各平台导入说明(Coze / Dify / Claude / 自定义框架)
同一份技能可按不同平台格式导出。
.skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用 下载
.skillpro 增强格式,额外含脚本 / 工具 / 依赖 / 钩子占位 下载
.json 纯 JSON 导出,只含 system_prompt 与模型参数 下载
Coze 带 frontmatter 的 Markdown,Coze 平台导入用 下载
Dify Dify DSL,创建应用后直接导入 下载

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。