Pular para conteúdo

Runbooks

Procedimentos operacionais associados aos alertas Prometheus. Todo alerta referencia uma destas páginas na annotation runbook.

Runbook Alertas
host-down HostDown
cpu HostCPUSaturation
memoria-swap HostMemoryExhaustion, OOMKillerInvoked, HostSwapPressure
disco HostDiskCritical, HostDiskWarning, HostInodesExhausted, FilesystemReadonly
containers CriticalContainerCrashLoop, CriticalContainerDown, ContainerOOMKilled, ContainerCPUThrottlingSevere
telemetria AlloyAgentDown
postgres PostgreSQLDown, PostgreSQLConnectionsExhausted, PostgreSQLDeadlocks
redis RedisDown, RedisMemoryExhausted
http-erros HTTPTotalFailure, HTTP5xxSustained
latencia HighLatencyP95
error-budget SLOFastBurn, SLOSlowBurn, SLOErrorBudget*

Convenções:

  • Acesso às VMs: SSH conforme inventário em ansible/inventory/.
  • Dashboards citados estão no Grafana da stack de operações.
  • Após resolver, registrar causa e ação no canal de incidentes.