Choosing useful monitoring and alerts
Monitoring is useful when an alert identifies an action someone can take. A running server does not necessarily mean that logins, payments or other important application functions work. Agree monitoring coverage and ownership as part of the service scope.
Select checks around the workload
- Infrastructure: capacity, storage health, network reachability and resource trends.
- Application: the public endpoint and a safe representative health check that does not create real transactions.
- Recovery: backup success, age of the latest usable copy and certificate renewal status where included.
- Operations: alert recipients, escalation path, coverage hours and suppression during approved maintenance.
Record what each alert means and the first investigation steps. Thresholds should reflect the application and its normal patterns. A monitoring check is not an availability guarantee, and monitoring outside the agreed scope should be requested explicitly.