問題文
The team is deciding what to alert on for the order API. Following the Prometheus project's alerting guidance, which approach should they take?
選択肢
- Alert on every possible cause of failure so the on-call engineer is told exactly which component broke before looking at any dashboard.
- Alert on symptoms associated with user pain, keeping the number of alerts as small as possible, and rely on consoles to pinpoint the cause.
- Alert on every threshold that resource dashboards display, because resource saturation always precedes user-visible failure.
- Alert on latency at every layer of the stack so that the slowest layer is always identified automatically.