Server health monitoring
Most server monitoring watches CPU, memory and disk, which are the three things least likely to be the actual problem. The signals that predict outages are duller and are usually not collected at all.
Engineering write-ups on it support from deployments KYCONNECTS has run, with the reasoning behind each approach.
Page 2 of 2
Most server monitoring watches CPU, memory and disk, which are the three things least likely to be the actual problem. The signals that predict outages are duller and are usually not collected at all.
Most dashboards are built to show everything, which is why nobody looks at them. A dashboard is a diagnostic tool with one question to answer, and the discipline that makes it useful is deciding what to leave off.
Most monitoring tells you a server is up while customers cannot use it. The difference between a dashboard and a control is whether it watches what users experience or what machines report.
We typically respond within 4–8 business hours.