Three pillars of observability
- Metrics — numeric measurements over time: RPS, p99 latency, error rate, CPU. Prometheus + Grafana is the standard
- Logs — structured event records. Centralised collection: ELK Stack (Elasticsearch, Logstash, Kibana), Loki
- Traces — tracking a request's path through multiple services. Jaeger, Zipkin, OpenTelemetry
The Four Golden Signals (Google SRE)
Latency, Traffic, Errors, Saturation — four metrics that describe the health of any service. Alerting on them covers 80% of incidents.