Sheet 13
Observability & Telemetry Sizing
Metrics ingestion volume, log retention disk storage, distributed OpenTelemetry trace spans, and SLO error budgets.
Daily Log Volume114.4 GB/day (30d: 3433 GB)
Prometheus Metrics
57.6B pts/day
429.1 GB/day storage (100 metrics/req)
OpenSearch Logs
114.4 GB/day
3.4 TB across 30-day retention
OpenTelemetry Traces
2.86 GB/day
1% sampling rate (5 KB / span)
Service Level Objectives (SLOs) & Error Budget Allocation
| SLO Category | Target Objective | Monthly Error Budget | Telemetry Source |
|---|---|---|---|
| Availability | 99.95% (Three and a half 9s) | 21.9 minutes outage / mo | HTTP 5xx error rate from Gateway |
| P95 Latency | < 200 ms (Order checkout) | 5% of requests may exceed 200ms | Envoy service mesh timer histograms |
| Kafka Ingestion Lag | < 1000 messages / partition | Max lag duration < 60 seconds | kafka_consumergroup_lag exporter |
| Payment Success Rate | 99.99% (Four 9s) | 0.01% unrecoverable failure allowed | Dead-letter queue topic counter |