Sheet 13

Observability & Telemetry Sizing

Metrics ingestion volume, log retention disk storage, distributed OpenTelemetry trace spans, and SLO error budgets.

Daily Log Volume114.4 GB/day (30d: 3433 GB)
Prometheus Metrics
57.6B pts/day

429.1 GB/day storage (100 metrics/req)

OpenSearch Logs
114.4 GB/day

3.4 TB across 30-day retention

OpenTelemetry Traces
2.86 GB/day

1% sampling rate (5 KB / span)

Service Level Objectives (SLOs) & Error Budget Allocation

SLO CategoryTarget ObjectiveMonthly Error BudgetTelemetry Source
Availability99.95% (Three and a half 9s)21.9 minutes outage / moHTTP 5xx error rate from Gateway
P95 Latency< 200 ms (Order checkout)5% of requests may exceed 200msEnvoy service mesh timer histograms
Kafka Ingestion Lag< 1000 messages / partitionMax lag duration < 60 secondskafka_consumergroup_lag exporter
Payment Success Rate99.99% (Four 9s)0.01% unrecoverable failure allowedDead-letter queue topic counter