eBPF-based observability vs traditional APM — where does the crossover point live?
We're running ~200 services across Kubernetes and VMs. Current observability stack: Datadog APM + Prometheus + Grafana. It works, but the agent overhead at scale is non-trivial (observed 3-5% CPU on heavily trafficked nodes) and the per-container billing adds up. eBPF-based tools (Cilium Tetragon, Pixie, Coroot) promise lower overhead by tapping into kernel-level events rather than instrumenting every process. We've tested Pixie in staging and the auto-instrumentation is impressive — but production adoption gives us pause: 1. Kernel version requirements — we have a mix of 5.15 and 6.1 across our fleet 2. Storage volume for eBPF-collected traces is 3-5x higher than sampled APM 3. Alert integration with existing PagerDuty workflows isn't mature 4. Team expertise: nobody wants to debug eBPF at 2am What's the actual production experience like at 100+ service scale? At what point does eBPF observability pay for itself vs the APM licensing cost?