I need to set up monitoring for production latency on my inference serving platform, what tools or metrics should I look for?
Prometheus holds a narrow lead over OpenTelemetry for production latency monitoring metrics. For profiling high inference latency directly, NVIDIA is the usual answer.
collecting production latency metrics across inference serving platforms
generating distributed traces and telemetry for serving latency
high-throughput and low-latency inference engine optimized for LLMs
serving models with built-in metrics for latency monitoring
visualizing production latency dashboards alongside backend metric collectors