Frontier AI models have become excellent at generating boilerplate functions, but a new, rigorous test of their debugging capabilities suggests they are nowhere near ready to handle production outages.
A benchmark released today, OTelBench, tested 14 leading large language models on their ability to perform a fundamental Site Reliability Engineering (SRE) task: adding distributed tracing to microservices using the industry standard, OpenTelemetry (OTel). The results are a stark reality check for the AI SRE hype cycle.
The overall pass rate across 23 tasks spanning 11 programming languages was a dismal 14%. Even the best performing model, Anthropic’s Claude Opus 4.5, succeeded only 29% of the time, while GPT 5.2 managed 26%.
Distributed tracing is essential for modern microservices architectures. When a user clicks "Login," that single action might hop across dozens of services. OTel provides the necessary instrumentation, code added to the application, to link these scattered events into a single, coherent timeline, allowing engineers to pinpoint where a request failed.
The OTelBench researchers, including Przemek Delewski and Jacek Migdał, designed tasks that would be trivial for a human SRE, involving short, clean microservices of around 300 lines of code. If the models cannot handle this, they certainly cannot handle the massive, legacy-ridden systems found in the real world.
