OpenAI’s sophisticated AI models rely on robust data infrastructure, often built with C++ for performance. However, C++’s lack of memory safety can lead to critical crashes. Recently, the company grappled with inexplicable crashes within its ChatGPT data infrastructure, specifically the Rockset service. These incidents involved functions appearing to return to invalid memory addresses, often with corrupted stack frames or misaligned stack pointers.
The initial debugging efforts, focusing on individual core dumps, proved fruitless. Hypotheses about bugs in custom C++ code, compilers, or even the Linux kernel were systematically ruled out, leading engineers to believe the problem was uniquely strange. The crashes seemed to occur on return from a function, with the return address slot in the stack frame sometimes being NULL or the stack pointer register misaligned.
Doctor or Epidemiologist?
The debugging team realized their conventional, case-by-case approach was insufficient. They shifted to an epidemiological mindset, seeking patterns across the entire population of crashes.
This pivot required building a high-quality dataset of crash information. Previous attempts to analyze logs failed due to corruption in stack traces. The team developed a pipeline to automatically process core dumps, extracting critical data like registers and filtering false positives.
The Unveiling of Two Bugs
This population-level analysis revealed not one, but two distinct sets of crashes. The first cluster, initially exhibiting a return-to-null behavior, was eventually traced to a silent hardware corruption issue on a single Azure host. The problematic host was denylisted, and improved monitoring was implemented.