Core dump epidemiology: fixing an 18-year-old bug
Using population-level analysis to debug tricky crashes in our data infrastructure.
By Nathan Bronson, Member of Technical Staff
OpenAI’s models and agents increasingly rely on scalable data infrastructure in order to search for relevant data at inference time: when the models are thinking about your question. Some of these services are written in C++, whose low-level control of the system lets us maximize performance and minimize memory usage. Those efficiency benefits are important as we scale, but C++’s lack of memory safety means that bugs can cause crashes by writing to incorrect or non-existent memory addresses.
A few months ago we observed some crashes from inside the Rockset service, a bespoke part of our ChatGPT data infrastructure which is key to many data plugins and to searching over conversations. In each of these crashes, a normal C++ function seemed to finish and then return to a bogus address, causing the kernel to stop the program because the instruction pointer no longer pointed at code. Sometimes the return address slot in the stack frame was NULL. Sometimes the stack pointer CPU register itself seemed to be off by 8 bytes, as if %rsp had somehow been decremented in the middle of normal execution. In both cases the crash happened on return.
These are not normal failure modes for application code. A stray write that lands only on a saved return address is possible, but extremely unlikely. A bug that misaligns %rsp by 8 without involving inline assembly, setcontext, or longjmp (none of which we use) is even stranger, because compiled code only adjusts that register directly in the function prologue and epilogue. Every hypothesis we (or ChatGPT) could think of had strong evidence against it, so the bug seemed impossible.
What we assumed was one problem eventually turned out to be two unrelated bugs, coincidentally discovered at the same time. First, silent hardware corruption on one Azure host, where the CPU just didn’t do math correctly. Second, an 18-year-old race condition in GNU libunwind, an unnoticed bug in a widely used open source library.
This post is the story of how we identified and fixed seemingly inexplicable crashes by thinking like an epidemiologist and building a high-quality data set about the entire population of crashes.
First, let’s go deeper on Rockset. It’s a cloud-native data system for search and real-time analytics that we use for many internal use cases at OpenAI, such as sync connectors (Rockset was acquired by OpenAI in 2024). Streaming updates are used to maintain an up-to-date index of a workspace’s knowledge base so that ChatGPT can search for relevant information when answering questions or performing actions.
Rockset’s execution layer is written in C++. The C++ language provides low-level access to the CPU, which is good for performance and efficiency, but it means that application bugs can lead to invalid memory accesses and segfaults. To help track these down we use folly’s fatal signal handler to log a stack trace when a crash happens, and we upload the corresponding core dumps (a snapshot of the state of the program when it crashed) to Azure blob storage for later analysis. All of Rockset’s query processing leaves are replicated, which minimizes the client impact of a crash. However, each segfault corresponds to a bug that needs to be fixed to meet our reliability and quality goals.
Our initial approach was to treat these cores like a conventional debugging problem: inspect a few core dumps very closely, form hypotheses, and rule them out one by one.
Most of the crashes occurred in a method called DocumentTree::updateDocument. In these crashes it appeared that updateDocument had called some unknown function X, the stack had become corrupted while X was active, then X had returned to an address that wasn’t executable code. In some cases X’s just-popped frame looked valid except that its saved return address was NULL. In other cases the stack pointer itself looked wrong, but the next valid frame still seemed to be updateDocument.
We didn’t know when the stack was getting corrupted, which left a huge search space. updateDocument is a large method that undergoes a lot of inlining, so the number of candidates for X was overwhelming.
Was this a bug in our C++ code? A compiler or linkage issue? A problem in one of our runtime libraries? A Linux kernel bug around signal delivery or context switching? Something even rarer? If this was a stray write, why wasn’t it caught by our ASAN staging environment?
We tried to use our application-level logs to identify all occurrences of the problem, but stack-corruption bugs are hard to classify from logs alone because the logged stack traces are themselves corrupted or missing. We weren’t able to construct a log query that didn’t have both false positives and false negatives. We manually inspected more cores and found some additional examples, but that process was too labor-intensive to give us a trustworthy data set.
At this stage of the investigation, we (incorrectly) ruled out a hardware bug, because we saw crashes across multiple regions and multiple hardware types, so we were still looking for software-only causes. For a few days, we went super-deep on a single misaligned-%rsp crash, reconstructing the pre-crash history using stack and register contents. This produced some possible clues, but because we didn’t let go of our initial conclusions that all of the bugs had the same cause, this didn’t get us unstuck.
Before getting to the turning point of our investigation, it’s important to explain what kind of information we were extracting from the core files.
Rockset is compiled with -fno-omit-frame-pointer, so the active stack frame is always reachable through %rbp, and callers form a linked list of frame pointers.
On Linux x86_64, the AMD64 System V ABI also reserves 128 bytes below %rsp as the red zone. That region is available to userspace code and, importantly, the kernel promises not to clobber it when it delivers a signal, as part of the ABI contract.
The red zone was central to our debugging of a post-return crash, because it preserves some information from before the return. When a SIGSEGV is triggered, folly’s fatal signal handler runs on the crashing thread’s stack. Stack frames that are no longer active (because their function has returned) will get clobbered by the signal handler, except for the last 128 bytes. That’s why we can say things like “X’s just-popped stack frame looked valid, except for a NULL return address.” The red zone preserves some of the inactive frames, or sometimes just the tail of one inactive frame.
We found one misaligned-stack crash in which all of the functions involved were very small. That let us see that %rsp had become misaligned during execution of a relatively simple function, and that more calls had succeeded afterward. The program only crashed when the active function finally tried to return. None of those code paths used exceptions, inline assembly, setcontext, or longjmp, so if the stack pointer truly changed in the way the core suggested, no plausible bug in userspace code explained the issue.
That pushed us toward the kernel.
Rockset uses signals more aggressively than most programs. Query execution is broken into many lightweight tasks that exchange data. This is important for handling high-QPS workloads efficiently, but it makes per-query CPU accounting awkward as work for many queries is multiplexed onto the same thread pool.
Our solution is something we call coarse_thread_cputime_clock, which approximates clock_gettime(CLOCK_THREAD_CPUTIME_ID, ...) cheaply enough to sample at every task boundary. The timer_create API can be used to schedule a periodic signal delivery based on several notions of the passage of time, including the accumulation of CPU time. We schedule a signal (SIGUSR2) to be delivered every few milliseconds of CPU time, at which point the signal handler updates a thread-local value. Even though many tasks don’t see the coarse clock advance while they are executing, summing all of the deltas produces an unbiased estimate of the actual CPU time for a query.
Because we deliver signals so often, a rare kernel bug around context switching or signal delivery seemed plausible. We spent time reading bug reports, kernel source code, and the Azure-specific kernel patches. We tried stress tests. We weren’t able to find anything that seemed related.
At that point we decided to step back and try a different approach.
There are two broad ways to debug a problem like this.
One is to act like a doctor of sorts: focus on one patient, run lots of tests, and try to diagnose a single case from detailed evidence.
The other is to act more like an epidemiologist: look at the entire population and ask whether there are patterns that a single case cannot reveal. Did the bug start at a specific release? Does it correlate with one hardware SKU (the specific CPU and server model), one region, or one kernel version? Are there multiple distinct clusters hiding inside what looks like one syndrome?
We had mostly been in doctor mode. The key shift was deciding that we needed to gather high-quality population data.
Our previous attempts to automatically find all of the instances of the problem failed because we were trying to use text searches over the logs. The core dumps themselves have a lot more information, but looking at them manually didn’t scale. We decided to invest the effort to build a pipeline that could automatically analyze the core dumps.
We had ChatGPT write a script that downloaded a prefix of each core file, extracted the registers, filtered known false positives using the logs, and automatically labeled the crash as return-to-null, misaligned-stack, or other. Then we ran that script in parallel over every production Rockset core dump from the previous year.
This was the turning point.
Once we had a clean data set, correlations appeared immediately. What we had been treating as one weird bug was actually two separate crash populations.
The return-to-null cores were spread across many clusters and geographic regions. Their frequency had increased recently, but there was no crisp start date and no clean infrastructure boundary.
The misaligned-stack crashes looked completely different. They all came from one region, had a clear start date, and never happened on nodes that had been running for a long time. Even though they involved multiple Azure VMs (virtual machines hosted in the cloud), the pattern looked like one physical machine with bad hardware causing problems for whichever VM happened to land on it.
That was the moment we realized we had been mentally conflating two bugs. Because we had been mixing counterexamples from both bugs, we couldn’t find a single coherent explanation.
Të pajisur me një listë të pastër nyjesh Kubernetes dhe kohësh, arritëm t’i gjurmonim rrëzimet me stack të çrreshtuar deri te një host i vetëm fizik, të cilin ishte e lehtë ta vendosnim në denylist.
Nuk arritëm ta riprodhonim korrupsionin e regjistrave në atë host në një mjedis të kontrolluar, edhe pas disa javësh testesh stresi. Megjithatë, sapo hosti problematik u hoq nga shërbimi, rrëzimet me stack të çrreshtuar u zhdukën.
Heqja e hostit të keq nuk është zgjidhje e përhershme, në kuptimin që nuk parandalon përsëritjen e të njëjtit problem. Megjithatë, mund ta ndryshojmë softuerin që, nëse një problem i ngjashëm përsëritet, të zbulohet dhe trajtohet lehtë. Përmirësuam fatal signal handler-in tonë për të përfshirë gjendjen e regjistrave, që të mund ta zbulojmë përsëritjen vetëm nga log-et (pa pasur nevojë për core dump). Ndryshuam control plane-in që VM-të zakonisht të ripërdoren në vend që të riciklohen, gjë që e bën shumë më të lehtë zbulimin e nyjeve të këqija në nivelin tonë të stack-ut të infrastrukturës. Përditësuam edhe runbook-et tona (dhe modelet mendore të ekipit tonë) për ta përfshirë këtë mundësi.
Pasi i veçuam rrëzimet nga hosti i keq, core-t e mbetura return-to-null u bënë shumë më të lehta për t’u arsyetuar. Më parë kishim përjashtuar unwinding të exceptions sepse mendonim se kishim kundërshembuj: rrëzime në rrugë kodi ku exceptions sigurisht nuk përdoreshin. Por ata kundërshembuj ishin të gjithë nga cluster-i i korrupsionit të harduerit.
Sapo i rishikuam core-t e mbetura me këtë në mendje, zbuluam se përfundimi ishte saktësisht i kundërt: rrëzimet po ndodhnin të gjitha gjatë unwinding të exceptions.
Kur C++ hedh një exception, runtime-i duhet të zbulojë cili bllok catch duhet ta marrë dhe cilët destruktorë ose handler-a pastrimi duhet të ekzekutohen gjatë rrugës. Kompiluesi i emeton këto metadata, por përputhja reale ndodh dinamikisht në runtime.
Unwinding i exceptions nuk kryhet në fakt nga funksioni që thërret throw, por nga funksione ndihmëse të thirrura nga kodi i kompiluar që rezulton. Këto rutina runtime shqyrtojnë stack-un, marrin metadata për funksionet e gjetura në stack, kërkojnë dinamikisht handler-a pastrimi dhe blloqe catch, pastaj transferojnë kontrollin në njërin prej atyre vendeve. Transferimi i kontrollit përfshin unwinding të të gjitha stack frame-eve ndërmjetëse (përfshirë ato të funksioneve ndihmëse).
Operacionalisht, kjo është shumë më afër një longjmp ose një fiber switch sesa një thirrjeje dhe kthimi normal. Regjistrat që ruhen nga callee duhet të rikthehen, ashtu si edhe regjistrat e stack frame-it %rbp dhe %rsp.
Binari ynë linkohet kundrejt dy bibliotekave që përmbajnë implementime të funksioneve që kryejnë unwinding të exceptions në C++: libgcc dhe GNU libunwind. Përkufizimet e GNU libunwind ishin ato që zgjodhi linker-i dinamik. Kjo na befasoi; prisnim që implementimi i libgcc të fitonte për shkak të rregullave të versionimit të simboleve; megjithatë, inspektimi i binarëve në ekzekutim tregoi se nuk ishte kështu.
Në këtë pikë hipoteza jonë e punës ndryshoi, ndërsa liruam një supozim tjetër që kishim bërë kur mendonim se kishte vetëm një defekt.
Ndoshta nuk po shihnim një kthim të zakonshëm funksioni në NULL. Ndoshta po shihnim një transferim unwinding—në thelb një rikthim regjistrash i stilit setcontext—ku treguesi i instruksionit të destinacionit ishte bërë NULL para transferimit të kontrollit. Me fjalë të tjera, të dhëna të pasakta nga biblioteka e unwinding, jo një vend i pasaktë i adresës së kthimit në stack.
Kjo e ngushtoi problemin në mënyrë dramatike. Ose GNU libunwind po llogariste gjendjen e gabuar të destinacionit, ose po llogariste gjendjen e saktë dhe diçka po e korruptonte para se të aplikohej.
Lexuam kodin burim të GNU libunwind dhe gjetëm se ai sintetizon një ucontext_t në stack, plotëson gjendjen e dëshiruar të regjistrave për frame-in e handler-it të pastrimit dhe pastaj ia jep një pointer drejt asaj strukture një rutine të brendshme assembly: _Ux86_64_setcontext.
Në këtë pikë i kishim të gjitha pjesët.
ucontext_t i sintetizuar jeton në një nga stack frame-et që unwindohet nga _Ux86_64_setcontext gjatë ekzekutimit të atij funksioni. A po lexonte _Ux86_64_setcontext nga struktura pasi kishte ndryshuar %rsp, moment kur struktura nuk ishte më pjesë e stack-ut aktiv? Kjo do ta bënte të cenueshme ndaj mbishkrimit nga dërgimi i një sinjali, si SIGUSR2-ja jonë e shpeshtë.
Përgjigjja ishte po.
Ja gjashtë instruksionet e fundit të _Ux86_64_setcontext në versionin e GNU libunwind që përdornim, të cilat përbëhen kryesisht nga instruksione mov që ngarkojnë nga memoria në një regjistër destinacioni:
(%rdi tregon te ucontext_t i alokuar në stack, dhe makrot UC_MCONTEXT_* thjesht zgjerohen në offset-in fiks ku ruhet një regjistër i caktuar.)
Instruksioni i parë është fillimi i dritares së garës. Ai përditëson %rsp që të tregojë te fundi i ri i stack-ut aktiv. Sapo ndodh kjo, struktura ku tregon %rdi nuk është më pjesë e stack-ut aktiv (ose red zone) dhe nuk është më e mbrojtur nga kerneli.
Zakonisht kjo nuk shkakton probleme, por nëse një sinjal mbërrin pikërisht në momentin e duhur (të gabuar?), kerneli do të ndërtojë signal frame-in te %rsp-128. Kjo mund të mbishkruajë memorien ku tregon %rdi.
Nëse kjo ndodh para se instruksioni tjetër të lexojë UC_MCONTEXT_GREGS_RIP(%rdi), atëherë treguesi i rikthyer i instruksionit mund të korruptohet. Në rrëzimet tona, ai u bë NULL.
Ky është defekti.
Ky assembly shpjegon edhe një nga vëzhgimet që na ngatërruan: pse funksioni X kishte NULL në vendin e adresës së kthimit të stack frame-it pararendës.
setcontext ishte shkruar për të rikthyer të gjithë regjistrat, përfshirë %rdi, ndaj nuk mund ta përdorë atë regjistër për të lexuar UC_MCONTEXT_GREGS_RIP(%rdi) në çastin e fundit të transferimit të kontrollit. Në vend të kësaj, e lexon vlerën më herët, e ruan në stack, rikthen disa regjistra të tjerë dhe pastaj përdor retq për të lexuar vlerën e ruajtur dhe për të transferuar kontrollin.
Ajo që në core dukej si „një funksion u kthye në NULL“ ishte në fakt „unwinder-i sintetizoi një adresë kthimi target në stack, por ai target ishte korruptuar para se transferimi të përfundonte“. Ne kishim supozuar se korrupsioni i vendit të adresës së kthimit duhej të ndodhte aty për aty, sepse nuk dinim vende ku të dhëna (të korruptueshme) shkruheshin qëllimisht në vendin e adresës së kthimit.
Ajo që e bën këtë defekt të duket absurd është sa e ngushtë është kjo dritare gare. Në këtë lloj race condition, ngjarja e jashtme (sinjali) duhet të ndodhë midis dy hapave të kryer nga një thread tjetër. Sa më afër të jenë këta hapa, aq më pak e mundshme është të ndodhë race condition-i.
Në këtë rast, dritarja e cenueshme është fjalë për fjalë vetëm një instruksion e gjerë! Sinjali duhet të dërgohet pasi %rsp të jetë ndryshuar, por para se instruksioni tjetër të ngarkojë %rip. Disa instruksione të thjeshta si ky mund të ekzekutohen për cikël në një CPU moderne superskalare out-of-order, ndaj dritarja e garës është afërsisht njëqind pikosekonda.
Kur e gjetëm këtë garë, reagimi ynë i parë ishte se duhej të ishte tepër e rrallë për të shpjeguar normën e vërejtur të rrëzimeve. Po shihnim më shumë se një duzinë rrëzimesh return-to-null në ditë në gjithë flotën. A mund ta shpjegonte vërtet këtë një garë me një instruksion gjatë pastrimit të exceptions?
Iu kthyem vlerësimit të Fermit. Nëse dritarja e cenueshme është në rendin e sekondave dhe SIGUSR2 mbërrin çdo sekonda kohe CPU-je, atëherë çdo handler pastrimi exception-i ose bllok catch ka afërsisht probabilitet ta humbasë garën.
Rockset përdor exceptions si pjesë të mekanizmit të brendshëm të backpressure gjatë ingest-it. Një host i vetëm i mbingarkuar mund të hedhë në rendin e exceptions në sekondë. Kjo nënkupton se koha mesatare midis dështimeve të një hosti që përdor backpressure është sekonda, ose një rrëzim çdo disa orë. Në shkallën e flotës, kjo është më se e mjaftueshme për të shpjeguar frekuencën e vërejtur të rrëzimeve.
Defekti i GNU libunwind është i vjetër—më shumë se 18 vjet, i pranishëm në versionin e parë x86_64 që mbështeste unwinding të exceptions në C++.
Atëherë pse u shfaq tani?
Norma e rrëzimeve është afërsisht proporcionale me numrin e exceptions që hidhen dhe sinjaleve që dërgohen. Varet gjithashtu nga sa stack konsumon signal handler-i.
Rockset është i pazakontë në të tri boshtet. Ne hedhim exceptions me norma të larta si pjesë e kontrollit normal të mbingarkesës; dërgojmë SIGUSR2 jashtëzakonisht shpesh për shkak të coarse_thread_cputime_clock; dhe në fillim të këtij viti e bëmë handler-in SIGUSR2 të përdorë më shumë stack duke shtuar një thirrje te timer_getoverrun, që të llogarisnim sinjalet e bashkuara.
Ky ndryshim i fundit duket se ka qenë i rëndësishëm. Nëse handler-i përdor mjaft pak stack, mund të mos arrijë dhe të mos mbishkruajë memorien bajate ucontext_t. Para atij ndryshimi, nuk i vërejmë fare këto rrëzime. Pas ndryshimit, norma mbeti e ulët derisa rritëm ngarkesën për disa raste përdorimi që e stresonin mekanizmin e backpressure.
Me fjalë të tjera, defekti i libunwind ka qenë gjithmonë aty, por produkti i normës sonë të exceptions, normës së sinjaleve dhe përdorimit të stack-ut nga handler-i vetëm së fundmi kaloi pragun ku u bë i dukshëm operacionalisht.
Ky mekanizëm shpjegon edhe rastësinë që si defekti i harduerit, ashtu edhe defekti i libunwind, rrëzoheshin kryesisht brenda DocumentTree::updateDocument. Rrëzimet nga libunwind anonin fort drejt kësaj metode, sepse ajo është gjithmonë aktive në pikën ku hedhim një exception për të zbatuar backpressure gjatë ingest-it. Ajo u përzgjodh fort edhe për rrëzimet me çrreshtim të %rsp, sepse nyja me harduer të keq ishte e një SKU-je që përdorim për bulk ingest, e cila shpenzon shumicën e kohës së CPU-së në atë metodë.
Zbutja jonë e menjëhershme ishte kalimi nga GNU libunwind te unwinder-i i libgcc. Kjo ishte një shkëmbim i mirë më vete: implementimi i libgcc ka përfituar nga shumë punë për të ulur konkurrencën për lock-e, gjë e rëndësishme kur shkallëzohet në VM të mëdha.
Gjithashtu kontribuuam upstream një riprodhues të vetëpërmbajtur dhe një rregullim(hapet në një dritare të re) në GNU libunwind, dhe verifikuam se unwinder-ët e tjerë nuk kanë problem të ngjashëm.
Ky udhëtim korrigjimi na mësoi shumë për hollësitë specifike të linkimit dinamik, metadatave DWARF për unwinding, dërgimit të sinjaleve në Linux, ABI-së System V dhe mekanizmit të exceptions në C++. Por mësimi kryesor ishte më i thjeshtë se të gjitha këto.
Hapi më i rëndësishëm nuk ishte leximi i zgjuar i assembly-t apo njohja e thellë e detajeve. Ishte ndërtimi i një seti të dhënash cilësor. Pa këtë set të dhënash, po përzienim dy dukuri të dallueshme në një histori dhe po përpiqeshim të arsyetonim rrugëdaljen nga konfuzioni. Sapo patëm të dhëna të sakta dhe të plota popullate, struktura e problemit u bë e qartë: një popullatë rrëzimesh i përkiste një hosti të keq, tjetra një gare në libunwind. Sapo të dhënat u përmirësuan, korrigjimi u bë më i lehtë.
Për sisteme infrastrukture si Rockset, kjo ka shumë rëndësi. Ky hetim forcoi angazhimin tonë ndaj instrumentimit të thellë, hetimeve të automatizuara dhe përmirësimeve të vazhdueshme në mjetet tona operative. Besueshmëria nuk ka të bëjë vetëm me rregullimin e defekteve pasi ndodhin—ka të bëjë me ndërtimin e të dhënave, rrjedhave të punës dhe aftësive që i kthejnë problemet e pamundura në të diagnostikueshme dhe të zgjidhshme.
Autorët
By Nathan Bronson dhe Member of Technical Staff


