Repository navigation
engine: darwin-arm64 solves in parallel — the residual crash was unfenced node publication - #1861
Conversation
…nced node publication The arm64 parallel crash that survived the seqlock entry fix was found with a TSAN build of the engine: 66 atomic-vs-plain write pairs, every one a btree split publishing a freshly built node with PLAIN pointer stores (insert_inner's getChildren()[pos + 1] = newNode and grow_parent's *root = new_root). On a weakly-ordered CPU the pointer can become visible before the node's own field stores — including its lock's start_write — so an optimistic reader descends into memory whose lock reads free and whose children are whatever the allocator left there: zeros on a fresh page (survivable), garbage on recycled memory. That is exactly the observed shape: only under memory pressure, a different rule each crash. seqlock-fix-2: a release fence before each of the two publication stores, applied by the same header overlay as fix-1, anchor-checked so an upstream header drift fails the build loudly; the engine id salt moves to +seqlock-fix-2 so no unpatched cache entry or package is ever taken for a patched one. The reader side needs nothing: the dereference is address-dependent, which arm64 orders by itself. Evidence, 4,461-file subject (1.6GB facts): unpatched -par crashed 2 of 2 runs; patched held 15 of 15 including two concurrent, relations sorted-identical to the serial flavor every run. Suites under the parallel default: typescript 239/264 (the same 25 const-object checks that fail on the base), python 306/306 through its own -par engine, front_door 21/21. Solve: 99s serial -> 66s at -j8; with this branch's rule work, 192s -> 66s (2.9x). darwin-arm64 joins linux-x64 as parallel by default; linux-arm64 carries the same fences but stays serial until validated on that hardware. AXIOM_SOLVE_PARALLEL=0 remains the one-variable rollback to serial, no rebuild. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
A contended profile put half the bundle stage in per-row INSERT dispatch (63.6s of 127s), so rows now go in 128 per statement (AXIOM_BUNDLE_BATCH tunes it; 1 is the old path, one prepared statement reused per row). The interleaved A/B on a quiet box then corrected two things at once: the stage is ~55s, not the 102-126s every contended run reported, and batching is worth 5-6% (55s -> 52s, 3 of 3 rounds), not the half the profile promised — the write is sqlite's own C work plus 1.4GB of IO, which statement dispatch barely moves. Kept because it is small, consistently faster, and gate-proven: every table identical to the per-row build, only run.created_at differs. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…on anchors too The fix-2 overlay anchors on BTree.h's two link stores and refuses an include dir without them — correctly, since skipping would mint an unpatched engine under the fix-2 id. The test's stub install only carried ParallelUtil.h, so prepare failed on CI (engine (java), first to reach it). The stub now carries both headers, and the assertions follow the overlay to include-seqlock-fix-2 and check the two fences landed. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…es through the fences A peer's repro — 6,139 files, 906MB of facts, three spinners and a concurrent reference solve — crashed 1 of 3 runs on the fix-2 binary. The two fences cover the child-pointer publications; the census of BTree.h shows three parent-pointer publications they do not: split() reparents existing children to the brand-new sibling before either fence runs, grow_parent publishes new_root through this->parent and sibling->parent before the *root fence, and insert_inner sets newNode->parent only after linking newNode in. The lock-parents walk reads exactly those pointers. The fences and the engine-id salt stay (they closed the first repro outright); the default does not flip until every publication path is fenced and the harder repro runs clean. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
|
Independent validation on a heavier repro (6,139-file subject, 906MB staged facts, 3 CPU spinners + a concurrent serial solve as load):
Suggestion before darwin goes default-on: add a Repro details: cassandra tree, facts staged with AXIOM_KEEP_FACTS=1, binary id 3786dac7…-par, loads = 3× yes + one concurrent -j1 solve of the same facts. |
… reachable fix-2 fenced the child-pointer links; the peer's harder repro (6,139 files, 906MB facts, spinners + a concurrent solve) crashed 1 of 3 through them. The census of BTree.h shows why: three parent-pointer publications were unfenced — split() reparented existing children to the brand-new sibling before any fence, grow_parent published new_root via this->parent and sibling->parent ahead of the *root fence, and insert_inner set newNode->parent only after linking newNode in. The lock-parents walk reads exactly those pointers. fix-3: split() fills the sibling completely (children, count), fences, THEN reparents; grow_parent fences before the first parent store; insert_inner gives newNode its parent and position before the fence and the link. Four fences, each before the first store that publishes a completed node. Salt +seqlock-fix-3; the engine-package stub carries the new anchors and asserts all four fences land. Validation state: engine-package java+typescript ok; the local brutal hammer and the peer's repro matrix are the gates before darwin parallel-by-default returns. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
|
fix-3 validation on the heavier repro (same 6,139-file subject, same staged facts, the concurrent -j1 side-solve that produced the fix-2 crash): 6/6 clean, 6/6 sort-equal to serial. Same conditions scored ~50% crashes on fix-1 alone and 1-of-6 on fix-2 — this is the first fully clean sheet on this repro across the whole hunt. With this, darwin default-on (reverting d6bc9b7) looks justified once your arm64 run agrees; an x64 cloud matrix + full E2E is finishing now and I'll append it. Still worth taking the |
fix-3 validation (commit
|
| Leg | Result |
|---|---|
darwin-arm64 (M-series), fix-3 -j8 |
10/10 loaded rounds clean (walls 92-286s), 58/58 relations every round |
| darwin-arm64 isolated run | 58/58 relations, sorted-identical to serial |
| linux-arm64 (GCP t2a-standard-8, gcc, souffle 2.5 built from source, same fences) | 10/10 loaded rounds clean + serial-vs-parallel parity 58/58 identical (the serial flavor alone printed the no-OpenMP warning, confirming the parallel binary really ran parallel) |
Caveats kept honest: two earlier local hammer instances overlapped in one directory and produced confusing partial-file counts — those runs are discarded; the numbers above are from a single clean instance plus the isolated checks. And my subject still isn't yours: your 6,139-file repro remains the gate. Please rerun your matrix on c4c55cfa with AXIOM_SOLVE_PARALLEL=1 — a fresh compile picks the overlay up automatically; any solve whose engine id lacks seqlock-fix-3 is the wrong binary. If your matrix runs clean, the darwin default flip comes back as its own commit on top; if it still crashes, the next census target is the rebalance path's sibling moves (iright->children[i]->parent = ileft), which fix-3 deliberately left unfenced because the destination node is pre-existing.
…the seqlock sed The release build patched only the write-entry RMW into the headers every platform compiles against, while the engine id salt moved to +seqlock-fix-3 — a dispatched build would have produced binaries labeled fix-3 but built without the publication fences, the exact mislabeling the salt exists to prevent. The gen step now applies the same anchored patch the local overlay uses, extracted from run-souffle.sh so the two cannot drift, and asserts all four fences landed. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…ed layout The release gen step sed'ed gen/souffle/souffle/utility/ParallelUtil.h — homebrew's doubled layout — while the pinned Ubuntu .deb installs a single level, so the step has failed on every runner since the seqlock sed landed; this workflow simply had not been dispatched since. Both patched headers are now located with find and a miss fails the build loudly. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
|
x64 cloud leg (n2-standard-8, fresh Ubuntu 24.04, same subject cloned at tip): 5/5 clean at -j8, 5/5 sort-equal, 65–67s per solve, and a full end-to-end index in 186s. Instance created for the run and deleted after. Full grid from my side now: darwin 6/6 (the repro that broke fix-1 and fix-2), linux-x64 5/5 + E2E. With your arm64 run green, reverting d6bc9b7 to default-on has a complete evidence sheet. |
The default gate stops naming platforms: any machine whose toolchain takes OpenMP solves parallel, because the evidence now covers the whole matrix — darwin-arm64 10/10 and linux-arm64 10/10 under load (spinners + a concurrent solve, relations identical to serial), linux-x64 5/5 on a fresh cloud box, and x64 is TSO where neither closed hole is observable. probe_openmp learns the MSYS/MinGW/Cygwin case (g++, -fopenmp as Linux). The release build follows: linux compiles OpenMP on BOTH architectures now that the publication fences are in the vendored headers, and the win32 MSVC leg gets /openmp plus the .parallel marker run-souffle reads to pass a real -j. Two caveats carried openly: MSVC's /openmp links vcomp dynamically, so the release-e2e clean-VM pass must confirm the exe starts without a toolchain installed, and the Windows stability matrix (in flight on a GCP VM) is the evidence bar before these binaries ship — same bar every other platform cleared. AXIOM_SOLVE_PARALLEL=0 stays the universal no-rebuild rollback. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
… natively souffle maps __builtin_popcountll onto MSVC's __popcnt64 and includes intrin.h for all of _WIN32, which breaks g++ on MinGW — the toolchain a Windows developer who clones this repository actually compiles with, so local engine builds (any flavor) never worked there and Windows was packaged-engines-only. Both constructs are MSVC-only concerns; the overlay narrows their guards to _MSC_VER (MiscUtil.h, PiggyList.h), found by glob in whichever include layout the install uses. Every other platform's preprocessed output is bit-identical, which is why the overlay keeps the fix-3 name. Proven by cross-compiling the full fenced TypeScript engine to a Windows executable with x86_64-w64-mingw32-g++ 16.2: fails on pristine headers at Brie.h's popcount, links clean at 18.4MB with the guards. The engine-package stub carries the new anchors. Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
|
Five-language threading proof complete on darwin-arm64 (fix-3 engines, real subjects, every pair content-gated against serial):
Zero crashes since fix-3 across all 14 loaded runs + every pair; every single comparison sort-equal or byte-equal. (*rule-hoisted serial; †under shared-machine contention, quiet-box ≈155/≈60s.) Plus the x64 leg earlier (5/5 + E2E). From here the platform story is complete minus Windows /openmp. |
The residual darwin-arm64 parallel crash, found with a TSAN build of the engine and fixed with two release fences: a btree split published a freshly built node with plain pointer stores (
getChildren()[pos + 1] = newNode,*root = new_root), so on a weakly-ordered CPU the pointer could become visible before the node's own field stores — including its lock — and an optimistic reader descended into recycled memory. Exactly the observed shape: only under memory pressure, a different rule each crash.Applied by the same header overlay as the seqlock entry fix, anchor-checked, engine-id salt bumped to
+seqlock-fix-2so no unpatched cache entry or package is ever reused.Evidence (4,461-file subject, 1.6GB facts): unpatched
-parcrashed 2 of 2 runs; patched held 15 of 15 including two concurrent, relations sorted-identical to serial every run. Suites under the parallel default: typescript 239/264 (same 25 pre-existing), python 306/306, java 333/333, front_door 21/21.Effect: darwin-arm64 joins linux-x64 as parallel-by-default — big-subject solve 99s → 66s at
-j8; stacked with the merged rule work, 192s → 66s (2.9x). All five languages' engines share the patched headers. linux-arm64 keeps the fences but stays serial until validated on that hardware (VM matrix next).AXIOM_SOLVE_PARALLEL=0is the no-rebuild rollback.Follow-up worth scheduling: upstream both fixes to souffle-lang with a minimal repro, then the overlay deletes itself once a release containing them is pinned.