Skip to content

engine: darwin-arm64 solves in parallel — the residual crash was unfenced node publication - #1861

Merged
swapnilpaliwal-sd merged 9 commits into
apps/integration-0.1.9from
perf/ts-datalog-rules
Oct 8, 2026
Merged

swapnilpaliwal-sd merged 9 commits into
apps/integration-0.1.9from
perf/ts-datalog-rules

Conversation

@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor

The residual darwin-arm64 parallel crash, found with a TSAN build of the engine and fixed with two release fences: a btree split published a freshly built node with plain pointer stores (getChildren()[pos + 1] = newNode, *root = new_root), so on a weakly-ordered CPU the pointer could become visible before the node's own field stores — including its lock — and an optimistic reader descended into recycled memory. Exactly the observed shape: only under memory pressure, a different rule each crash.

Applied by the same header overlay as the seqlock entry fix, anchor-checked, engine-id salt bumped to +seqlock-fix-2 so no unpatched cache entry or package is ever reused.

Evidence (4,461-file subject, 1.6GB facts): unpatched -par crashed 2 of 2 runs; patched held 15 of 15 including two concurrent, relations sorted-identical to serial every run. Suites under the parallel default: typescript 239/264 (same 25 pre-existing), python 306/306, java 333/333, front_door 21/21.

Effect: darwin-arm64 joins linux-x64 as parallel-by-default — big-subject solve 99s → 66s at -j8; stacked with the merged rule work, 192s → 66s (2.9x). All five languages' engines share the patched headers. linux-arm64 keeps the fences but stays serial until validated on that hardware (VM matrix next). AXIOM_SOLVE_PARALLEL=0 is the no-rebuild rollback.

Follow-up worth scheduling: upstream both fixes to souffle-lang with a minimal repro, then the overlay deletes itself once a release containing them is pinned.

…nced node publication

The arm64 parallel crash that survived the seqlock entry fix was found with a TSAN build of
the engine: 66 atomic-vs-plain write pairs, every one a btree split publishing a freshly
built node with PLAIN pointer stores (insert_inner's getChildren()[pos + 1] = newNode and
grow_parent's *root = new_root). On a weakly-ordered CPU the pointer can become visible
before the node's own field stores — including its lock's start_write — so an optimistic
reader descends into memory whose lock reads free and whose children are whatever the
allocator left there: zeros on a fresh page (survivable), garbage on recycled memory. That
is exactly the observed shape: only under memory pressure, a different rule each crash.

seqlock-fix-2: a release fence before each of the two publication stores, applied by the
same header overlay as fix-1, anchor-checked so an upstream header drift fails the build
loudly; the engine id salt moves to +seqlock-fix-2 so no unpatched cache entry or package
is ever taken for a patched one. The reader side needs nothing: the dereference is
address-dependent, which arm64 orders by itself.

Evidence, 4,461-file subject (1.6GB facts): unpatched -par crashed 2 of 2 runs; patched
held 15 of 15 including two concurrent, relations sorted-identical to the serial flavor
every run. Suites under the parallel default: typescript 239/264 (the same 25 const-object
checks that fail on the base), python 306/306 through its own -par engine, front_door
21/21. Solve: 99s serial -> 66s at -j8; with this branch's rule work, 192s -> 66s (2.9x).

darwin-arm64 joins linux-x64 as parallel by default; linux-arm64 carries the same fences
but stays serial until validated on that hardware. AXIOM_SOLVE_PARALLEL=0 remains the
one-variable rollback to serial, no rebuild.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
swapnilpaliwal-sd and others added 3 commits October 7, 2026 21:06
A contended profile put half the bundle stage in per-row INSERT dispatch (63.6s of 127s),
so rows now go in 128 per statement (AXIOM_BUNDLE_BATCH tunes it; 1 is the old path, one
prepared statement reused per row). The interleaved A/B on a quiet box then corrected two
things at once: the stage is ~55s, not the 102-126s every contended run reported, and
batching is worth 5-6% (55s -> 52s, 3 of 3 rounds), not the half the profile promised —
the write is sqlite's own C work plus 1.4GB of IO, which statement dispatch barely moves.
Kept because it is small, consistently faster, and gate-proven: every table identical to
the per-row build, only run.created_at differs.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…on anchors too

The fix-2 overlay anchors on BTree.h's two link stores and refuses an include dir without
them — correctly, since skipping would mint an unpatched engine under the fix-2 id. The
test's stub install only carried ParallelUtil.h, so prepare failed on CI (engine (java),
first to reach it). The stub now carries both headers, and the assertions follow the
overlay to include-seqlock-fix-2 and check the two fences landed.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…es through the fences

A peer's repro — 6,139 files, 906MB of facts, three spinners and a concurrent reference
solve — crashed 1 of 3 runs on the fix-2 binary. The two fences cover the child-pointer
publications; the census of BTree.h shows three parent-pointer publications they do not:
split() reparents existing children to the brand-new sibling before either fence runs,
grow_parent publishes new_root through this->parent and sibling->parent before the *root
fence, and insert_inner sets newNode->parent only after linking newNode in. The
lock-parents walk reads exactly those pointers. The fences and the engine-id salt stay
(they closed the first repro outright); the default does not flip until every publication
path is fenced and the harder repro runs clean.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

Independent validation on a heavier repro (6,139-file subject, 906MB staged facts, 3 CPU spinners + a concurrent serial solve as load):

  • fix-2 parallel binary, -j8, 6 runs: 5/6 clean, 1 segfault (the crash landed under the heaviest overlap window). Before fix-2 the same conditions crashed ~half the runs; your fences clearly close most of the remaining window but not all of it on this subject.
  • Content: every completed run sort-equal to serial — correctness holds; the residual defect is availability only.

Suggestion before darwin goes default-on: add a -j1 retry when a parallel solve exits nonzero in run-souffle — same binary, single thread, so the race class cannot fire and no second compile is needed. That converts the residual crash (here ~1/6 under pathological load, likely far rarer in normal use) into a slow-but-correct solve, and makes parallel-by-default safe on every platform independent of whether a third publication site exists. Happy to land that line if you prefer; otherwise suggest darwin stays opt-in until this repro runs clean.

Repro details: cassandra tree, facts staged with AXIOM_KEEP_FACTS=1, binary id 3786dac7…-par, loads = 3× yes + one concurrent -j1 solve of the same facts.

… reachable

fix-2 fenced the child-pointer links; the peer's harder repro (6,139 files, 906MB facts,
spinners + a concurrent solve) crashed 1 of 3 through them. The census of BTree.h shows
why: three parent-pointer publications were unfenced — split() reparented existing
children to the brand-new sibling before any fence, grow_parent published new_root via
this->parent and sibling->parent ahead of the *root fence, and insert_inner set
newNode->parent only after linking newNode in. The lock-parents walk reads exactly those
pointers.

fix-3: split() fills the sibling completely (children, count), fences, THEN reparents;
grow_parent fences before the first parent store; insert_inner gives newNode its parent
and position before the fence and the link. Four fences, each before the first store that
publishes a completed node. Salt +seqlock-fix-3; the engine-package stub carries the new
anchors and asserts all four fences land.

Validation state: engine-package java+typescript ok; the local brutal hammer and the
peer's repro matrix are the gates before darwin parallel-by-default returns.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

fix-3 validation on the heavier repro (same 6,139-file subject, same staged facts, the concurrent -j1 side-solve that produced the fix-2 crash):

6/6 clean, 6/6 sort-equal to serial. Same conditions scored ~50% crashes on fix-1 alone and 1-of-6 on fix-2 — this is the first fully clean sheet on this repro across the whole hunt.

With this, darwin default-on (reverting d6bc9b7) looks justified once your arm64 run agrees; an x64 cloud matrix + full E2E is finishing now and I'll append it. Still worth taking the -j1-retry-on-nonzero as cheap belt-and-braces for whatever hardware none of us owns.

@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

fix-3 validation (commit c4c55cfa)

Responding to the loaded-repro crash on fix-2 (1 of 3 under spinners + concurrent solve): the census of BTree.h found three parent-pointer publications the two fix-2 fences did not cover — split() reparents existing children to the brand-new sibling before any fence, grow_parent publishes new_root through this->parent/sibling->parent ahead of the *root fence, and insert_inner set newNode->parent only after linking the node in. The lock-parents walk reads exactly those pointers. fix-3 puts a release fence before the first store that makes a completed node reachable, on every path (4 fences; insert_inner reordered so the node is whole before it is linked). Salt +seqlock-fix-3; the engine-package stub carries the new anchors.

Evidence so far, 4,461-file TS subject (1.6GB facts), 3 CPU spinners + a concurrent solve in every round:

Leg Result
darwin-arm64 (M-series), fix-3 -j8 10/10 loaded rounds clean (walls 92-286s), 58/58 relations every round
darwin-arm64 isolated run 58/58 relations, sorted-identical to serial
linux-arm64 (GCP t2a-standard-8, gcc, souffle 2.5 built from source, same fences) 10/10 loaded rounds clean + serial-vs-parallel parity 58/58 identical (the serial flavor alone printed the no-OpenMP warning, confirming the parallel binary really ran parallel)

Caveats kept honest: two earlier local hammer instances overlapped in one directory and produced confusing partial-file counts — those runs are discarded; the numbers above are from a single clean instance plus the isolated checks. And my subject still isn't yours: your 6,139-file repro remains the gate. Please rerun your matrix on c4c55cfa with AXIOM_SOLVE_PARALLEL=1 — a fresh compile picks the overlay up automatically; any solve whose engine id lacks seqlock-fix-3 is the wrong binary. If your matrix runs clean, the darwin default flip comes back as its own commit on top; if it still crashes, the next census target is the rebalance path's sibling moves (iright->children[i]->parent = ileft), which fix-3 deliberately left unfenced because the destination node is pre-existing.

swapnilpaliwal-sd and others added 2 commits October 7, 2026 23:25
…the seqlock sed

The release build patched only the write-entry RMW into the headers every platform
compiles against, while the engine id salt moved to +seqlock-fix-3 — a dispatched build
would have produced binaries labeled fix-3 but built without the publication fences, the
exact mislabeling the salt exists to prevent. The gen step now applies the same anchored
patch the local overlay uses, extracted from run-souffle.sh so the two cannot drift, and
asserts all four fences landed.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
…ed layout

The release gen step sed'ed gen/souffle/souffle/utility/ParallelUtil.h — homebrew's
doubled layout — while the pinned Ubuntu .deb installs a single level, so the step has
failed on every runner since the seqlock sed landed; this workflow simply had not been
dispatched since. Both patched headers are now located with find and a miss fails the
build loudly.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

x64 cloud leg (n2-standard-8, fresh Ubuntu 24.04, same subject cloned at tip): 5/5 clean at -j8, 5/5 sort-equal, 65–67s per solve, and a full end-to-end index in 186s. Instance created for the run and deleted after.

Full grid from my side now: darwin 6/6 (the repro that broke fix-1 and fix-2), linux-x64 5/5 + E2E. With your arm64 run green, reverting d6bc9b7 to default-on has a complete evidence sheet.

swapnilpaliwal-sd and others added 2 commits October 7, 2026 23:43
The default gate stops naming platforms: any machine whose toolchain takes OpenMP solves
parallel, because the evidence now covers the whole matrix — darwin-arm64 10/10 and
linux-arm64 10/10 under load (spinners + a concurrent solve, relations identical to
serial), linux-x64 5/5 on a fresh cloud box, and x64 is TSO where neither closed hole is
observable. probe_openmp learns the MSYS/MinGW/Cygwin case (g++, -fopenmp as Linux).

The release build follows: linux compiles OpenMP on BOTH architectures now that the
publication fences are in the vendored headers, and the win32 MSVC leg gets /openmp plus
the .parallel marker run-souffle reads to pass a real -j. Two caveats carried openly:
MSVC's /openmp links vcomp dynamically, so the release-e2e clean-VM pass must confirm the
exe starts without a toolchain installed, and the Windows stability matrix (in flight on
a GCP VM) is the evidence bar before these binaries ship — same bar every other platform
cleared. AXIOM_SOLVE_PARALLEL=0 stays the universal no-rebuild rollback.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
… natively

souffle maps __builtin_popcountll onto MSVC's __popcnt64 and includes intrin.h for all of
_WIN32, which breaks g++ on MinGW — the toolchain a Windows developer who clones this
repository actually compiles with, so local engine builds (any flavor) never worked there
and Windows was packaged-engines-only. Both constructs are MSVC-only concerns; the overlay
narrows their guards to _MSC_VER (MiscUtil.h, PiggyList.h), found by glob in whichever
include layout the install uses. Every other platform's preprocessed output is
bit-identical, which is why the overlay keeps the fix-3 name. Proven by cross-compiling
the full fenced TypeScript engine to a Windows executable with x86_64-w64-mingw32-g++
16.2: fails on pristine headers at Brie.h's popcount, links clean at 18.4MB with the
guards. The engine-package stub carries the new anchors.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd

Copy link
Copy Markdown
Contributor Author

Five-language threading proof complete on darwin-arm64 (fix-3 engines, real subjects, every pair content-gated against serial):

Language Subject -j1 -j8 Scaling Loaded matrix
Java 6,139-file subject 94.5s* 57.9s 2.5× 6/6 clean
Python (stacked rules+threads) 526k-LOC subject 183.4s† 73.1s† 2.5× 4/4 clean
TypeScript 921k-LOC subject 111.0s 51.7s 2.15× 4/4 clean
C# 801 / 1,145-file subjects 21.8 / 23.0s 7.6 / 12.1s 2.9× / 1.9× —
JavaScript 638-file subject 30.5s 8.4s 3.63× —

Zero crashes since fix-3 across all 14 loaded runs + every pair; every single comparison sort-equal or byte-equal. (*rule-hoisted serial; †under shared-machine contention, quiet-box ≈155/≈60s.) Plus the x64 leg earlier (5/5 + E2E). From here the platform story is complete minus Windows /openmp.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant