Conversation
maleadt
added this pull request to stack #944
September 24, 2026 15:38
maleadt
removed this pull request from stack #944
September 25, 2026 06:14
maleadt
changed the base branch from
tb/throw-arguments
to
tb/remove-atomics-demotion
September 25, 2026 06:14
maleadt
added this pull request to stack #951
September 25, 2026 06:14
`lower_fences!` used thread scope for `singlethread` fences and device scope for all others. Map the scopes the LLVM SPIR-V back-end uses to their MSL equivalents instead: `subgroup` to simdgroup, `workgroup` to threadgroup, and `device` and the system scope to device (Metal has no wider scope). Reject other scopes in `validate_ir`, as the NVPTX and AMDGPU back-ends do, rather than guessing what they mean. The scope and memory-order mappings are shared with the atomics lowering that follows, as are the target-independent helpers in `atomics.jl`.
Metal has no LLVM back-end, so LLVM's atomic instructions reached Apple's compiler as-is. It accepts the common cases, but it cannot expand what AIR lacks (8- and 16-bit operations, nand, floating-point min/max), ignores synchronization scopes, and only has ordered atomics from MSL 4.1. Do what `AtomicExpand` and instruction selection do for other targets, so that front-ends can emit plain LLVM atomics with an ordering and a scope: - `validate_ir` rejects the atomics that cannot be lowered: outside device and threadgroup memory, wider than 32 bits (except non-fetching 64-bit umin/umax on device memory), misaligned, with an unknown scope, or ordered below MSL 3.2; - below MSL 4.1, ordered operations are relaxed ones bracketed with fences (AtomicExpand's default leading/trailing scheme); - floating-point loads, stores and exchanges are cast to integers, 8- and 16-bit operations become masked operations on the containing word, and operations AIR lacks become compare-exchange loops; - the rest are selected to `air.atomic.*` calls, in the form of the target's AIR and MSL versions. `cmpxchg` becomes a weak compare-exchange with the success flag derived from the old value, as MSL does. One legality function, `metal_atomic_action`, decides both: like the rule tables of LLVM's legalizers, it returns how to lower an operation, or why it cannot be, which `validate_ir` reports. Front-ends that call `air.atomic.*` directly (e.g. to pass memory flags) do so in the MSL 4.1 form, which is now legalized for older targets here instead of in Metal.jl. The intrinsics are declared like Apple does (`mustprogress nounwind willreturn`), no longer `argmemonly` and `readonly`/`writeonly`, which would allow LLVM to move other memory accesses across ordered atomics.
When `AllocOpt` moves an object with `@atomic` fields to the stack, it keeps the compare-exchange loops on its fields (#934). MSL has no atomics on thread memory, so the lowering rejected those, and before it Apple's compiler handled them inconsistently: macOS 15's silently produced wrong results. They don't need to be atomic, though: no thread can access another's stack, so plain accesses behave the same, whatever the ordering or scope. Lower atomics on pointers that are derived from allocas only to plain loads and stores.
From Julia 1.13, `@atomic` modifications of fields are calls to `julia.atomicmodify`, which AllocOpt treats as an escape, so the object in the test stays on the heap and there are no atomics on the stack to lower. The IR-level test of the lowering still covers every version.
…ites. MSL 4.1 clears the volatile bit of atomics on non-`volatile` objects, and we did the same for non-volatile LLVM atomics. Apple's compiler then treats the atomic like a plain access: a relaxed load of memory the kernel doesn't write is hoisted into the uniform preamble (out of a spin loop), and a read-modify-write that leaves memory unchanged (adding or or-ing 0, whatever the ordering) becomes a load. On the M1 (macOS 27), a spin-wait on such a read-modify-write never observes the store another threadgroup makes, while the volatile one does. LLVM atomics allow neither transformation, so set the bit on loads and read-modify-writes, as MSL does on every atomic before 4.1. Stores and compare-exchanges compile the same either way, as does any read-modify-write that does change memory, and the bit costs nothing measurable.
From MSL 4.1, we select an LLVM seq_cst atomic to the MSL 4.1 seq_cst intrinsic, but Apple's compiler emits a seq_cst store as a release store that doesn't wait for its write. A later seq_cst load can then be performed before it: in the store-buffering litmus test (`store x; load y` against `store y; load x`, across threadgroups), both loads return the old values about 0.5% of the time on an M1, for LLVM atomics and MSL's own seq_cst atomics alike. Follow seq_cst stores by a seq_cst fence, like AtomicExpand does for targets whose `shouldInsertTrailingSeqCstFenceForAtomicStore` returns true. Read-modify-writes and compare-exchanges wait for their result and already behave; below MSL 4.1, the fences that bracket every ordered operation already did. On the M1, this removes all store-buffering violations (none in 6.5M runs, also for the RMW and compare-exchange variants), and costs 4-11% on a seq_cst-store-bound kernel. Fences before seq_cst loads work too, but cost 50% on a load-bound one.
LLVM requires that a thread repeatedly loading an address monotonically eventually sees the stores of other threads. MSL doesn't: device memory is only coherent within a threadgroup by default, and relaxed atomics only guarantee atomicity and a consistent modification order (MSL 4.1 §4.8, §6.16.1.1). Apple's compiler emits relaxed atomic loads of device memory as cached loads, so on an M1 a spin loop that relaxed-loads a flag another threadgroup sets never sees the store: not in 20M iterations, on every MSL version, with MSL's own relaxed atomics too. (Within a threadgroup it does, as the threadgroup shares the cache.) Acquire loads invalidate the cache after loading, and the same spin loop exits after a few hundred iterations. Lower device-scope monotonic loads of device memory as acquire loads, which below MSL 4.1 means a trailing fence. Loads at a narrower scope only need to see stores from their threadgroup, so they are unaffected; front-ends that want MSL's semantics (and cost) can use the workgroup scope. Relaxed loads of device memory are selected at device scope regardless, like MSL does: the scope doesn't change the code, but it is recorded. This costs nothing when the loads miss the cache, but a kernel of cache-resident device-scope relaxed loads runs 6x (MSL 4.1) to 7x (older) slower. Only strengthening loads that may repeat on the same address avoids part of that, but needs cycle analysis that isn't worth its code.
This was referenced Sep 25, 2026
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## tb/remove-atomics-demotion #942 +/- ##
==============================================================
+ Coverage 86.38% 87.15% +0.77%
==============================================================
Files 29 30 +1
Lines 5794 6277 +483
==============================================================
+ Hits 5005 5471 +466
- Misses 789 806 +17 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This was referenced Sep 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Metal has no LLVM back-end, so LLVM's atomic instructions reached Apple's compiler as-is. It handles the common cases, but it can't expand what AIR lacks (8- and 16-bit operations,
nand, floating-point min/max), ignores synchronization scopes, and only has ordered atomics from MSL 4.1. This PR does for Metal whatAtomicExpandand instruction selection do for targets with an LLVM back-end, so that front-ends (e.g. UnsafeAtomics, Atomix, or Metal.jl's own atomics) can emit plain LLVM atomics with an ordering and a synchronization scope:metal_atomic_action, decides both whatvalidate_irrejects and how the lowering handles each operation, like the rule tables of LLVM's legalizers. Rejected are atomics outside device and threadgroup memory; wider than 32 bits (except non-fetching 64-bitumin/umaxon device memory); misaligned; with an unknown scope; or ordered below MSL 3.2.air.atomic.fence.nand,fmax/fmin/fmaximum/fminimum,uinc_wrap/udec_wrap,usub_cond/usub_sat, and threadgroupfadd/fsubbelow MSL 4.1) become compare-exchange loops.air.atomic.*calls in the form of the target's AIR and MSL versions.cmpxchgbecomes a weak compare-exchange whose success flag is derived from the old value, as in MSL.air.atomic.*calls that front-ends emit directly (in the MSL 4.1 form, e.g. to pass memory flags) are legalized for older targets here, rather than in Metal.jl. This is idempotent for current Metal.jl, which legalizes its calls itself.mustprogress nounwind willreturn). Before, they were declaredargmemonlyandreadonly/writeonly, which allows LLVM to move other memory accesses across ordered atomics.Two related changes, as separate commits:
workgroupto threadgroup scope (it used to be device),subgroupto simdgroup scope,deviceand the system scope to device scope. Unknown scopes are rejected instead of guessed.AllocOptmoves an object with@atomicfields to the stack, it keeps the compare-exchange loops on its fields. MSL has no atomics on thread memory, and on macOS 15 Apple's compiler produced wrong results for them ([2.7.0] Metal: atomic RMW on thread-private memory is silently dropped #934): it selects a threadgroup compare-exchange at address 0, which never touches the stack slot. Plain accesses are equivalent, since no thread can access another thread's stack. Fixes [2.7.0] Metal: atomic RMW on thread-private memory is silently dropped #934.Three more commits implement LLVM's memory model where MSL guarantees less. I checked them with litmus tests on an M1 (macOS 27, code built for MSL 3.2, 4.0 and 4.1), and offline against the macOS 15, 26 and 27 back-ends:
volatileobjects, and without it the back-end does two things LLVM atomics don't allow. It hoists a relaxed load of memory the kernel doesn't write into the uniform preamble, i.e. out of any spin loop. It also turns a read-modify-write that doesn't change memory (e.g. adding 0) into such a load. Stores and compare-exchanges compile the same either way, and there's no measurable cost.seq_cstas acquire-release plus a single total order of modifications, and Apple compiles aseq_cststore as a release store that doesn't wait for its write. So store buffering is observable (about 0.5% of runs on an M1), which LLVM'sseq_cstforbids. With the fence it's 0. Read-modify-writes and compare-exchanges wait for their result and don't need one, and the bracketing below MSL 4.1 was already sequentially consistent. This costs 4–11% on kernels that are bound byseq_cststores.