Skip to content

Metal: lower LLVM atomics to AIR atomic intrinsics - #942

Open
maleadt wants to merge 7 commits into
tb/remove-atomics-demotionfrom
tb/atomics
Open

maleadt wants to merge 7 commits into
tb/remove-atomics-demotionfrom
tb/atomics

Conversation

@maleadt

@maleadt maleadt commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Metal has no LLVM back-end, so LLVM's atomic instructions reached Apple's compiler as-is. It handles the common cases, but it can't expand what AIR lacks (8- and 16-bit operations, nand, floating-point min/max), ignores synchronization scopes, and only has ordered atomics from MSL 4.1. This PR does for Metal what AtomicExpand and instruction selection do for targets with an LLVM back-end, so that front-ends (e.g. UnsafeAtomics, Atomix, or Metal.jl's own atomics) can emit plain LLVM atomics with an ordering and a synchronization scope:

  • One legality function, metal_atomic_action, decides both what validate_ir rejects and how the lowering handles each operation, like the rule tables of LLVM's legalizers. Rejected are atomics outside device and threadgroup memory; wider than 32 bits (except non-fetching 64-bit umin/umax on device memory); misaligned; with an unknown scope; or ordered below MSL 3.2.
  • Below MSL 4.1, ordered operations become relaxed ones bracketed with fences (AtomicExpand's default leading/trailing scheme), lowered to air.atomic.fence.
  • Floating-point loads, stores and exchanges are cast to integers. 8- and 16-bit operations become masked operations on the containing word. Operations AIR lacks (nand, fmax/fmin/fmaximum/fminimum, uinc_wrap/udec_wrap, usub_cond/usub_sat, and threadgroup fadd/fsub below MSL 4.1) become compare-exchange loops.
  • What's left is selected to air.atomic.* calls in the form of the target's AIR and MSL versions. cmpxchg becomes a weak compare-exchange whose success flag is derived from the old value, as in MSL.
  • air.atomic.* calls that front-ends emit directly (in the MSL 4.1 form, e.g. to pass memory flags) are legalized for older targets here, rather than in Metal.jl. This is idempotent for current Metal.jl, which legalizes its calls itself.
  • The intrinsics are declared like Apple does (mustprogress nounwind willreturn). Before, they were declared argmemonly and readonly/writeonly, which allows LLVM to move other memory accesses across ordered atomics.

Two related changes, as separate commits:

  • Fence scopes are mapped like MSL does: workgroup to threadgroup scope (it used to be device), subgroup to simdgroup scope, device and the system scope to device scope. Unknown scopes are rejected instead of guessed.
  • Atomics on thread-private memory become plain accesses. When Julia's AllocOpt moves an object with @atomic fields to the stack, it keeps the compare-exchange loops on its fields. MSL has no atomics on thread memory, and on macOS 15 Apple's compiler produced wrong results for them ([2.7.0] Metal: atomic RMW on thread-private memory is silently dropped #934): it selects a threadgroup compare-exchange at address 0, which never touches the stack slot. Plain accesses are equivalent, since no thread can access another thread's stack. Fixes [2.7.0] Metal: atomic RMW on thread-private memory is silently dropped #934.

Three more commits implement LLVM's memory model where MSL guarantees less. I checked them with litmus tests on an M1 (macOS 27, code built for MSL 3.2, 4.0 and 4.1), and offline against the macOS 15, 26 and 27 back-ends:

  • The volatile bit is set on all atomic loads and read-modify-writes. Since MSL 4.1, Apple only sets it on atomics of volatile objects, and without it the back-end does two things LLVM atomics don't allow. It hoists a relaxed load of memory the kernel doesn't write into the uniform preamble, i.e. out of any spin loop. It also turns a read-modify-write that doesn't change memory (e.g. adding 0) into such a load. Stores and compare-exchanges compile the same either way, and there's no measurable cost.
  • From MSL 4.1, a sequentially consistent store is followed by a sequentially consistent fence. This is AtomicExpand's trailing-fence scheme, as AArch64 uses it for MSVC. MSL defines seq_cst as acquire-release plus a single total order of modifications, and Apple compiles a seq_cst store as a release store that doesn't wait for its write. So store buffering is observable (about 0.5% of runs on an M1), which LLVM's seq_cst forbids. With the fence it's 0. Read-modify-writes and compare-exchanges wait for their result and don't need one, and the bracketing below MSL 4.1 was already sequentially consistent. This costs 4–11% on kernels that are bound by seq_cst stores.
  • Device-scope relaxed loads of device memory become acquire loads. MSL only makes device memory coherent within a threadgroup unless you synchronize (MSL 4.1 §4.8, §6.16.1.1), and relaxed loads are served from a cache that other threadgroups' stores don't update. A spin loop waiting for another threadgroup's store never sees it, not in 20M iterations, and that includes MSL's own relaxed loads. LLVM requires monotonic loads to eventually see other threads' stores. This costs nothing when the loads miss the cache, but up to 6–7× when they hit it. Front-ends that want MSL's semantics (and cost) can use the workgroup scope, as Metal.jl does for its own relaxed loads. Relaxed loads of device memory are selected at device scope regardless, like MSL does.

@maleadt
maleadt added this pull request to stack #944 September 24, 2026 15:38
@maleadt
maleadt removed this pull request from stack #944 September 25, 2026 06:14
@maleadt
maleadt changed the base branch from tb/throw-arguments to tb/remove-atomics-demotion September 25, 2026 06:14
@maleadt
maleadt added this pull request to stack #951 September 25, 2026 06:14
`lower_fences!` used thread scope for `singlethread` fences and device
scope for all others. Map the scopes the LLVM SPIR-V back-end uses to
their MSL equivalents instead: `subgroup` to simdgroup, `workgroup` to
threadgroup, and `device` and the system scope to device (Metal has no
wider scope). Reject other scopes in `validate_ir`, as the NVPTX and
AMDGPU back-ends do, rather than guessing what they mean.

The scope and memory-order mappings are shared with the atomics lowering
that follows, as are the target-independent helpers in `atomics.jl`.
Metal has no LLVM back-end, so LLVM's atomic instructions reached Apple's
compiler as-is. It accepts the common cases, but it cannot expand what AIR
lacks (8- and 16-bit operations, nand, floating-point min/max), ignores
synchronization scopes, and only has ordered atomics from MSL 4.1. Do
what `AtomicExpand` and instruction selection do for other targets, so
that front-ends can emit plain LLVM atomics with an ordering and a scope:

- `validate_ir` rejects the atomics that cannot be lowered: outside
  device and threadgroup memory, wider than 32 bits (except non-fetching
  64-bit umin/umax on device memory), misaligned, with an unknown scope,
  or ordered below MSL 3.2;
- below MSL 4.1, ordered operations are relaxed ones bracketed with
  fences (AtomicExpand's default leading/trailing scheme);
- floating-point loads, stores and exchanges are cast to integers,
  8- and 16-bit operations become masked operations on the containing
  word, and operations AIR lacks become compare-exchange loops;
- the rest are selected to `air.atomic.*` calls, in the form of the
  target's AIR and MSL versions. `cmpxchg` becomes a weak compare-exchange
  with the success flag derived from the old value, as MSL does.

One legality function, `metal_atomic_action`, decides both: like the rule
tables of LLVM's legalizers, it returns how to lower an operation, or why
it cannot be, which `validate_ir` reports.

Front-ends that call `air.atomic.*` directly (e.g. to pass memory flags)
do so in the MSL 4.1 form, which is now legalized for older targets here
instead of in Metal.jl. The intrinsics are declared like Apple does
(`mustprogress nounwind willreturn`), no longer `argmemonly` and
`readonly`/`writeonly`, which would allow LLVM to move other memory
accesses across ordered atomics.
When `AllocOpt` moves an object with `@atomic` fields to the stack, it
keeps the compare-exchange loops on its fields (#934). MSL has no atomics
on thread memory, so the lowering rejected those, and before it Apple's
compiler handled them inconsistently: macOS 15's silently produced wrong
results. They don't need to be atomic, though: no thread can access
another's stack, so plain accesses behave the same, whatever the ordering
or scope. Lower atomics on pointers that are derived from allocas only to
plain loads and stores.
From Julia 1.13, `@atomic` modifications of fields are calls to
`julia.atomicmodify`, which AllocOpt treats as an escape, so the object in
the test stays on the heap and there are no atomics on the stack to lower.
The IR-level test of the lowering still covers every version.
…ites.

MSL 4.1 clears the volatile bit of atomics on non-`volatile` objects, and
we did the same for non-volatile LLVM atomics. Apple's compiler then
treats the atomic like a plain access: a relaxed load of memory the
kernel doesn't write is hoisted into the uniform preamble (out of a
spin loop), and a read-modify-write that leaves memory unchanged (adding
or or-ing 0, whatever the ordering) becomes a load. On the M1 (macOS 27),
a spin-wait on such a read-modify-write never observes the store another
threadgroup makes, while the volatile one does. LLVM atomics allow
neither transformation, so set the bit on loads and read-modify-writes,
as MSL does on every atomic before 4.1. Stores and compare-exchanges
compile the same either way, as does any read-modify-write that does
change memory, and the bit costs nothing measurable.
From MSL 4.1, we select an LLVM seq_cst atomic to the MSL 4.1 seq_cst
intrinsic, but Apple's compiler emits a seq_cst store as a release store
that doesn't wait for its write. A later seq_cst load can then be performed
before it: in the store-buffering litmus test (`store x; load y` against
`store y; load x`, across threadgroups), both loads return the old values
about 0.5% of the time on an M1, for LLVM atomics and MSL's own seq_cst
atomics alike. Follow seq_cst stores by a seq_cst fence, like AtomicExpand
does for targets whose `shouldInsertTrailingSeqCstFenceForAtomicStore`
returns true. Read-modify-writes and compare-exchanges wait for their
result and already behave; below MSL 4.1, the fences that bracket every
ordered operation already did.

On the M1, this removes all store-buffering violations (none in 6.5M
runs, also for the RMW and compare-exchange variants), and costs 4-11%
on a seq_cst-store-bound kernel. Fences before seq_cst loads work too,
but cost 50% on a load-bound one.
LLVM requires that a thread repeatedly loading an address monotonically
eventually sees the stores of other threads. MSL doesn't: device memory
is only coherent within a threadgroup by default, and relaxed atomics
only guarantee atomicity and a consistent modification order (MSL 4.1
§4.8, §6.16.1.1). Apple's compiler emits relaxed atomic loads of device
memory as cached loads, so on an M1 a spin loop that relaxed-loads a flag
another threadgroup sets never sees the store: not in 20M iterations, on
every MSL version, with MSL's own relaxed atomics too. (Within a
threadgroup it does, as the threadgroup shares the cache.) Acquire loads
invalidate the cache after loading, and the same spin loop exits after a
few hundred iterations. Lower device-scope monotonic loads of device
memory as acquire loads, which below MSL 4.1 means a trailing fence.

Loads at a narrower scope only need to see stores from their threadgroup,
so they are unaffected; front-ends that want MSL's semantics (and cost)
can use the workgroup scope. Relaxed loads of device memory are selected
at device scope regardless, like MSL does: the scope doesn't change the
code, but it is recorded.

This costs nothing when the loads miss the cache, but a kernel of
cache-resident device-scope relaxed loads runs 6x (MSL 4.1) to 7x
(older) slower. Only strengthening loads that may repeat on the same
address avoids part of that, but needs cycle analysis that isn't worth
its code.
@codecov

codecov Bot commented Sep 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.63492% with 22 lines in your changes missing coverage. Please review.
✅ Project coverage is 87.15%. Comparing base (78d617d) to head (082d689).

Files with missing lines Patch % Lines
src/atomics.jl 77.41% 14 Missing ⚠️
src/metal.jl 98.19% 8 Missing ⚠️
Additional details and impacted files
@@                      Coverage Diff                       @@
##           tb/remove-atomics-demotion     #942      +/-   ##
==============================================================
+ Coverage                       86.38%   87.15%   +0.77%     
==============================================================
  Files                              29       30       +1     
  Lines                            5794     6277     +483     
==============================================================
+ Hits                             5005     5471     +466     
- Misses                            789      806      +17     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[2.7.0] Metal: atomic RMW on thread-private memory is silently dropped

1 participant