Skip to content

Fix ballot on wave64 GPUs and with GPUCompiler 2.9 - #1112

Merged
luraess merged 2 commits into
mainfrom
lr/ballot
Sep 27, 2026
Merged

luraess merged 2 commits into
mainfrom
lr/ballot

Conversation

@luraess

@luraess luraess commented Sep 26, 2026

Copy link
Copy Markdown
Member

The Documentation job has failed on main since build 4275, and so on every PR since. The four doctests that use ballot fail with "unsupported call to an unknown LLVM intrinsic". The only difference between the last green and the first red docs environment is GPUCompiler 2.8.2 → 2.9.0, which reports llvmcalls of unknown intrinsics at compile time (JuliaGPU/GPUCompiler.jl#958). This is not only a docs problem: ballot, activemask, ballot_sync, any_sync and all_sync all go through ballot, which had two bugs:

  • The wave64 branch called llvm.amdgcn.ballot.w64. That is not a valid name for the overloaded intrinsic, whose suffix is the result type (.i64). Julia compiles the llvmcall to jl_error("llvmcall only supports intrinsic calls"), and GPUCompiler 2.9.0 now rejects that on every target, wave32 included. The branches now use .i64 and .i32.
  • LLVM doesn't fold llvm.amdgcn.wavefrontsize before instruction selection, so the backend also selects the branch that never runs. On wave64, the i32 ballot can't be selected (Cannot select: i32 = AMDGPUISD::SETCC), so ballot has never compiled on MI250/MI300, with any GPUCompiler. finish_module! now folds llvm.amdgcn.wavefrontsize to the job's wavefront size, so that branch goes away before optimization. shfl, random and the hostcall and printing code also get it as a constant now.

Only the doctests covered ballot, and they run on the wave32 Buildkite GPUs, where the broken branch was dead code. The new "Wavefront Ballot" testset in device/wavefront runs ballot at the native wavefront size and in wave64 mode, so it covers both branches on wave32 GPUs too.

ballot over every other lane, on an MI250 (gfx90a, Julia 1.12). gfx1100 was only compiled, not run:

main, GPUCompiler 2.8.2 main, GPUCompiler 2.9.0 This PR, GPUCompiler 2.9.0
gfx90a (wave64), run Cannot select unknown intrinsic 0x5555555555555555
gfx1100 wave32, compiled ok unknown intrinsic ok
gfx1100 wave64, compiled Cannot select unknown intrinsic ok

The full suite passes on the MI250 (2 GCDs, Julia 1.12, GPUCompiler 2.9.0): 17149 passed, 21 broken.

🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMDGPU.jl Benchmarks

Details
Benchmark suite Current: 0840adf Previous: 8a66547 Ratio
amdgpu/synchronization/context/device 552.5 ns 550 ns 1.00
amdgpu/synchronization/stream/blocking 230 ns 222.5 ns 1.03
amdgpu/synchronization/stream/nonblocking 305 ns 307.5 ns 0.99
applications/bitonic_sort 1132481.25 ns 1160103.75 ns 0.98
applications/convolution 105429 ns 72376 ns 1.46
applications/floyd_warshall 9014979.25 ns 9100856.25 ns 0.99
applications/histogram 820511.75 ns 798998.75 ns 1.03
applications/prefix_sum 187982.75 ns 190535 ns 0.99
array/accumulate/Float32/1d 82093.75 ns 79346 ns 1.03
array/accumulate/Float32/dims=1 273718.75 ns 271536.25 ns 1.01
array/accumulate/Float32/dims=1L 107561.5 ns 72536 ns 1.48
array/accumulate/Float32/dims=2 102481.5 ns 74041 ns 1.38
array/accumulate/Float32/dims=2L 3031194 ns 2834326.75 ns 1.07
array/accumulate/Int64/1d 91161.25 ns 79098.5 ns 1.15
array/accumulate/Int64/dims=1 249486.25 ns 248393.25 ns 1.00
array/accumulate/Int64/dims=1L 115196.5 ns 89128.75 ns 1.29
array/accumulate/Int64/dims=2 98814 ns 76041 ns 1.30
array/accumulate/Int64/dims=2L 3389189 ns 3139866 ns 1.08
array/broadcast 49500.5 ns 50443.25 ns 0.98
array/construct 2097.5 ns 2247.5 ns 0.93
array/copy 39135.5 ns 39278 ns 1.00
array/copyto!/cpu_to_gpu 89113.75 ns 85676.25 ns 1.04
array/copyto!/gpu_to_cpu 89573.75 ns 85591.25 ns 1.05
array/copyto!/gpu_to_gpu 34480.25 ns 58288.25 ns 0.59
array/iteration/findall/bool 129201.75 ns 125954.25 ns 1.03
array/iteration/findall/int 146174.5 ns 146617 ns 1.00
array/iteration/findfirst/bool 157204.75 ns 155292.25 ns 1.01
array/iteration/findfirst/int 191912.75 ns 189260 ns 1.01
array/iteration/findmin/1d 103654 ns 103258.75 ns 1.00
array/iteration/findmin/2d 90551.25 ns 90081.25 ns 1.01
array/iteration/logical 227973.25 ns 226565.5 ns 1.01
array/iteration/scalar 307657 ns 295931.5 ns 1.04
array/permutedims/2d 47783.25 ns 70736 ns 0.68
array/permutedims/3d 57961 ns 69253.5 ns 0.84
array/permutedims/4d 74281 ns 50278.25 ns 1.48
array/random/rand/Float32 44853.25 ns 40065.75 ns 1.12
array/random/rand/Int64 54655.75 ns 48863 ns 1.12
array/random/rand!/Float32 39685.75 ns 65541 ns 0.61
array/random/rand!/Int64 49058 ns 69623.25 ns 0.70
array/random/randn/Float32 70631.25 ns 71161 ns 0.99
array/random/randn!/Float32 57285.75 ns 58721 ns 0.98
array/reductions/mapreduce/Float32/1d 82101 ns 81621 ns 1.01
array/reductions/mapreduce/Float32/dims=1 82861.25 ns 56660.75 ns 1.46
array/reductions/mapreduce/Float32/dims=1L 838612 ns 844514.25 ns 0.99
array/reductions/mapreduce/Float32/dims=2 81718.75 ns 74073.5 ns 1.10
array/reductions/mapreduce/Float32/dims=2L 136354.25 ns 136524.5 ns 1.00
array/reductions/mapreduce/Int64/1d 82266 ns 81663.5 ns 1.01
array/reductions/mapreduce/Int64/dims=1 77741.25 ns 73933.5 ns 1.05
array/reductions/mapreduce/Int64/dims=1L 842112.25 ns 839174.25 ns 1.00
array/reductions/mapreduce/Int64/dims=2 80413.5 ns 75486.25 ns 1.07
array/reductions/mapreduce/Int64/dims=2L 137434.5 ns 137611.75 ns 1.00
array/reductions/reduce/Float32/1d 81826.25 ns 81541.25 ns 1.00
array/reductions/reduce/Float32/dims=1 83023.75 ns 70296 ns 1.18
array/reductions/reduce/Float32/dims=1L 839754.5 ns 836984.25 ns 1.00
array/reductions/reduce/Float32/dims=2 81046.25 ns 75123.5 ns 1.08
array/reductions/reduce/Float32/dims=2L 136959.5 ns 136774.25 ns 1.00
array/reductions/reduce/Int64/1d 78481.25 ns 81368.75 ns 0.96
array/reductions/reduce/Int64/dims=1 85493.75 ns 74831 ns 1.14
array/reductions/reduce/Int64/dims=1L 844699.5 ns 844554 ns 1.00
array/reductions/reduce/Int64/dims=2 81663.75 ns 75411.25 ns 1.08
array/reductions/reduce/Int64/dims=2L 137884.25 ns 137047 ns 1.01
array/reverse/1d 42298.25 ns 43458 ns 0.97
array/reverse/1dL 54366 ns 70606 ns 0.77
array/reverse/1dL_inplace 55555.75 ns 79671 ns 0.70
array/reverse/1d_inplace 36373 ns 60390.75 ns 0.60
array/reverse/2d 47145.75 ns 47625.75 ns 0.99
array/reverse/2dL 90248.75 ns 57993.5 ns 1.56
array/reverse/2dL_inplace 66211 ns 79096 ns 0.84
array/reverse/2d_inplace 36985.5 ns 61191 ns 0.60
array/sorting/1d 327267.25 ns 326449.5 ns 1.00
gemm/tiled 1958523.25 ns 1961659.75 ns 1.00
gemm/tiled_unbounded 1974238.25 ns 1960334.75 ns 1.01
integration/byval/reference 39901 ns 39821 ns 1.00
integration/byval/slices=1 39931 ns 40830 ns 0.98
integration/byval/slices=2 147002 ns 130842 ns 1.12
integration/byval/slices=3 239803 ns 241884 ns 0.99
integration/volumerhs 4893940 ns 4900148 ns 1.00
kernel/indexing 29640.5 ns 41880.5 ns 0.71
kernel/indexing_checked 29840.5 ns 37075.5 ns 0.80
kernel/launch 1170 ns 1130 ns 1.04
kernel/rand 98979 ns 70343.5 ns 1.41
latency/import 1449711665 ns 1456646355 ns 1.00
latency/precompile 22905651598 ns 22642954217 ns 1.01
latency/ttfp 2201433275 ns 2251455325 ns 0.98
stencil/diffusion3d 1619138.25 ns 1611230 ns 1.00
stencil/diffusion3d_checked 1658591.5 ns 1655470.5 ns 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@luraess
luraess merged commit 6612a33 into main Sep 27, 2026
6 of 7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant