Skip to content

Test @fastmath integer powers - #978

Merged
maleadt merged 1 commit into
mainfrom
tb/fastmath-powi
Sep 26, 2026
Merged

maleadt merged 1 commit into
mainfrom
tb/fastmath-powi

Conversation

@maleadt

@maleadt maleadt commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

@fastmath x^n with an integer n emits llvm.powi, which Apple's back-end can't compile, so the kernel fails with "Compilation to native code failed". JuliaGPU/GPUCompiler.jl#946 expands the intrinsic into multiplies, in the same order the CPU uses.

This PR bumps GPUCompiler to 2.9 and adds an integration test. The test compares @fastmath x^n on the GPU with the CPU result, bit for bit, for special bases (±0, ±Inf, NaN) and exponents including typemin/typemax(Int32). It covers run-time and constant exponents, plus Float16. With the current GPUCompiler it fails with the compilation error above.

@maleadt maleadt changed the title Test @fastmath integer powers Test @fastmath integer powers Sep 25, 2026
@maleadt
maleadt marked this pull request as ready for review September 26, 2026 06:29
`@fastmath x^n` with an integer `n` emits `llvm.powi`, which Apple's back-end
cannot compile. GPUCompiler 2.9 expands it into multiplies; check that the
results match the CPU bit for bit.
@codecov

codecov Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 86.47%. Comparing base (26c9e7d) to head (162dfdf).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #978      +/-   ##
==========================================
+ Coverage   85.37%   86.47%   +1.10%     
==========================================
  Files          77       92      +15     
  Lines        5660     6743    +1083     
==========================================
+ Hits         4832     5831     +999     
- Misses        828      912      +84     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Metal Benchmarks

Details
Benchmark suite Current: 162dfdf Previous: 26c9e7d Ratio
array/accumulate/Float32/1d 391000 ns 393834 ns 0.99
array/accumulate/Float32/dims=1 364292 ns 368959 ns 0.99
array/accumulate/Float32/dims=1L 8767416 ns 8770667 ns 1.00
array/accumulate/Float32/dims=2 432208 ns 429083 ns 1.01
array/accumulate/Float32/dims=2L 2530084 ns 2531958 ns 1.00
array/accumulate/Int64/1d 844166 ns 813000 ns 1.04
array/accumulate/Int64/dims=1 908125 ns 910333 ns 1.00
array/accumulate/Int64/dims=1L 9437250 ns 9446250 ns 1.00
array/accumulate/Int64/dims=2 1207833 ns 1210125 ns 1.00
array/accumulate/Int64/dims=2L 6424583 ns 6431042 ns 1.00
array/broadcast 239125 ns 238917 ns 1.00
array/construct 2333 ns 2292 ns 1.02
array/permutedims/2d 450541 ns 435125 ns 1.04
array/permutedims/3d 1032625 ns 997375 ns 1.04
array/permutedims/4d 1014625 ns 1092083 ns 0.93
array/private/copy 227000 ns 222750 ns 1.02
array/private/copyto!/cpu_to_gpu 213916 ns 212875 ns 1.00
array/private/copyto!/gpu_to_cpu 212209 ns 213375 ns 0.99
array/private/copyto!/gpu_to_gpu 215875 ns 211208 ns 1.02
array/private/iteration/findall/bool 1050000 ns 1027583 ns 1.02
array/private/iteration/findall/int 1213875 ns 1210334 ns 1.00
array/private/iteration/findfirst/bool 1148834 ns 1144291 ns 1.00
array/private/iteration/findfirst/int 1161125 ns 1151333 ns 1.01
array/private/iteration/findmin/1d 1205667 ns 1200334 ns 1.00
array/private/iteration/findmin/2d 1029625 ns 1023958 ns 1.01
array/private/iteration/logical 1653208 ns 1621875 ns 1.02
array/private/iteration/scalar 1388334 ns 1357334 ns 1.02
array/random/rand/Float32 415042 ns 418625 ns 0.99
array/random/rand/Int64 492833 ns 491042 ns 1.00
array/random/rand!/Float32 403584 ns 400584 ns 1.01
array/random/rand!/Int64 429958 ns 426750 ns 1.01
array/random/randn/Float32 385333 ns 387583 ns 0.99
array/random/randn!/Float32 373958 ns 372458 ns 1.00
array/reductions/mapreduce/Float32/1d 442625 ns 448792 ns 0.99
array/reductions/mapreduce/Float32/dims=1 300666 ns 348750 ns 0.86
array/reductions/mapreduce/Float32/dims=1L 608959 ns 612458 ns 0.99
array/reductions/mapreduce/Float32/dims=2 356458 ns 355000 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 1074250 ns 1064708 ns 1.01
array/reductions/mapreduce/Int64/1d 628959 ns 626667 ns 1.00
array/reductions/mapreduce/Int64/dims=1 633667 ns 457375 ns 1.39
array/reductions/mapreduce/Int64/dims=1L 1011875 ns 1007000 ns 1.00
array/reductions/mapreduce/Int64/dims=2 763125 ns 791542 ns 0.96
array/reductions/mapreduce/Int64/dims=2L 2192041 ns 2208042 ns 0.99
array/reductions/reduce/Float32/1d 443167 ns 448000 ns 0.99
array/reductions/reduce/Float32/dims=1 345292 ns 357291 ns 0.97
array/reductions/reduce/Float32/dims=1L 615250 ns 616542 ns 1.00
array/reductions/reduce/Float32/dims=2 251833 ns 245792 ns 1.02
array/reductions/reduce/Float32/dims=2L 471834 ns 476459 ns 0.99
array/reductions/reduce/Int64/1d 631375 ns 632917 ns 1.00
array/reductions/reduce/Int64/dims=1 631542 ns 453375 ns 1.39
array/reductions/reduce/Int64/dims=1L 1005916 ns 1010625 ns 1.00
array/reductions/reduce/Int64/dims=2 262167 ns 264125 ns 0.99
array/reductions/reduce/Int64/dims=2L 668250 ns 671250 ns 1.00
array/shared/copy 130167 ns 130250 ns 1.00
array/shared/copyto!/cpu_to_gpu 38042 ns 37500 ns 1.01
array/shared/copyto!/gpu_to_cpu 36833 ns 38500 ns 0.96
array/shared/copyto!/gpu_to_gpu 37292 ns 38875 ns 0.96
array/shared/iteration/findall/bool 1055625 ns 1026708 ns 1.03
array/shared/iteration/findall/int 1221959 ns 1209666 ns 1.01
array/shared/iteration/findfirst/bool 924334 ns 966125 ns 0.96
array/shared/iteration/findfirst/int 987417 ns 979125 ns 1.01
array/shared/iteration/findmin/1d 1058625 ns 1064333 ns 0.99
array/shared/iteration/findmin/2d 1032667 ns 1030625 ns 1.00
array/shared/iteration/logical 1518708 ns 1376500 ns 1.10
array/shared/iteration/scalar 3864.5 ns 3744.75 ns 1.03
array/sorting/1d 2087167 ns 2081667 ns 1.00
array/sorting/2d 8296000 ns 8338750 ns 0.99
integration/byval/reference 1119125 ns 1123500 ns 1.00
integration/byval/slices=1 1121833 ns 1120292 ns 1.00
integration/byval/slices=2 2022334 ns 2027333 ns 1.00
integration/byval/slices=3 6635458 ns 6583833 ns 1.01
integration/metaldevrt 392750 ns 373250 ns 1.05
kernel/indexing 156625 ns 212750 ns 0.74
kernel/indexing_checked 398750 ns 391792 ns 1.02
kernel/launch 1820.8 ns 1804.2 ns 1.01
kernel/rand 398416 ns 404541 ns 0.98
latency/import 1789705375 ns 1786979292 ns 1.00
latency/precompile 32875539792 ns 32696140792 ns 1.01
latency/ttfp 2268129084 ns 2268758291 ns 1.00
metal/synchronization/context 557.5752688172043 ns 545.1957671957672 ns 1.02
metal/synchronization/stream 345.2916666666667 ns 350.5492957746479 ns 0.99

This comment was automatically generated by workflow using github-action-benchmark.

@maleadt
maleadt merged commit 80a2aac into main Sep 26, 2026
19 checks passed
@maleadt
maleadt deleted the tb/fastmath-powi branch September 26, 2026 08:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant