Derive the autocast test backend and scope the bf16 gate to NCCL - #8559
Merged
Merged
Conversation
delock
force-pushed
the
pr-f-autocast-cpu
branch
from
September 17, 2026 00:32
71e6b84 to
c162ba4
Compare
bf16_required_version_check() requires torch >= 1.10, CUDA >= 11 and NCCL >= 2.10.3, so on the cpu accelerator, where bf16 collectives run over gloo/ccl and none of those dependencies exist, the check always fails and every bf16 test is skipped, about 40 call sites across 15 files. test_zero_autocast is worse off: it raises instead of skipping, so each case counts as a failure. Exempt the cpu accelerator inside the check itself so all call sites benefit: on cpu only the accelerator's own bf16 support is required. Every other accelerator evaluates the exact original floors, and test_zero_autocast now skips like every other caller instead of raising. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
delock
force-pushed
the
pr-f-autocast-cpu
branch
from
September 17, 2026 00:39
c162ba4 to
baf0caa
Compare
sfc-gh-truwase
approved these changes
Sep 17, 2026
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Sep 17, 2026
yermakoffivan
pushed a commit
to yermakoffivan/deepspeed
that referenced
this pull request
Sep 28, 2026
…dai#8684) ## Description Every fp16-config test that reaches `deepspeed.initialize` crashes its sanity check (`Type fp16 is not supported on your device.`) on accelerators whose `is_fp16_supported()` is false. On CPU that maps to the AVX512-FP16 capability of the host, and GitHub's `ubuntu-24.04` runners are hardware-heterogeneous, so these tests flip between failure and skip depending on which runner they land on. deepspeedai#8398 added the first skipifs; the multi-rank CPU run in deepspeedai#8381 flushed out six more files: | File | Shape of the gap | |---|---| | `checkpoint/test_universal_checkpoint.py` | fp16 parametrizations fail **inside the baseline DistributedFixture's distributed run** — pytest reports a setup **ERROR** for every dependent test (48 on the multi-rank run) instead of a skip | | `v1/zero/test_zero_coalesce_grad_reduction.py` | `TestCoalesceFP16` forces an fp16 config | | `runtime/test_no_sync_ctxt.py` | dtype=float16 parametrizations of three methods; stages 2/3 additionally never reach their expected no_sync AssertionError on such hosts | | `checkpoint/test_moe_checkpoint.py` | whole class hardcodes fp16 | | `runtime/zero/test_zero_offloadpp.py` | `TestZeroPartialOffloadConfigSweep` hardcodes fp16 | | `checkpoint/test_pipeline.py` | fp16 enabled for zero_stage > 0; only that parametrization skips, zero_stage=0 keeps running | With this PR, **every fp16-config test in the suite guards on `is_fp16_supported()`** — the capability gap is a skip, not a failure, on any accelerator. ## Validation (executed on real hardware) - 20-core x86_64 CPU without AVX512-FP16, gloo, 2-4 ranks (`LOCAL_SIZE=2/4`) - Before: 18 failures + 48 setup ERRORs across these files on the multi-rank CPU run - After: every affected parametrization skips; adjacent non-fp16 parametrizations keep passing (e.g. pipeline zero_stage=0 runs to completion) - Full-suite evidence in deepspeedai#8381: the multi-rank CPU run went 8 failures -> 0 with these guards Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409, deepspeedai#8559, deepspeedai#8648. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
bf16_required_version_check()(tests/unit/util.py) requires torch >= 1.10, CUDA >= 11.0 and NCCL >= 2.10.3. On the cpu accelerator, bf16 collectives run over gloo/ccl and none of those transport dependencies exist, so the check always returns False and every bf16 test is skipped — about 40 call sites across 15 files.test_zero_autocast.pyis worse off: it raises instead of skipping, so each of its cases counts as a failure (24 on the multi-rank CPU run in #8381).This PR scopes the version floors inside the check itself:
is_bf16_supported()) is required — the torch/CUDA/NCCL floors are transport dependencies that gloo/ccl does not havecpu_acceleratoris False so the expression is bit-identical to before, and the npu/hpu/xpu exemption branches are untouchedtest_zero_autocast.py: the bf16 gate goes back to the bare call every other caller uses, and skips instead of raisingThe hardcoded
init_distributed(dist_backend='nccl')in the same test is deliberately left alone: it is a no-op (the harness already initialized the process group,comm.py:838-839), and deriving the backend per accelerator would change behavior on non-cuda accelerators (npu/hpu resolve to hccl etc.). The baselineDDP(device_ids=[i])pinning is left for a follow-up (#8399 fixed the same pattern elsewhere).Validation (executed on real hardware)
bf16_required_version_check()on cpu returnsFalsebefore,Trueafter. Note:CPU_Accelerator.is_bf16_supported()is currently a stub that always returns True, so the cpu path does not gate on the hardware's bf16 instructions — giving it a real capability probe is left as a follow-upcpu_accelerator == Falsethe new expression reduces exactly to the originalA and C and N and PSibling PRs from the same series: #8397, #8398, #8399, #8407, #8409. Exposed by the
LOCAL_SIZE=4multi-rank CPU run in #8381.