Skip to content

Derive the autocast test backend and scope the bf16 gate to NCCL - #8559

Merged
delock merged 1 commit into
deepspeedai:masterfrom
delock:pr-f-autocast-cpu
Sep 17, 2026
Merged

delock merged 1 commit into
deepspeedai:masterfrom
delock:pr-f-autocast-cpu

Conversation

@delock

@delock delock commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Description

bf16_required_version_check() (tests/unit/util.py) requires torch >= 1.10, CUDA >= 11.0 and NCCL >= 2.10.3. On the cpu accelerator, bf16 collectives run over gloo/ccl and none of those transport dependencies exist, so the check always returns False and every bf16 test is skipped — about 40 call sites across 15 files. test_zero_autocast.py is worse off: it raises instead of skipping, so each of its cases counts as a failure (24 on the multi-rank CPU run in #8381).

This PR scopes the version floors inside the check itself:

if (cpu_accelerator and accelerator_pass) or (torch_version_available and cuda_version_available
                                              and nccl_version_available and accelerator_pass):
    return True
  • cpu: only the accelerator's own bf16 support (is_bf16_supported()) is required — the torch/CUDA/NCCL floors are transport dependencies that gloo/ccl does not have
  • every other accelerator (cuda, npu, hpu, xpu, mlu, ...): evaluates the exact original floors; cpu_accelerator is False so the expression is bit-identical to before, and the npu/hpu/xpu exemption branches are untouched
  • test_zero_autocast.py: the bf16 gate goes back to the bare call every other caller uses, and skips instead of raising

The hardcoded init_distributed(dist_backend='nccl') in the same test is deliberately left alone: it is a no-op (the harness already initialized the process group, comm.py:838-839), and deriving the backend per accelerator would change behavior on non-cuda accelerators (npu/hpu resolve to hccl etc.). The baseline DDP(device_ids=[i]) pinning is left for a follow-up (#8399 fixed the same pattern elsewhere).

Validation (executed on real hardware)

  • 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend
  • Direct call: bf16_required_version_check() on cpu returns False before, True after. Note: CPU_Accelerator.is_bf16_supported() is currently a stub that always returns True, so the cpu path does not gate on the hardware's bf16 instructions — giving it a real capability probe is left as a follow-up
  • Non-cpu equivalence: with cpu_accelerator == False the new expression reduces exactly to the original A and C and N and P
  • pre-commit (yapf / flake8 / check-torchdist / codespell) passes on both changed files

Sibling PRs from the same series: #8397, #8398, #8399, #8407, #8409. Exposed by the LOCAL_SIZE=4 multi-rank CPU run in #8381.

bf16_required_version_check() requires torch >= 1.10, CUDA >= 11 and
NCCL >= 2.10.3, so on the cpu accelerator, where bf16 collectives run
over gloo/ccl and none of those dependencies exist, the check always
fails and every bf16 test is skipped, about 40 call sites across 15
files. test_zero_autocast is worse off: it raises instead of skipping,
so each case counts as a failure.

Exempt the cpu accelerator inside the check itself so all call sites
benefit: on cpu only the accelerator's own bf16 support is required.
Every other accelerator evaluates the exact original floors, and
test_zero_autocast now skips like every other caller instead of
raising.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
@sfc-gh-truwase
sfc-gh-truwase added this pull request to the merge queue Sep 17, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 17, 2026
@delock
delock added this pull request to the merge queue Sep 17, 2026
Merged via the queue into deepspeedai:master with commit b7b197e Sep 17, 2026
13 checks passed
@delock
delock deleted the pr-f-autocast-cpu branch September 17, 2026 16:21
yermakoffivan pushed a commit to yermakoffivan/deepspeed that referenced this pull request Sep 28, 2026
…dai#8684)

## Description

Every fp16-config test that reaches `deepspeed.initialize` crashes its
sanity check (`Type fp16 is not supported on your device.`) on
accelerators whose `is_fp16_supported()` is false. On CPU that maps to
the AVX512-FP16 capability of the host, and GitHub's `ubuntu-24.04`
runners are hardware-heterogeneous, so these tests flip between failure
and skip depending on which runner they land on. deepspeedai#8398 added the first
skipifs; the multi-rank CPU run in deepspeedai#8381 flushed out six more files:

| File | Shape of the gap |
|---|---|
| `checkpoint/test_universal_checkpoint.py` | fp16 parametrizations fail
**inside the baseline DistributedFixture's distributed run** — pytest
reports a setup **ERROR** for every dependent test (48 on the multi-rank
run) instead of a skip |
| `v1/zero/test_zero_coalesce_grad_reduction.py` | `TestCoalesceFP16`
forces an fp16 config |
| `runtime/test_no_sync_ctxt.py` | dtype=float16 parametrizations of
three methods; stages 2/3 additionally never reach their expected
no_sync AssertionError on such hosts |
| `checkpoint/test_moe_checkpoint.py` | whole class hardcodes fp16 |
| `runtime/zero/test_zero_offloadpp.py` |
`TestZeroPartialOffloadConfigSweep` hardcodes fp16 |
| `checkpoint/test_pipeline.py` | fp16 enabled for zero_stage > 0; only
that parametrization skips, zero_stage=0 keeps running |

With this PR, **every fp16-config test in the suite guards on
`is_fp16_supported()`** — the capability gap is a skip, not a failure,
on any accelerator.

## Validation (executed on real hardware)

- 20-core x86_64 CPU without AVX512-FP16, gloo, 2-4 ranks
(`LOCAL_SIZE=2/4`)
- Before: 18 failures + 48 setup ERRORs across these files on the
multi-rank CPU run
- After: every affected parametrization skips; adjacent non-fp16
parametrizations keep passing (e.g. pipeline zero_stage=0 runs to
completion)
- Full-suite evidence in deepspeedai#8381: the multi-rank CPU run went 8 failures
-> 0 with these guards

Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409,
deepspeedai#8559, deepspeedai#8648.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants