Skip fp16-config tests on accelerators without fp16 support - #8398
Conversation
test_multiple_models/test_zero_optimizer and the MoE tests hardcode
"fp16": {"enabled": True} in their DeepSpeed configs. The engine's
sanity check then raises "Type fp16 is not supported on your device"
on any accelerator whose is_fp16_supported() is false - on CPU that
depends on the host's AVX512-FP16 capability, and GitHub's ubuntu-24.04
runners are heterogeneous enough that the same test passes on one
runner and fails on the next.
Skip these tests via a capability query instead of depending on runner
hardware, mirroring the existing bf16 skip precedent in
test_zero_user_backward.
Validated as part of the multi-rank CPU experiment in deepspeedai#8381: 146
failures of this class became deterministic skips, zero regressions.
Signed-off-by: Guokai Ma <guokai.ma@intel.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 31e794e4ed
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| import deepspeed | ||
| import deepspeed.comm as dist | ||
| import torch | ||
| from deepspeed import get_accelerator |
There was a problem hiding this comment.
Add the required Signed-off-by trailer
This is a non-merge commit, but its message has no Signed-off-by trailer. Add the author name and email from the Git configuration using git commit --signoff so the commit satisfies the repository's mandatory DCO/CI requirement.
AGENTS.md reference: AGENTS.md:L8-L8
Useful? React with 👍 / 👎.
…epspeedai#8407) ## Description `train_cifar` unconditionally passed `devices=[get_accelerator().current_device_name()]` to `fork_rng`. On CPU that becomes `devices=['cpu']`, and `torch.cpu` has no `get_rng_state`, so the context manager raises `AttributeError` before the test body starts — the CPU RNG lives in the global generator that `fork_rng` already forks. Probe the device module for a per-device RNG instead of matching accelerator names: - backends whose device module has `get_rng_state` (e.g. cuda) behave exactly as before; - backends without one (cpu) pass `devices=[]`, which changes nothing beyond the global-RNG save/restore `fork_rng` always performs. The crash is only reachable from the multi-rank tests that call `train_cifar` (`test_onebit.py`, `test_pipe.py`), which the CPU runner currently skips at the device gate; it was exposed by the `LOCAL_SIZE=4` experiment in deepspeedai#8381. Sibling fixes from the same series landed as deepspeedai#8397, deepspeedai#8398 and deepspeedai#8399. ## Test plan - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the changed file; - exercised under the `LOCAL_SIZE=4` multi-rank CPU run tracked in deepspeedai#8381: the `test_onebit` / `test_pipe` callers got past `fork_rng` and into their test bodies with this exact change. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…peedai#8409) ## Description The dynamic offload-state tests assert strict allocated-memory deltas around `offload_states()` / `reload_states()`: - `alloc_after_offload < alloc_before_offload` - `alloc_after_reload > alloc_after_offload` That contract assumes `memory_allocated()` is allocator bookkeeping, which holds on cuda (`torch.cuda.memory_allocated()`). On cpu, `CPU_Accelerator.memory_allocated()` reports process RSS (psutil), and RSS does not shrink when tensors are freed — so all 92 parameterized cases fail even when the offload itself is correct (the device-placement and data-integrity checks in the same tests pass). Gate only the memory-delta asserts on whether the accelerator's torch device module exposes `memory_allocated` (cuda does; `torch.cpu` does not), mirroring the capability probe used for `fork_rng` in `train_cifar` (deepspeedai#8407): - cuda and other allocator-backed backends: behavior unchanged - cpu: the unobservable deltas are skipped; all device-placement validations still run Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381 (92 of the 131 v1-half failures there). ## Validation (executed on real hardware) - 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend, 2 ranks (`LOCAL_SIZE=2`) - Before: `TestDynamicOffloadStatesZero12[False-1-False-False-optim_states]` fails on the persistent-state delta assert - After: 5 representative cases pass (persistent and grad paths, ZeRO stage 1/2/3, `static_offload_optimizer=True` branch) — 5 passed in 55.6s - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the changed file Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…pspeedai#8559) ## Description `bf16_required_version_check()` (`tests/unit/util.py`) requires torch >= 1.10, **CUDA >= 11.0 and NCCL >= 2.10.3**. On the cpu accelerator, bf16 collectives run over gloo/ccl and none of those transport dependencies exist, so the check always returns False and **every bf16 test is skipped — about 40 call sites across 15 files**. `test_zero_autocast.py` is worse off: it **raises** instead of skipping, so each of its cases counts as a failure (24 on the multi-rank CPU run in deepspeedai#8381). This PR scopes the version floors inside the check itself: ```python if (cpu_accelerator and accelerator_pass) or (torch_version_available and cuda_version_available and nccl_version_available and accelerator_pass): return True ``` - **cpu**: only the accelerator's own bf16 support (`is_bf16_supported()`) is required — the torch/CUDA/NCCL floors are transport dependencies that gloo/ccl does not have - **every other accelerator** (cuda, npu, hpu, xpu, mlu, ...): evaluates the exact original floors; `cpu_accelerator` is False so the expression is bit-identical to before, and the npu/hpu/xpu exemption branches are untouched - `test_zero_autocast.py`: the bf16 gate goes back to the bare call every other caller uses, and skips instead of raising The hardcoded `init_distributed(dist_backend='nccl')` in the same test is deliberately left alone: it is a no-op (the harness already initialized the process group, `comm.py:838-839`), and deriving the backend per accelerator would change behavior on non-cuda accelerators (npu/hpu resolve to hccl etc.). The baseline `DDP(device_ids=[i])` pinning is left for a follow-up (deepspeedai#8399 fixed the same pattern elsewhere). ## Validation (executed on real hardware) - 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend - Direct call: `bf16_required_version_check()` on cpu returns `False` before, `True` after. Note: `CPU_Accelerator.is_bf16_supported()` is currently a stub that always returns True, so the cpu path does not gate on the hardware's bf16 instructions — giving it a real capability probe is left as a follow-up - Non-cpu equivalence: with `cpu_accelerator == False` the new expression reduces exactly to the original `A and C and N and P` - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on both changed files Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409. Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The baseline DistributedFixture builds an fp16 engine config for every dtype=float16 parametrization, and deepspeed.initialize's sanity check raises "Type fp16 is not supported on your device." on accelerators without fp16 support. Because the failure happens inside the fixture's distributed run, pytest reports a setup ERROR for every dependent test (48 on the multi-rank CPU run) instead of a skip. Skip inside the fixture like deepspeedai#8398 did for the fp16-config tests; DistributedFixture propagates the skip to all dependent tests. Devices with fp16 support are unchanged. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
TestCoalesceFP16 forces an fp16 config, and deepspeed.initialize's sanity check raises "Type fp16 is not supported on your device." on accelerators without fp16 support — the same hardware lottery deepspeedai#8398 handled for the other fp16-config tests. Add the same class-level skipif so the gap is a skip, not three failures. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The dtype=float16 parametrizations of TestNoSyncCtxt crash
deepspeed.initialize's sanity check ("Type fp16 is not supported on
your device.") on accelerators without fp16 support — the same
hardware lottery deepspeedai#8398 handled elsewhere. Guard the three
dtype-parametrized methods so the gap is a skip, not eight failures;
stages 2/3 additionally never reach their expected no_sync
AssertionError on such hosts.
Verified locally on an fp16-incapable CPU: 8 fail -> 8 skip, all 17
remaining parametrizations pass.
Signed-off-by: Guokai Ma <guokai.ma@intel.com>
TestMoECheckpoint hardcodes fp16 configs, so deepspeed.initialize's sanity check fails on accelerators without fp16 support — the same hardware lottery deepspeedai#8398 handled elsewhere. Add the class-level skipif. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
TestZeroPartialOffloadConfigSweep hardcodes fp16 and test_checkpoint_pipe_engine enables it for zero_stage > 0; both trip initialize's sanity check on accelerators without fp16 support (deepspeedai#8398's hardware lottery). Skip those parametrizations so the gap is a skip. With this, every fp16-config test in the suite guards on is_fp16_supported(). Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…dai#8684) ## Description Every fp16-config test that reaches `deepspeed.initialize` crashes its sanity check (`Type fp16 is not supported on your device.`) on accelerators whose `is_fp16_supported()` is false. On CPU that maps to the AVX512-FP16 capability of the host, and GitHub's `ubuntu-24.04` runners are hardware-heterogeneous, so these tests flip between failure and skip depending on which runner they land on. deepspeedai#8398 added the first skipifs; the multi-rank CPU run in deepspeedai#8381 flushed out six more files: | File | Shape of the gap | |---|---| | `checkpoint/test_universal_checkpoint.py` | fp16 parametrizations fail **inside the baseline DistributedFixture's distributed run** — pytest reports a setup **ERROR** for every dependent test (48 on the multi-rank run) instead of a skip | | `v1/zero/test_zero_coalesce_grad_reduction.py` | `TestCoalesceFP16` forces an fp16 config | | `runtime/test_no_sync_ctxt.py` | dtype=float16 parametrizations of three methods; stages 2/3 additionally never reach their expected no_sync AssertionError on such hosts | | `checkpoint/test_moe_checkpoint.py` | whole class hardcodes fp16 | | `runtime/zero/test_zero_offloadpp.py` | `TestZeroPartialOffloadConfigSweep` hardcodes fp16 | | `checkpoint/test_pipeline.py` | fp16 enabled for zero_stage > 0; only that parametrization skips, zero_stage=0 keeps running | With this PR, **every fp16-config test in the suite guards on `is_fp16_supported()`** — the capability gap is a skip, not a failure, on any accelerator. ## Validation (executed on real hardware) - 20-core x86_64 CPU without AVX512-FP16, gloo, 2-4 ranks (`LOCAL_SIZE=2/4`) - Before: 18 failures + 48 setup ERRORs across these files on the multi-rank CPU run - After: every affected parametrization skips; adjacent non-fp16 parametrizations keep passing (e.g. pipeline zero_stage=0 runs to completion) - Full-suite evidence in deepspeedai#8381: the multi-rank CPU run went 8 failures -> 0 with these guards Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409, deepspeedai#8559, deepspeedai#8648. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Problem
TestMultipleModels::test_zero_optimizer,TestSimpleMoE,TestMoE,TestPRMoE, andTestMOETensorParallelhardcode"fp16": {"enabled": True}in their DeepSpeed configs. The engine's sanity check then raises:on any accelerator whose
is_fp16_supported()is false. On CPU that maps to the AVX512-FP16 capability of the host, and GitHub'subuntu-24.04runners are hardware-heterogeneous: the same test passes on one runner and fails on the next (observed directly in #8381 — 146 failures appeared on one runner generation and none on another, with identical code).Change
Skip these tests via a capability query:
@pytest.mark.skipif(not get_accelerator().is_fp16_supported(), reason="fp16 is not supported on this accelerator")tests/unit/v1/zero/test_zero_user_backward.py;Validation
Validated as part of the multi-rank CPU CI experiment in #8381: the 146 hardware-lottery failures became deterministic skips, zero regressions on previously-passing tests.