Do not pin DDP device_ids for CPU reference models - #8399
Conversation
test_zero_user_backward builds torch DDP reference models at five sites via DDP(model, device_ids=[rank], output_device=rank). torch's DDP contract only allows device_ids/output_device for single-device GPU modules; CPU modules live on one shared device and must omit them, so multi-rank CPU runs failed inside the DDP constructor. Route all five sites through one wrap_ddp_reference helper that checks the model's torch device type - the same precondition torch's own DDP constructor enforces - and pins device ids only for indexed devices. GPU behavior is unchanged. Validated as part of the multi-rank CPU experiment in deepspeedai#8381: all DDP-constructor failures were eliminated (the file's few remaining failures there are unrelated, see its triage table); zero regressions vs the same-commit baseline. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 41e1dc720b
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| return model, optimizer | ||
|
|
||
|
|
||
| def wrap_ddp_reference(model, device, rank): |
There was a problem hiding this comment.
Add the required sign-off trailer
This is a non-merge commit, but its message has no Signed-off-by trailer, so it does not satisfy the repository's commit requirements and may be rejected by the DCO/CI checks. Recreate the commit with git commit --signoff using the configured Git identity.
AGENTS.md reference: AGENTS.md:L8-L8
Useful? React with 👍 / 👎.
…epspeedai#8407) ## Description `train_cifar` unconditionally passed `devices=[get_accelerator().current_device_name()]` to `fork_rng`. On CPU that becomes `devices=['cpu']`, and `torch.cpu` has no `get_rng_state`, so the context manager raises `AttributeError` before the test body starts — the CPU RNG lives in the global generator that `fork_rng` already forks. Probe the device module for a per-device RNG instead of matching accelerator names: - backends whose device module has `get_rng_state` (e.g. cuda) behave exactly as before; - backends without one (cpu) pass `devices=[]`, which changes nothing beyond the global-RNG save/restore `fork_rng` always performs. The crash is only reachable from the multi-rank tests that call `train_cifar` (`test_onebit.py`, `test_pipe.py`), which the CPU runner currently skips at the device gate; it was exposed by the `LOCAL_SIZE=4` experiment in deepspeedai#8381. Sibling fixes from the same series landed as deepspeedai#8397, deepspeedai#8398 and deepspeedai#8399. ## Test plan - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the changed file; - exercised under the `LOCAL_SIZE=4` multi-rank CPU run tracked in deepspeedai#8381: the `test_onebit` / `test_pipe` callers got past `fork_rng` and into their test bodies with this exact change. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…peedai#8409) ## Description The dynamic offload-state tests assert strict allocated-memory deltas around `offload_states()` / `reload_states()`: - `alloc_after_offload < alloc_before_offload` - `alloc_after_reload > alloc_after_offload` That contract assumes `memory_allocated()` is allocator bookkeeping, which holds on cuda (`torch.cuda.memory_allocated()`). On cpu, `CPU_Accelerator.memory_allocated()` reports process RSS (psutil), and RSS does not shrink when tensors are freed — so all 92 parameterized cases fail even when the offload itself is correct (the device-placement and data-integrity checks in the same tests pass). Gate only the memory-delta asserts on whether the accelerator's torch device module exposes `memory_allocated` (cuda does; `torch.cpu` does not), mirroring the capability probe used for `fork_rng` in `train_cifar` (deepspeedai#8407): - cuda and other allocator-backed backends: behavior unchanged - cpu: the unobservable deltas are skipped; all device-placement validations still run Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381 (92 of the 131 v1-half failures there). ## Validation (executed on real hardware) - 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend, 2 ranks (`LOCAL_SIZE=2`) - Before: `TestDynamicOffloadStatesZero12[False-1-False-False-optim_states]` fails on the persistent-state delta assert - After: 5 representative cases pass (persistent and grad paths, ZeRO stage 1/2/3, `static_offload_optimizer=True` branch) — 5 passed in 55.6s - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the changed file Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…pspeedai#8559) ## Description `bf16_required_version_check()` (`tests/unit/util.py`) requires torch >= 1.10, **CUDA >= 11.0 and NCCL >= 2.10.3**. On the cpu accelerator, bf16 collectives run over gloo/ccl and none of those transport dependencies exist, so the check always returns False and **every bf16 test is skipped — about 40 call sites across 15 files**. `test_zero_autocast.py` is worse off: it **raises** instead of skipping, so each of its cases counts as a failure (24 on the multi-rank CPU run in deepspeedai#8381). This PR scopes the version floors inside the check itself: ```python if (cpu_accelerator and accelerator_pass) or (torch_version_available and cuda_version_available and nccl_version_available and accelerator_pass): return True ``` - **cpu**: only the accelerator's own bf16 support (`is_bf16_supported()`) is required — the torch/CUDA/NCCL floors are transport dependencies that gloo/ccl does not have - **every other accelerator** (cuda, npu, hpu, xpu, mlu, ...): evaluates the exact original floors; `cpu_accelerator` is False so the expression is bit-identical to before, and the npu/hpu/xpu exemption branches are untouched - `test_zero_autocast.py`: the bf16 gate goes back to the bare call every other caller uses, and skips instead of raising The hardcoded `init_distributed(dist_backend='nccl')` in the same test is deliberately left alone: it is a no-op (the harness already initialized the process group, `comm.py:838-839`), and deriving the backend per accelerator would change behavior on non-cuda accelerators (npu/hpu resolve to hccl etc.). The baseline `DDP(device_ids=[i])` pinning is left for a follow-up (deepspeedai#8399 fixed the same pattern elsewhere). ## Validation (executed on real hardware) - 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend - Direct call: `bf16_required_version_check()` on cpu returns `False` before, `True` after. Note: `CPU_Accelerator.is_bf16_supported()` is currently a stub that always returns True, so the cpu path does not gate on the hardware's bf16 instructions — giving it a real capability probe is left as a follow-up - Non-cpu equivalence: with `cpu_accelerator == False` the new expression reduces exactly to the original `A and C and N and P` - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on both changed files Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409. Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The torch_autocast baseline pins device_ids/output_device to the rank unconditionally, but CPU modules live on one shared device and torch requires device_ids=None there — the same pattern deepspeedai#8399 fixed for the reference models in test_zero_user_backward. Mirror its check so only indexed devices take device_ids. With the barrier gone the bf16 cases run to completion and pass on cpu (sampled stages 0/3, safe-modules and nested configurations). The fp16 cases now fail later on the loss-parity comparison instead — cpu fp16 autocast semantics are a separate follow-up. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…dai#8684) ## Description Every fp16-config test that reaches `deepspeed.initialize` crashes its sanity check (`Type fp16 is not supported on your device.`) on accelerators whose `is_fp16_supported()` is false. On CPU that maps to the AVX512-FP16 capability of the host, and GitHub's `ubuntu-24.04` runners are hardware-heterogeneous, so these tests flip between failure and skip depending on which runner they land on. deepspeedai#8398 added the first skipifs; the multi-rank CPU run in deepspeedai#8381 flushed out six more files: | File | Shape of the gap | |---|---| | `checkpoint/test_universal_checkpoint.py` | fp16 parametrizations fail **inside the baseline DistributedFixture's distributed run** — pytest reports a setup **ERROR** for every dependent test (48 on the multi-rank run) instead of a skip | | `v1/zero/test_zero_coalesce_grad_reduction.py` | `TestCoalesceFP16` forces an fp16 config | | `runtime/test_no_sync_ctxt.py` | dtype=float16 parametrizations of three methods; stages 2/3 additionally never reach their expected no_sync AssertionError on such hosts | | `checkpoint/test_moe_checkpoint.py` | whole class hardcodes fp16 | | `runtime/zero/test_zero_offloadpp.py` | `TestZeroPartialOffloadConfigSweep` hardcodes fp16 | | `checkpoint/test_pipeline.py` | fp16 enabled for zero_stage > 0; only that parametrization skips, zero_stage=0 keeps running | With this PR, **every fp16-config test in the suite guards on `is_fp16_supported()`** — the capability gap is a skip, not a failure, on any accelerator. ## Validation (executed on real hardware) - 20-core x86_64 CPU without AVX512-FP16, gloo, 2-4 ranks (`LOCAL_SIZE=2/4`) - Before: 18 failures + 48 setup ERRORs across these files on the multi-rank CPU run - After: every affected parametrization skips; adjacent non-fp16 parametrizations keep passing (e.g. pipeline zero_stage=0 runs to completion) - Full-suite evidence in deepspeedai#8381: the multi-rank CPU run went 8 failures -> 0 with these guards Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409, deepspeedai#8559, deepspeedai#8648. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Problem
tests/unit/v1/zero/test_zero_user_backward.pybuilds torch DDP reference models (the known-good baseline that ZeRO results are compared against) at five sites:device_ids=[rank]assumes rank ↔ GPU index. torch's DDP contract only allowsdevice_ids/output_devicefor single-device GPU modules; CPU modules live on one shared device and must omit them, so multi-rank CPU runs died inside the DDP constructor with:Change
Route all five sites through one helper:
Design notes:
Validation
Validated as part of the multi-rank CPU CI experiment in #8381: all DDP-constructor failures were eliminated (the file's few remaining failures there are unrelated — see the triage table in that PR), zero regressions vs the same-commit baseline.