Use device names, not rank ids, for device placement in test helpers - #8397
Conversation
get_accelerator().current_device() returns a CUDA device index on GPU
backends, but on CPU it is the LOCAL_RANK environment value, a plain
string like '1' that torch rejects as a device ("Invalid device
string"). reduce_boolean_flags and the autotp tests fed that value
straight into tensor/device placement, so multi-rank CPU runs failed
before reaching their first collective.
Use current_device_name() instead: it yields a full device string on
every backend ('cpu', 'cuda:N', ...) and is equivalent to the index on
GPU backends. Also carry the boolean flag in a 1-dim tensor because
gloo rejects 0-dim inputs to all_gather_into_tensor.
Validated as part of the multi-rank CPU experiment in deepspeedai#8381: this
failure class disappeared with zero regressions vs the same-commit
baseline.
Signed-off-by: Guokai Ma <guokai.ma@intel.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8ceca37333
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| device = get_accelerator().current_device() | ||
| tensor_flag = torch.tensor(1 if flag else 0, dtype=torch.int, device=device) | ||
| # current_device() is a rank id on CPU, not a valid torch device; use the device name. | ||
| device = get_accelerator().current_device_name() |
There was a problem hiding this comment.
Add the required sign-off trailer
This is a non-merge commit, but its message has no Signed-off-by trailer, so it violates the repository's mandatory commit requirement and may be rejected by DCO/CI checks. Recreate the commit with --signoff using the configured Git identity.
AGENTS.md reference: AGENTS.md:L8-L8
Useful? React with 👍 / 👎.
The reference-input randn() call exceeds the 119-column limit, failing the formatting CI. Split one argument per line as yapf requires. Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
…epspeedai#8407) ## Description `train_cifar` unconditionally passed `devices=[get_accelerator().current_device_name()]` to `fork_rng`. On CPU that becomes `devices=['cpu']`, and `torch.cpu` has no `get_rng_state`, so the context manager raises `AttributeError` before the test body starts — the CPU RNG lives in the global generator that `fork_rng` already forks. Probe the device module for a per-device RNG instead of matching accelerator names: - backends whose device module has `get_rng_state` (e.g. cuda) behave exactly as before; - backends without one (cpu) pass `devices=[]`, which changes nothing beyond the global-RNG save/restore `fork_rng` always performs. The crash is only reachable from the multi-rank tests that call `train_cifar` (`test_onebit.py`, `test_pipe.py`), which the CPU runner currently skips at the device gate; it was exposed by the `LOCAL_SIZE=4` experiment in deepspeedai#8381. Sibling fixes from the same series landed as deepspeedai#8397, deepspeedai#8398 and deepspeedai#8399. ## Test plan - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the changed file; - exercised under the `LOCAL_SIZE=4` multi-rank CPU run tracked in deepspeedai#8381: the `test_onebit` / `test_pipe` callers got past `fork_rng` and into their test bodies with this exact change. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…peedai#8409) ## Description The dynamic offload-state tests assert strict allocated-memory deltas around `offload_states()` / `reload_states()`: - `alloc_after_offload < alloc_before_offload` - `alloc_after_reload > alloc_after_offload` That contract assumes `memory_allocated()` is allocator bookkeeping, which holds on cuda (`torch.cuda.memory_allocated()`). On cpu, `CPU_Accelerator.memory_allocated()` reports process RSS (psutil), and RSS does not shrink when tensors are freed — so all 92 parameterized cases fail even when the offload itself is correct (the device-placement and data-integrity checks in the same tests pass). Gate only the memory-delta asserts on whether the accelerator's torch device module exposes `memory_allocated` (cuda does; `torch.cpu` does not), mirroring the capability probe used for `fork_rng` in `train_cifar` (deepspeedai#8407): - cuda and other allocator-backed backends: behavior unchanged - cpu: the unobservable deltas are skipped; all device-placement validations still run Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381 (92 of the 131 v1-half failures there). ## Validation (executed on real hardware) - 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend, 2 ranks (`LOCAL_SIZE=2`) - Before: `TestDynamicOffloadStatesZero12[False-1-False-False-optim_states]` fails on the persistent-state delta assert - After: 5 representative cases pass (persistent and grad paths, ZeRO stage 1/2/3, `static_offload_optimizer=True` branch) — 5 passed in 55.6s - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the changed file Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…pspeedai#8559) ## Description `bf16_required_version_check()` (`tests/unit/util.py`) requires torch >= 1.10, **CUDA >= 11.0 and NCCL >= 2.10.3**. On the cpu accelerator, bf16 collectives run over gloo/ccl and none of those transport dependencies exist, so the check always returns False and **every bf16 test is skipped — about 40 call sites across 15 files**. `test_zero_autocast.py` is worse off: it **raises** instead of skipping, so each of its cases counts as a failure (24 on the multi-rank CPU run in deepspeedai#8381). This PR scopes the version floors inside the check itself: ```python if (cpu_accelerator and accelerator_pass) or (torch_version_available and cuda_version_available and nccl_version_available and accelerator_pass): return True ``` - **cpu**: only the accelerator's own bf16 support (`is_bf16_supported()`) is required — the torch/CUDA/NCCL floors are transport dependencies that gloo/ccl does not have - **every other accelerator** (cuda, npu, hpu, xpu, mlu, ...): evaluates the exact original floors; `cpu_accelerator` is False so the expression is bit-identical to before, and the npu/hpu/xpu exemption branches are untouched - `test_zero_autocast.py`: the bf16 gate goes back to the bare call every other caller uses, and skips instead of raising The hardcoded `init_distributed(dist_backend='nccl')` in the same test is deliberately left alone: it is a no-op (the harness already initialized the process group, `comm.py:838-839`), and deriving the backend per accelerator would change behavior on non-cuda accelerators (npu/hpu resolve to hccl etc.). The baseline `DDP(device_ids=[i])` pinning is left for a follow-up (deepspeedai#8399 fixed the same pattern elsewhere). ## Validation (executed on real hardware) - 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend - Direct call: `bf16_required_version_check()` on cpu returns `False` before, `True` after. Note: `CPU_Accelerator.is_bf16_supported()` is currently a stub that always returns True, so the cpu path does not gate on the hardware's bf16 instructions — giving it a real capability probe is left as a follow-up - Non-cpu equivalence: with `cpu_accelerator == False` the new expression reduces exactly to the original `A and C and N and P` - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on both changed files Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409. Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…dai#8684) ## Description Every fp16-config test that reaches `deepspeed.initialize` crashes its sanity check (`Type fp16 is not supported on your device.`) on accelerators whose `is_fp16_supported()` is false. On CPU that maps to the AVX512-FP16 capability of the host, and GitHub's `ubuntu-24.04` runners are hardware-heterogeneous, so these tests flip between failure and skip depending on which runner they land on. deepspeedai#8398 added the first skipifs; the multi-rank CPU run in deepspeedai#8381 flushed out six more files: | File | Shape of the gap | |---|---| | `checkpoint/test_universal_checkpoint.py` | fp16 parametrizations fail **inside the baseline DistributedFixture's distributed run** — pytest reports a setup **ERROR** for every dependent test (48 on the multi-rank run) instead of a skip | | `v1/zero/test_zero_coalesce_grad_reduction.py` | `TestCoalesceFP16` forces an fp16 config | | `runtime/test_no_sync_ctxt.py` | dtype=float16 parametrizations of three methods; stages 2/3 additionally never reach their expected no_sync AssertionError on such hosts | | `checkpoint/test_moe_checkpoint.py` | whole class hardcodes fp16 | | `runtime/zero/test_zero_offloadpp.py` | `TestZeroPartialOffloadConfigSweep` hardcodes fp16 | | `checkpoint/test_pipeline.py` | fp16 enabled for zero_stage > 0; only that parametrization skips, zero_stage=0 keeps running | With this PR, **every fp16-config test in the suite guards on `is_fp16_supported()`** — the capability gap is a skip, not a failure, on any accelerator. ## Validation (executed on real hardware) - 20-core x86_64 CPU without AVX512-FP16, gloo, 2-4 ranks (`LOCAL_SIZE=2/4`) - Before: 18 failures + 48 setup ERRORs across these files on the multi-rank CPU run - After: every affected parametrization skips; adjacent non-fp16 parametrizations keep passing (e.g. pipeline zero_stage=0 runs to completion) - Full-suite evidence in deepspeedai#8381: the multi-rank CPU run went 8 failures -> 0 with these guards Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409, deepspeedai#8559, deepspeedai#8648. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Problem
get_accelerator().current_device()returns a device index on GPU backends (torch.cuda.current_device()→ int), but on CPU it returns theLOCAL_RANKenvironment value — a plain string like'1'. Two test-side consumers fed that value straight into tensor/device placement:reduce_boolean_flagsintests/unit/common.py(backbone ofallclose_on_all_ranks, the "all ranks succeed or fail together" check)tests/unit/v1/autotp/test_autotp_training.pyOn CPU this fails immediately with
RuntimeError: Invalid device string: '1'— before the first collective even runs.Change
current_device_name(), which returns a full device string on every backend ('cpu','cuda:N','mps:0', …) and is equivalent to the index on GPU backends.reduce_boolean_flags, carry the flag in a 1-dim tensor: gloo rejects 0-dim inputs toall_gather_into_tensor(NCCL tolerates them), so the previous form would have failed on the very next line.Validation
Validated as part of the multi-rank CPU CI experiment in #8381 (same-commit baseline comparison): this failure class disappeared, zero regressions on previously-passing tests.