Skip to content

Use device names, not rank ids, for device placement in test helpers - #8397

Merged
delock merged 2 commits into
deepspeedai:masterfrom
delock:pr-a-device-names
Sep 3, 2026
Merged

delock merged 2 commits into
deepspeedai:masterfrom
delock:pr-a-device-names

Conversation

@delock

@delock delock commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator

Problem

get_accelerator().current_device() returns a device index on GPU backends (torch.cuda.current_device() → int), but on CPU it returns the LOCAL_RANK environment value — a plain string like '1'. Two test-side consumers fed that value straight into tensor/device placement:

  • reduce_boolean_flags in tests/unit/common.py (backbone of allclose_on_all_ranks, the "all ranks succeed or fail together" check)
  • 15 call sites in tests/unit/v1/autotp/test_autotp_training.py

On CPU this fails immediately with RuntimeError: Invalid device string: '1' — before the first collective even runs.

Change

  • Use current_device_name(), which returns a full device string on every backend ('cpu', 'cuda:N', 'mps:0', …) and is equivalent to the index on GPU backends.
  • In reduce_boolean_flags, carry the flag in a 1-dim tensor: gloo rejects 0-dim inputs to all_gather_into_tensor (NCCL tolerates them), so the previous form would have failed on the very next line.

Validation

Validated as part of the multi-rank CPU CI experiment in #8381 (same-commit baseline comparison): this failure class disappeared, zero regressions on previously-passing tests.

get_accelerator().current_device() returns a CUDA device index on GPU
backends, but on CPU it is the LOCAL_RANK environment value, a plain
string like '1' that torch rejects as a device ("Invalid device
string"). reduce_boolean_flags and the autotp tests fed that value
straight into tensor/device placement, so multi-rank CPU runs failed
before reaching their first collective.

Use current_device_name() instead: it yields a full device string on
every backend ('cpu', 'cuda:N', ...) and is equivalent to the index on
GPU backends. Also carry the boolean flag in a 1-dim tensor because
gloo rejects 0-dim inputs to all_gather_into_tensor.

Validated as part of the multi-rank CPU experiment in deepspeedai#8381: this
failure class disappeared with zero regressions vs the same-commit
baseline.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8ceca37333

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread tests/unit/common.py
device = get_accelerator().current_device()
tensor_flag = torch.tensor(1 if flag else 0, dtype=torch.int, device=device)
# current_device() is a rank id on CPU, not a valid torch device; use the device name.
device = get_accelerator().current_device_name()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Add the required sign-off trailer

This is a non-merge commit, but its message has no Signed-off-by trailer, so it violates the repository's mandatory commit requirement and may be rejected by DCO/CI checks. Recreate the commit with --signoff using the configured Git identity.

AGENTS.md reference: AGENTS.md:L8-L8

Useful? React with 👍 / 👎.

@delock
delock enabled auto-merge September 3, 2026 09:55
@delock
delock disabled auto-merge September 3, 2026 09:55
The reference-input randn() call exceeds the 119-column limit, failing
the formatting CI. Split one argument per line as yapf requires.

Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
@delock
delock enabled auto-merge September 3, 2026 10:34
@delock
delock added this pull request to the merge queue Sep 3, 2026
Merged via the queue into deepspeedai:master with commit 4fd3c82 Sep 3, 2026
13 checks passed
@delock
delock deleted the pr-a-device-names branch September 3, 2026 11:50
banxingmjj pushed a commit to openanolis/DeepSpeed that referenced this pull request Sep 15, 2026
…epspeedai#8407)

## Description

`train_cifar` unconditionally passed
`devices=[get_accelerator().current_device_name()]` to `fork_rng`. On
CPU that becomes `devices=['cpu']`, and `torch.cpu` has no
`get_rng_state`, so the context manager raises `AttributeError` before
the test body starts — the CPU RNG lives in the global generator that
`fork_rng` already forks.

Probe the device module for a per-device RNG instead of matching
accelerator names:

- backends whose device module has `get_rng_state` (e.g. cuda) behave
exactly as before;
- backends without one (cpu) pass `devices=[]`, which changes nothing
beyond the global-RNG save/restore `fork_rng` always performs.

The crash is only reachable from the multi-rank tests that call
`train_cifar` (`test_onebit.py`, `test_pipe.py`), which the CPU runner
currently skips at the device gate; it was exposed by the `LOCAL_SIZE=4`
experiment in deepspeedai#8381. Sibling fixes from the same series landed as deepspeedai#8397,
deepspeedai#8398 and deepspeedai#8399.

## Test plan

- pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the
changed file;
- exercised under the `LOCAL_SIZE=4` multi-rank CPU run tracked in
deepspeedai#8381: the `test_onebit` / `test_pipe` callers got past `fork_rng` and
into their test bodies with this exact change.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
banxingmjj pushed a commit to openanolis/DeepSpeed that referenced this pull request Sep 15, 2026
…peedai#8409)

## Description

The dynamic offload-state tests assert strict allocated-memory deltas
around `offload_states()` / `reload_states()`:

- `alloc_after_offload < alloc_before_offload`
- `alloc_after_reload > alloc_after_offload`

That contract assumes `memory_allocated()` is allocator bookkeeping,
which holds on cuda (`torch.cuda.memory_allocated()`). On cpu,
`CPU_Accelerator.memory_allocated()` reports process RSS (psutil), and
RSS does not shrink when tensors are freed — so all 92 parameterized
cases fail even when the offload itself is correct (the device-placement
and data-integrity checks in the same tests pass).

Gate only the memory-delta asserts on whether the accelerator's torch
device module exposes `memory_allocated` (cuda does; `torch.cpu` does
not), mirroring the capability probe used for `fork_rng` in
`train_cifar` (deepspeedai#8407):

- cuda and other allocator-backed backends: behavior unchanged
- cpu: the unobservable deltas are skipped; all device-placement
validations still run

Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381 (92 of the 131
v1-half failures there).

## Validation (executed on real hardware)

- 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend, 2 ranks
(`LOCAL_SIZE=2`)
- Before:
`TestDynamicOffloadStatesZero12[False-1-False-False-optim_states]` fails
on the persistent-state delta assert
- After: 5 representative cases pass (persistent and grad paths, ZeRO
stage 1/2/3, `static_offload_optimizer=True` branch) — 5 passed in 55.6s
- pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the
changed file

Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
pull Bot pushed a commit to QSLee-Net/DeepSpeed that referenced this pull request Sep 17, 2026
…pspeedai#8559)

## Description

`bf16_required_version_check()` (`tests/unit/util.py`) requires torch >=
1.10, **CUDA >= 11.0 and NCCL >= 2.10.3**. On the cpu accelerator, bf16
collectives run over gloo/ccl and none of those transport dependencies
exist, so the check always returns False and **every bf16 test is
skipped — about 40 call sites across 15 files**. `test_zero_autocast.py`
is worse off: it **raises** instead of skipping, so each of its cases
counts as a failure (24 on the multi-rank CPU run in deepspeedai#8381).

This PR scopes the version floors inside the check itself:

```python
if (cpu_accelerator and accelerator_pass) or (torch_version_available and cuda_version_available
                                              and nccl_version_available and accelerator_pass):
    return True
```

- **cpu**: only the accelerator's own bf16 support
(`is_bf16_supported()`) is required — the torch/CUDA/NCCL floors are
transport dependencies that gloo/ccl does not have
- **every other accelerator** (cuda, npu, hpu, xpu, mlu, ...): evaluates
the exact original floors; `cpu_accelerator` is False so the expression
is bit-identical to before, and the npu/hpu/xpu exemption branches are
untouched
- `test_zero_autocast.py`: the bf16 gate goes back to the bare call
every other caller uses, and skips instead of raising

The hardcoded `init_distributed(dist_backend='nccl')` in the same test
is deliberately left alone: it is a no-op (the harness already
initialized the process group, `comm.py:838-839`), and deriving the
backend per accelerator would change behavior on non-cuda accelerators
(npu/hpu resolve to hccl etc.). The baseline `DDP(device_ids=[i])`
pinning is left for a follow-up (deepspeedai#8399 fixed the same pattern
elsewhere).

## Validation (executed on real hardware)

- 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend
- Direct call: `bf16_required_version_check()` on cpu returns `False`
before, `True` after. Note: `CPU_Accelerator.is_bf16_supported()` is
currently a stub that always returns True, so the cpu path does not gate
on the hardware's bf16 instructions — giving it a real capability probe
is left as a follow-up
- Non-cpu equivalence: with `cpu_accelerator == False` the new
expression reduces exactly to the original `A and C and N and P`
- pre-commit (yapf / flake8 / check-torchdist / codespell) passes on
both changed files

Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409.
Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
yermakoffivan pushed a commit to yermakoffivan/deepspeed that referenced this pull request Sep 28, 2026
…dai#8684)

## Description

Every fp16-config test that reaches `deepspeed.initialize` crashes its
sanity check (`Type fp16 is not supported on your device.`) on
accelerators whose `is_fp16_supported()` is false. On CPU that maps to
the AVX512-FP16 capability of the host, and GitHub's `ubuntu-24.04`
runners are hardware-heterogeneous, so these tests flip between failure
and skip depending on which runner they land on. deepspeedai#8398 added the first
skipifs; the multi-rank CPU run in deepspeedai#8381 flushed out six more files:

| File | Shape of the gap |
|---|---|
| `checkpoint/test_universal_checkpoint.py` | fp16 parametrizations fail
**inside the baseline DistributedFixture's distributed run** — pytest
reports a setup **ERROR** for every dependent test (48 on the multi-rank
run) instead of a skip |
| `v1/zero/test_zero_coalesce_grad_reduction.py` | `TestCoalesceFP16`
forces an fp16 config |
| `runtime/test_no_sync_ctxt.py` | dtype=float16 parametrizations of
three methods; stages 2/3 additionally never reach their expected
no_sync AssertionError on such hosts |
| `checkpoint/test_moe_checkpoint.py` | whole class hardcodes fp16 |
| `runtime/zero/test_zero_offloadpp.py` |
`TestZeroPartialOffloadConfigSweep` hardcodes fp16 |
| `checkpoint/test_pipeline.py` | fp16 enabled for zero_stage > 0; only
that parametrization skips, zero_stage=0 keeps running |

With this PR, **every fp16-config test in the suite guards on
`is_fp16_supported()`** — the capability gap is a skip, not a failure,
on any accelerator.

## Validation (executed on real hardware)

- 20-core x86_64 CPU without AVX512-FP16, gloo, 2-4 ranks
(`LOCAL_SIZE=2/4`)
- Before: 18 failures + 48 setup ERRORs across these files on the
multi-rank CPU run
- After: every affected parametrization skips; adjacent non-fp16
parametrizations keep passing (e.g. pipeline zero_stage=0 runs to
completion)
- Full-suite evidence in deepspeedai#8381: the multi-rank CPU run went 8 failures
-> 0 with these guards

Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409,
deepspeedai#8559, deepspeedai#8648.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants