Probe the device module for train_cifar's fork_rng device entries - #8407
Conversation
train_cifar always listed the current device in fork_rng's devices=, but torch.cpu has no get_rng_state and fork_rng already saves the CPU/global RNG. Probe the device module for a per-device RNG instead of matching the accelerator name, so backends without get_rng_state (cpu) pass no device entries at all. The sibling fixes from the original commits (DDP device_ids pinning, device names, fp16 skips) landed separately as deepspeedai#8397, deepspeedai#8398 and deepspeedai#8399. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ade99f4704
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| with get_accelerator().random().fork_rng(devices=[get_accelerator().current_device_name()], **fork_kwargs): | ||
| # fork_rng only needs entries for backends with per-device generators: the global | ||
| # CPU RNG is always saved, and torch.cpu has no get_rng_state to call anyway. | ||
| device_mod = torch.get_device_module(get_accelerator().device_name()) |
There was a problem hiding this comment.
Guard the device-module lookup on older PyTorch
When these tests run with the supported minimum PyTorch 2.0 (requirements/requirements.txt specifies torch>=2.0.0), torch.get_device_module does not exist, so every train_cifar call now raises AttributeError before entering fork_rng, including CUDA paths that previously worked. Please use a lookup available on older supported releases or guard this API by PyTorch version.
Useful? React with 👍 / 👎.
| else: | ||
| fork_kwargs = {} | ||
| with get_accelerator().random().fork_rng(devices=[get_accelerator().current_device_name()], **fork_kwargs): | ||
| # fork_rng only needs entries for backends with per-device generators: the global |
There was a problem hiding this comment.
Add the required commit sign-off
This is a non-merge commit, but its message has no Signed-off-by trailer. Recreate the commit with --signoff so it satisfies the repository's commit requirements.
AGENTS.md reference: AGENTS.md:L8-L8
Useful? React with 👍 / 👎.
0306300
…peedai#8409) ## Description The dynamic offload-state tests assert strict allocated-memory deltas around `offload_states()` / `reload_states()`: - `alloc_after_offload < alloc_before_offload` - `alloc_after_reload > alloc_after_offload` That contract assumes `memory_allocated()` is allocator bookkeeping, which holds on cuda (`torch.cuda.memory_allocated()`). On cpu, `CPU_Accelerator.memory_allocated()` reports process RSS (psutil), and RSS does not shrink when tensors are freed — so all 92 parameterized cases fail even when the offload itself is correct (the device-placement and data-integrity checks in the same tests pass). Gate only the memory-delta asserts on whether the accelerator's torch device module exposes `memory_allocated` (cuda does; `torch.cpu` does not), mirroring the capability probe used for `fork_rng` in `train_cifar` (deepspeedai#8407): - cuda and other allocator-backed backends: behavior unchanged - cpu: the unobservable deltas are skipped; all device-placement validations still run Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381 (92 of the 131 v1-half failures there). ## Validation (executed on real hardware) - 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend, 2 ranks (`LOCAL_SIZE=2`) - Before: `TestDynamicOffloadStatesZero12[False-1-False-False-optim_states]` fails on the persistent-state delta assert - After: 5 representative cases pass (persistent and grad paths, ZeRO stage 1/2/3, `static_offload_optimizer=True` branch) — 5 passed in 55.6s - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the changed file Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…pspeedai#8559) ## Description `bf16_required_version_check()` (`tests/unit/util.py`) requires torch >= 1.10, **CUDA >= 11.0 and NCCL >= 2.10.3**. On the cpu accelerator, bf16 collectives run over gloo/ccl and none of those transport dependencies exist, so the check always returns False and **every bf16 test is skipped — about 40 call sites across 15 files**. `test_zero_autocast.py` is worse off: it **raises** instead of skipping, so each of its cases counts as a failure (24 on the multi-rank CPU run in deepspeedai#8381). This PR scopes the version floors inside the check itself: ```python if (cpu_accelerator and accelerator_pass) or (torch_version_available and cuda_version_available and nccl_version_available and accelerator_pass): return True ``` - **cpu**: only the accelerator's own bf16 support (`is_bf16_supported()`) is required — the torch/CUDA/NCCL floors are transport dependencies that gloo/ccl does not have - **every other accelerator** (cuda, npu, hpu, xpu, mlu, ...): evaluates the exact original floors; `cpu_accelerator` is False so the expression is bit-identical to before, and the npu/hpu/xpu exemption branches are untouched - `test_zero_autocast.py`: the bf16 gate goes back to the bare call every other caller uses, and skips instead of raising The hardcoded `init_distributed(dist_backend='nccl')` in the same test is deliberately left alone: it is a no-op (the harness already initialized the process group, `comm.py:838-839`), and deriving the backend per accelerator would change behavior on non-cuda accelerators (npu/hpu resolve to hccl etc.). The baseline `DDP(device_ids=[i])` pinning is left for a follow-up (deepspeedai#8399 fixed the same pattern elsewhere). ## Validation (executed on real hardware) - 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend - Direct call: `bf16_required_version_check()` on cpu returns `False` before, `True` after. Note: `CPU_Accelerator.is_bf16_supported()` is currently a stub that always returns True, so the cpu path does not gate on the hardware's bf16 instructions — giving it a real capability probe is left as a follow-up - Non-cpu equivalence: with `cpu_accelerator == False` the new expression reduces exactly to the original `A and C and N and P` - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on both changed files Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409. Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…dai#8684) ## Description Every fp16-config test that reaches `deepspeed.initialize` crashes its sanity check (`Type fp16 is not supported on your device.`) on accelerators whose `is_fp16_supported()` is false. On CPU that maps to the AVX512-FP16 capability of the host, and GitHub's `ubuntu-24.04` runners are hardware-heterogeneous, so these tests flip between failure and skip depending on which runner they land on. deepspeedai#8398 added the first skipifs; the multi-rank CPU run in deepspeedai#8381 flushed out six more files: | File | Shape of the gap | |---|---| | `checkpoint/test_universal_checkpoint.py` | fp16 parametrizations fail **inside the baseline DistributedFixture's distributed run** — pytest reports a setup **ERROR** for every dependent test (48 on the multi-rank run) instead of a skip | | `v1/zero/test_zero_coalesce_grad_reduction.py` | `TestCoalesceFP16` forces an fp16 config | | `runtime/test_no_sync_ctxt.py` | dtype=float16 parametrizations of three methods; stages 2/3 additionally never reach their expected no_sync AssertionError on such hosts | | `checkpoint/test_moe_checkpoint.py` | whole class hardcodes fp16 | | `runtime/zero/test_zero_offloadpp.py` | `TestZeroPartialOffloadConfigSweep` hardcodes fp16 | | `checkpoint/test_pipeline.py` | fp16 enabled for zero_stage > 0; only that parametrization skips, zero_stage=0 keeps running | With this PR, **every fp16-config test in the suite guards on `is_fp16_supported()`** — the capability gap is a skip, not a failure, on any accelerator. ## Validation (executed on real hardware) - 20-core x86_64 CPU without AVX512-FP16, gloo, 2-4 ranks (`LOCAL_SIZE=2/4`) - Before: 18 failures + 48 setup ERRORs across these files on the multi-rank CPU run - After: every affected parametrization skips; adjacent non-fp16 parametrizations keep passing (e.g. pipeline zero_stage=0 runs to completion) - Full-suite evidence in deepspeedai#8381: the multi-rank CPU run went 8 failures -> 0 with these guards Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407, deepspeedai#8409, deepspeedai#8559, deepspeedai#8648. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Description
train_cifarunconditionally passeddevices=[get_accelerator().current_device_name()]tofork_rng. On CPU that becomesdevices=['cpu'], andtorch.cpuhas noget_rng_state, so the context manager raisesAttributeErrorbefore the test body starts — the CPU RNG lives in the global generator thatfork_rngalready forks.Probe the device module for a per-device RNG instead of matching accelerator names:
get_rng_state(e.g. cuda) behave exactly as before;devices=[], which changes nothing beyond the global-RNG save/restorefork_rngalways performs.The crash is only reachable from the multi-rank tests that call
train_cifar(test_onebit.py,test_pipe.py), which the CPU runner currently skips at the device gate; it was exposed by theLOCAL_SIZE=4experiment in #8381. Sibling fixes from the same series landed as #8397, #8398 and #8399.Test plan
LOCAL_SIZE=4multi-rank CPU run tracked in [DON'T MERGE] Run multi-rank CPU unit tests in CI via LOCAL_SIZE #8381: thetest_onebit/test_pipecallers got pastfork_rngand into their test bodies with this exact change.