Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: af3841378c
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| # multi-rank test. CPU ranks are plain processes over gloo, so advertise 4 | ||
| # local devices to let world_size<=4 tests run. The test harness re-sets | ||
| # LOCAL_SIZE per worker, so this value only affects the launch gate. | ||
| LOCAL_SIZE: '4' |
There was a problem hiding this comment.
Keep CUDA-only distributed tests out of the CPU run
In the cpu-torch-latest job, advertising four devices admits every distributed test with world_size <= 4, not only CPU-safe tests. For example, tests/unit/ulysses_alst/test_ulysses_sp_hf.py:240-263 defines an unguarded two-rank test that creates tensors on cuda:<rank>; because CPU_Accelerator.is_available() returns true, the harness does not skip it, and the all-unit pytest invocation at workflow line 283 will fail on the CPU-only PyTorch installation. Scope this override to an explicitly CPU-compatible subset or add CPU capability checks before enabling the previously skipped tests.
Useful? React with 👍 / 👎.
| # multi-rank test. CPU ranks are plain processes over gloo, so advertise 4 | ||
| # local devices to let world_size<=4 tests run. The test harness re-sets | ||
| # LOCAL_SIZE per worker, so this value only affects the launch gate. | ||
| LOCAL_SIZE: '4' |
There was a problem hiding this comment.
Add the required sign-off trailer
This is a non-merge commit, but its commit message has no Signed-off-by trailer. Add the author sign-off so the commit satisfies the repository's commit and CI requirements.
AGENTS.md reference: AGENTS.md:L8-L8
Useful? React with 👍 / 👎.
|
This is an on going work. Note I'll rebase when revealed bugs are fixed. The goal is fix all issues exposed by this CI, then turn on multi-rank test in CPU workflow. |
Experiment results: enabling multi-rank CPU tests via
|
| Count | Test | Root cause |
|---|---|---|
| 92 | v1/zero/test_offload_states.py |
Asserts memory_allocated() drops after offload — a VRAM concept. CPU_Accelerator.memory_allocated() returns RSS, which never shrinks on free. Test/accelerator semantic gap; needs discussion (gate on device semantics or a better CPU metric). |
| 24 | v1/zero/test_zero_autocast.py |
Hardcoded dist_backend='nccl' + a CUDA/NCCL-oriented bf16 gate. Needs accelerator-derived backend + capability-based skip. |
| 25 | ulysses SP tests | Numerical mismatches on CPU (real correctness questions, need deep dive). |
| 16 | onebit optimizer tests | NoneType.size at onebit/adam.py:108 — runtime bug on CPU path. |
| ~20 | pipe / zeropp / moe-checkpoint / coalesce | Smaller follow-ups, same GPU-assumption patterns. |
Test-side fixes already in this branch (validated by CI)
- All 5 DDP reference sites route through
wrap_ddp_reference()(CPU modules must not pindevice_ids) — dropped this class from 114 failures to 7. reduce_boolean_flags/ autotp tests usecurrent_device_name()instead ofcurrent_device()(a LOCAL_RANK string on CPU, not a device).- fp16-config tests get a capability skipif (
not get_accelerator().is_fp16_supported()) — GHubuntu-24.04runners are hardware-heterogeneous w.r.t. AVX512-FP16, so hardcoded fp16 was a runner lottery. fork_rngprobes the device module for a per-device RNG instead of matching accelerator names.
CI infrastructure findings (need maintainer decisions)
--maxfail=100inPYTEST_OPTS+ no timeout inDistributedExec._close_pool= permanent wedge: when the 100th failure trips the interrupt, teardown (pool.starmap(_dist_destroy),close/join) can block forever on wedged gloo peers; the run then dies at the 6h job limit without printing any summary. This is what cancelled the first three attempts. (This branch works around it by splitting the suite and raising maxfail.)- Suite size vs 4 vCPU: the full multi-rank suite does not fit one 6h job. Options: split invocations (done here), lower
-n, or a dedicated larger runner. CPU_Accelerator.device_count()'s NUMA-node semantics remain the root cause of the original gap; the long-term fix could be a test-harness-level exemption for CPU (ranks are processes, not devices).
Happy to split the commits into separate PRs (workflow change / mechanical test fixes / triage follow-ups) per maintainer preference — the failure inventory above is intended as the working list.
|
Given the CPU UT would run ~hrs after enabling LOCAL_SIZE=4, I wouldn't suggest to turn on this in master. However it is beneficial to use this branch to reveal problems in CPU accelerator and fix them. I'm mark this PR as do not merge. |
…dai#8398) ## Problem `TestMultipleModels::test_zero_optimizer`, `TestSimpleMoE`, `TestMoE`, `TestPRMoE`, and `TestMOETensorParallel` hardcode `"fp16": {"enabled": True}` in their DeepSpeed configs. The engine's sanity check then raises: ``` ValueError: Type fp16 is not supported on your device. ``` on any accelerator whose `is_fp16_supported()` is false. On CPU that maps to the AVX512-FP16 capability of the host, and GitHub's `ubuntu-24.04` runners are hardware-heterogeneous: **the same test passes on one runner and fails on the next** (observed directly in deepspeedai#8381 — 146 failures appeared on one runner generation and none on another, with identical code). ## Change Skip these tests via a capability query: ```python @pytest.mark.skipif(not get_accelerator().is_fp16_supported(), reason="fp16 is not supported on this accelerator") ``` - capability only, no accelerator-name matching; - mirrors the existing bf16 skip precedent in `tests/unit/v1/zero/test_zero_user_backward.py`; - deliberately a **skip** rather than silently running bf16 — these tests exist to cover the fp16 paths. ## Validation Validated as part of the multi-rank CPU CI experiment in deepspeedai#8381: the 146 hardware-lottery failures became deterministic skips, zero regressions on previously-passing tests. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
## Problem
`tests/unit/v1/zero/test_zero_user_backward.py` builds torch DDP
**reference** models (the known-good baseline that ZeRO results are
compared against) at five sites:
```python
model_ddp = DDP(model_ddp, device_ids=[rank], output_device=rank)
```
`device_ids=[rank]` assumes rank ↔ GPU index. torch's DDP contract only
allows `device_ids`/`output_device` for single-device GPU modules; **CPU
modules live on one shared device and must omit them**, so multi-rank
CPU runs died inside the DDP constructor with:
```
ValueError: DistributedDataParallel device_ids and output_device arguments only work with single-device/multiple-device GPU modules or CPU modules, ...
```
## Change
Route all five sites through one helper:
```python
def wrap_ddp_reference(model, device, rank):
# Only indexed devices take device_ids/output_device; CPU modules live on one shared device.
if torch.device(device).type == 'cpu':
return DDP(model)
return DDP(model, device_ids=[rank], output_device=rank)
```
Design notes:
- the condition restates the exact precondition torch's own DDP
constructor enforces, in torch's device vocabulary — it follows the
model's actual device rather than the global accelerator configuration;
- a single helper means new reference-model sites cannot forget the
branch (the first fix round in deepspeedai#8381 missed 4 of the 5 sites for exactly
this reason);
- GPU behavior is unchanged.
## Validation
Validated as part of the multi-rank CPU CI experiment in deepspeedai#8381: all
DDP-constructor failures were eliminated (the file's few remaining
failures there are unrelated — see the triage table in that PR), zero
regressions vs the same-commit baseline.
Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…eepspeedai#8397) ## Problem `get_accelerator().current_device()` returns a **device index** on GPU backends (`torch.cuda.current_device()` → int), but on CPU it returns the `LOCAL_RANK` environment value — a plain **string** like `'1'`. Two test-side consumers fed that value straight into tensor/device placement: - `reduce_boolean_flags` in `tests/unit/common.py` (backbone of `allclose_on_all_ranks`, the "all ranks succeed or fail together" check) - 15 call sites in `tests/unit/v1/autotp/test_autotp_training.py` On CPU this fails immediately with `RuntimeError: Invalid device string: '1'` — before the first collective even runs. ## Change - Use `current_device_name()`, which returns a full device string on every backend (`'cpu'`, `'cuda:N'`, `'mps:0'`, …) and is equivalent to the index on GPU backends. - In `reduce_boolean_flags`, carry the flag in a 1-dim tensor: gloo rejects 0-dim inputs to `all_gather_into_tensor` (NCCL tolerates them), so the previous form would have failed on the very next line. ## Validation Validated as part of the multi-rank CPU CI experiment in deepspeedai#8381 (same-commit baseline comparison): this failure class disappeared, zero regressions on previously-passing tests. --------- Signed-off-by: Guokai Ma <guokai.ma@intel.com> Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
cpu-torch-latest runs on a single-socket runner where CPU_Accelerator.device_count() reports 1 NUMA node, so the per-device gate in tests/unit/common.py skips every test that needs more than one rank. CPU ranks are plain processes over gloo and need no per-rank hardware, so advertise 4 local devices via LOCAL_SIZE, the env var device_count() reads first. The test harness re-sets LOCAL_SIZE per worker, so this value only affects the launch gate. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
With LOCAL_SIZE=4 the suite now runs to ~63% and then all xdist workers go silent for hours until the 6h job limit cancels the run - the pool worker cleanup hang that DS_DISABLE_REUSE_DIST_ENV was added for. Fresh pools per test cost some wall time but let the run finish and print the failure summary. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Both full-suite attempts wedge before printing a summary: once the 100th failure trips PYTEST_OPTS' --maxfail, pytest-xdist's interrupt path stalls forever in mp pool teardown (no timeout guards _close_pool), and the 6h job limit cancels the run. Run the suite as two fresh-worker halves, override maxfail so all failures are listed, and cap each half with timeout so the sequential tail always runs. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
f46af32 to
f2793e7
Compare
…epspeedai#8407) ## Description `train_cifar` unconditionally passed `devices=[get_accelerator().current_device_name()]` to `fork_rng`. On CPU that becomes `devices=['cpu']`, and `torch.cpu` has no `get_rng_state`, so the context manager raises `AttributeError` before the test body starts — the CPU RNG lives in the global generator that `fork_rng` already forks. Probe the device module for a per-device RNG instead of matching accelerator names: - backends whose device module has `get_rng_state` (e.g. cuda) behave exactly as before; - backends without one (cpu) pass `devices=[]`, which changes nothing beyond the global-RNG save/restore `fork_rng` always performs. The crash is only reachable from the multi-rank tests that call `train_cifar` (`test_onebit.py`, `test_pipe.py`), which the CPU runner currently skips at the device gate; it was exposed by the `LOCAL_SIZE=4` experiment in deepspeedai#8381. Sibling fixes from the same series landed as deepspeedai#8397, deepspeedai#8398 and deepspeedai#8399. ## Test plan - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the changed file; - exercised under the `LOCAL_SIZE=4` multi-rank CPU run tracked in deepspeedai#8381: the `test_onebit` / `test_pipe` callers got past `fork_rng` and into their test bodies with this exact change. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…peedai#8409) ## Description The dynamic offload-state tests assert strict allocated-memory deltas around `offload_states()` / `reload_states()`: - `alloc_after_offload < alloc_before_offload` - `alloc_after_reload > alloc_after_offload` That contract assumes `memory_allocated()` is allocator bookkeeping, which holds on cuda (`torch.cuda.memory_allocated()`). On cpu, `CPU_Accelerator.memory_allocated()` reports process RSS (psutil), and RSS does not shrink when tensors are freed — so all 92 parameterized cases fail even when the offload itself is correct (the device-placement and data-integrity checks in the same tests pass). Gate only the memory-delta asserts on whether the accelerator's torch device module exposes `memory_allocated` (cuda does; `torch.cpu` does not), mirroring the capability probe used for `fork_rng` in `train_cifar` (deepspeedai#8407): - cuda and other allocator-backed backends: behavior unchanged - cpu: the unobservable deltas are skipped; all device-placement validations still run Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381 (92 of the 131 v1-half failures there). ## Validation (executed on real hardware) - 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend, 2 ranks (`LOCAL_SIZE=2`) - Before: `TestDynamicOffloadStatesZero12[False-1-False-False-optim_states]` fails on the persistent-state delta assert - After: 5 representative cases pass (persistent and grad paths, ZeRO stage 1/2/3, `static_offload_optimizer=True` branch) — 5 passed in 55.6s - pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the changed file Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Status updateFixes extracted so far (spin-off PRs)Investigating the failures exposed by this CI change surfaced three real bugs, each with a standalone fix:
Current failures on this branch (84 total: 75 in the
|
The shm-based inference_all_reduce keys its /dev/shm segments by MASTER_ADDR/MASTER_PORT. The bind-and-release probe in get_master_port() hands every pool the same first-free port when tests run one process each (--forked), so concurrent pools share segments: ranks of different pools cross-write the collective state and produce wrong results or spin forever (the CI worker wedge). Also the op never unlinks its ~66MB-per- rank segments, which exhausts /dev/shm over a full run. Offset the port probe per process and per pool so each pool gets a distinct key, and remove the segments once the pool (or spawned procs) is done. Verified with two concurrent gloo pools calling inference_all_reduce: shared key gives wrong results (rank sees 2.0 instead of 3.0), distinct keys are correct; three rounds of -n 2 --forked TestDistInferenceAllReduce pass with zero /dev/shm leftovers. Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
Concurrent xdist workers each downloading bert-base-uncased race on the transformers cache lock and fail with PermissionError. Populate a shared HF_HOME once before the suite and point both unit-test halves at it (the sequential phase already used it), so tests only read the cache. Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
huggingface-cli is now a deprecation stub in the runner's huggingface_hub: it prints a warning and exits 1, killing the job in the pre-download step before any test runs. Use its replacement, hf download, which is already installed on the runner. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
bert-base-uncased is stored on Xet, and the anonymous xet-read-token request 404s, so hf download dies before populating the cache. Fall back to the classic resolve endpoint, which works without a token. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…eepspeedai#8539) ## Problem While investigating multi-rank CPU CI (deepspeedai#8381), `TestUnmanagedGradientAccumulation::test_unmanaged_varying_backward_count[3]` failed intermittently with parameters turned to NaN, and `test_zero_coalesce_grad_reduction` showed mismatches with denormal garbage values (`4.8e+37`, `1e-38`) — the signature of uninitialized memory. ## Root cause `_allgather_params_coalesced()` launches one async all_gather per parameter, but only waits on the **last** handle before pointing `param.data` at the `torch.empty` flat buffers: ```python launch_handles[-1].wait() param.data = gathered_tensor.narrow(...) if not get_accelerator().resolves_data_dependency(): get_accelerator().synchronize() ``` On CUDA this is safe by construction: the ops are enqueued from a single thread onto the same stream (FIFO), and `torch.cuda.synchronize()` waits for everything anyway. On gloo each async handle runs on an independent background thread with **no ordering between handles** — waiting for the last one says nothing about the earlier ones — and the CPU accelerator's `synchronize()` is a no-op. Reading `param.data` can therefore race the gather and expose uninitialized memory. The corruption is timing-dependent (locally 9/20 runs fail; deepspeedai#8382's comparison-helper fix made the numeric assertions reachable), which is why it was never caught on GPU CI. ## Fix Wait on every handle. On CUDA the earlier handles are guaranteed complete by stream ordering, so the extra waits are immediate status checks with no real waiting; the change makes correctness independent of that structural coincidence. ## Verification (CPU/gloo, world sizes 2–3) | Test | Before | After | |---|---|---| | `test_unmanaged_varying_backward_count[3]` | 9/20 failed | 0/20 failed | | `TestZero3ParamPartitioningBase` / `TestGradientAllreduceOp` / `TestUnmanagedGradientAccumulation` (regression, 57 cases) | — | 55 pass + 2 skip | GPU behavior is unchanged by construction (same buckets, same order, stream-FIFO + full synchronize); GPU CI will re-confirm. Signed-off-by: Guokai Ma <guokai.ma@intel.com> Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
The baseline DistributedFixture builds an fp16 engine config for every dtype=float16 parametrization, and deepspeed.initialize's sanity check raises "Type fp16 is not supported on your device." on accelerators without fp16 support. Because the failure happens inside the fixture's distributed run, pytest reports a setup ERROR for every dependent test (48 on the multi-rank CPU run) instead of a skip. Skip inside the fixture like deepspeedai#8398 did for the fp16-config tests; DistributedFixture propagates the skip to all dependent tests. Devices with fp16 support are unchanged. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
…dai#8610) Follow up deepspeedai#8381 @delock **Fix zeropp 9 (sequential)** in deepspeedai#8381 (comment) ``` unit/runtime/zero/test_zeropp.py ``` <img width="624" height="349" alt="image" src="https://github.com/user-attachments/assets/85750e15-91ca-4290-ae90-734302f20324" /> **passed in https://github.com/deepspeedai/DeepSpeed/actions/runs/35571162595/job/106242965025** This pull request updates the `tests/unit/runtime/zero/test_zeropp.py` test suite to ensure that certain tests are only run on accelerators that support FP16 precision. This prevents test failures on hardware that does not support FP16. Test robustness improvements: * Added `@pytest.mark.skipif` decorators to the `test`, `test_eval`, and `test_gradient_accumulation` methods in `TestZeroPPConfigSweep` and a `test` method in another test class, so these tests are skipped if the accelerator does not support FP16. This uses `get_accelerator().is_fp16_supported()` to check hardware capability. [[1]](diffhunk://#diff-e3c08df84c5c79ec0d57c31a7817cddb1ba2f31c620c482eaf4561698ffc3f27R173) [[2]](diffhunk://#diff-e3c08df84c5c79ec0d57c31a7817cddb1ba2f31c620c482eaf4561698ffc3f27R216) [[3]](diffhunk://#diff-e3c08df84c5c79ec0d57c31a7817cddb1ba2f31c620c482eaf4561698ffc3f27R258) [[4]](diffhunk://#diff-e3c08df84c5c79ec0d57c31a7817cddb1ba2f31c620c482eaf4561698ffc3f27R384) * Imported `get_accelerator` from `deepspeed.accelerator` to enable the FP16 support check. --------- Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
Follow up deepspeedai#8381 @delock fix ``` unit/pipe/test_pipe_module.py unit/runtime/pipe/test_pipe.py ``` This pull request refactors device handling in the activation checkpointing code to simplify and standardize how device variables are used. The main change is to consistently use a `device` variable instead of `cuda_device` and to remove unused or redundant stream-related code. This improves code clarity and maintainability without changing functionality. **Device handling simplification:** * Replaced all occurrences of `cuda_device` with a unified `device` variable obtained from `get_accelerator().current_device_name()` in `checkpointing.py`. [[1]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL604-R604) [[2]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL838-R831) [[3]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL941-R933) * Updated all device argument usages in calls to `copy_to_device`, `gather_partitioned_activations`, and `move_to_device` to use the new `device` variable. [[1]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL618-R617) [[2]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL709-R714) [[3]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL851-R843) [[4]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL952-R950) **Stream and code cleanup:** * Removed creation and usage of the `transport_stream` variable, and deleted commented-out or unused stream synchronization code. These changes make the device management more straightforward and reduce potential confusion around device and stream handling in activation checkpointing. Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
Follow up deepspeedai#8381 @delock Fix CPU CI failures in https://github.com/deepspeedai/DeepSpeed/actions/runs/35675316100/job/106580600974?pr=8381 ``` FAILED unit/inference/quantization/test_intX_quantization.py::TestQuantizedInt::test_half_int8_quantization FAILED unit/checkpoint/test_convert_checkpoint.py::TestCheckpointConvert::test_convert_zero_checkpoint_to_fp32_state_dict ``` --------- Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
TestCoalesceFP16 forces an fp16 config, and deepspeed.initialize's sanity check raises "Type fp16 is not supported on your device." on accelerators without fp16 support — the same hardware lottery deepspeedai#8398 handled for the other fp16-config tests. Add the same class-level skipif so the gap is a skip, not three failures. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The multiple-backward checkpointing test compares DDP and engine gradients over three optimizer iterations at dtype-default tolerances. The engine keeps fp32 accounting while the DDP reference steps the bf16 weights directly, so the two paths diverge by roughly one bf16 ulp of weight per step even from exactly equal gradients; by iteration 2 that crosses a reduction/relu rounding boundary and produces ~8e-3 absolute gradient differences on O(1) gradients. Whether the difference survives bf16 quantization depends on the CPU bf16 kernel's rounding, so the test flips between pass and fail across torch versions (fails on 2.10/2.11, passes on 2.13) without any behavioral change. Give iterations >= 2 an absolute tolerance floor (atol=1e-2 alongside the bf16-default rtol): the test already documents that late iterations drift, and the floor absorbs quantization-boundary chaos while still catching real gradient errors. Verified on both kernel families: 6/6 fail -> 6/6 pass on torch 2.11, unchanged 6/6 pass on torch 2.13. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
deepspeedai#8648) ## Description `TestZeroUserBackwardWithCheckpointing::test_checkpointed_multiple_backward` compares DDP and engine gradients over three optimizer iterations at dtype-default tolerances, and flips between pass and fail across torch versions on cpu (fails on 2.10/2.11 kernels, passes on 2.13) without any behavioral change. A minimal two-rank repro isolates the mechanism: ``` torch 2.13 torch 2.11 (= CI family) iter 0 grads exact match exact match post-step — linear2.weight differs by 1.7e-6 <- the seed iter 1 grads exact match exact match (quantization absorbs it) iter 2 grads exact match 7.8e-3 <- exceeds rtol final weights 1 bf16 ulp 1 bf16 ulp ``` The engine keeps fp32 accounting while the DDP reference steps the bf16 weights directly, so the two mathematically-equivalent paths diverge by ~one bf16 ulp of weight per optimizer step even from bit-identical gradients; by iteration 2 that crosses a reduction/relu rounding boundary. Whether the difference survives bf16 quantization depends on the CPU bf16 kernel rounding — hence the version dependence. This PR gives iterations >= 2 an absolute tolerance floor (`atol=1e-2` alongside the bf16-default `rtol=1.6e-2`). The test already documents "small differences at later iterations are expected due to bfloat16 precision"; the floor absorbs quantization-boundary chaos while still catching real gradient errors. ## Validation (executed on real hardware) - 20-core x86_64 CPU, gloo, 2 ranks (`LOCAL_SIZE=2`) - torch 2.11 (the failing kernel family): 6/6 parametrizations **fail -> pass** - torch 2.13 (the passing family): 6/6 pass before and after (no regression) - pre-commit passes on the changed file Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
With torch_autocast the engine reduces gradients in the autocast dtype while the fp32 DDP baseline reduces in fp32. On gloo an fp16 collective round-trips through fp32, so after one optimizer step the step-1 loss carries fp16-ulp relative differences (observed 2e-4..8e-4) that the default tolerance rejects even though step-0 matches bit-exact. Allow a low-precision tolerance for fp16 below ZeRO-3; the ZeRO-3 fp16 variants keep the strict tolerance (their CPU divergence is much larger and under investigation). Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
Every rank worker JIT-loads deepspeed_shm_comm on first init_distributed. Concurrent loads serialize on torch's FileBaton, which has no stale-lock handling: a worker killed while holding the baton (exec timeout, forked test teardown, the 150m phase timeout) leaves the lock file behind and every later load blocks forever in file_baton.wait — the xdist workers wedge in place and the phase times out without a summary. The wedge followed whichever test first hit a leftover lock, which is why it seemed to move between runs. Warm the op in the launcher (and clear a stale baton when the op is already built) so workers only hit the disk cache, and pre-build the hot ops in CI before the suite starts. Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
The dtype=float16 parametrizations of TestNoSyncCtxt crash
deepspeed.initialize's sanity check ("Type fp16 is not supported on
your device.") on accelerators without fp16 support — the same
hardware lottery deepspeedai#8398 handled elsewhere. Guard the three
dtype-parametrized methods so the gap is a skip, not eight failures;
stages 2/3 additionally never reach their expected no_sync
AssertionError on such hosts.
Verified locally on an fp16-incapable CPU: 8 fail -> 8 skip, all 17
remaining parametrizations pass.
Signed-off-by: Guokai Ma <guokai.ma@intel.com>
ZeRO-3 also rounds the partitioned parameters through the autocast dtype on every all-gather, so the engine's forward runs on fp16-rounded weights while the fp32 baseline rounds per op; the loss carries percent-level dtype noise (measured 0.4%-2.7% across runs) instead of the sub-ulp noise covered below ZeRO-3. Extend the parity tolerance accordingly (rtol=5e-2, atol=5e-2 with 2x margin over the measured band); bf16 and fp32 paths are unchanged. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The default preset (2.10.0) predates four stable releases and carries a known CPU SDPA integer-divide-by-zero on zero-head tensors, which kills a rank with SIGFPE in the uneven-TP tests (trailing ranks hold size-0 attention shards; fixed in newer kernels, verified on 2.13). Add the 2.14.0-cpu preset case and make it the default for this branch so the experiment runs against a current torch; the older presets stay selectable for manual dispatch. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The plain cpu-torch-latest run already guards every single-rank test, so this branch's budget should go only to the multi-rank tests its LOCAL_SIZE gate admits. Gate a collection filter on DS_MULTIRANK_ONLY: it deselects non-DistributedTest items and world_size<=1 classes (honoring the per-test world_size mark with the launcher's precedence) and the workflow sets it. This also gives the non-v1 half enough of its time budget to expose the full multi-rank backlog instead of a slice. Groundwork for a dedicated cpu-torch-multi-latest workflow that will own exactly this selection. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Two latent bugs surfaced by the multi-rank CPU run (gloo): - register_with_transformers validated core_attn_implementation against a bare AutoConfig's _attn_implementation when given a string model path. A config that has not gone through model loading still holds the unresolved 'eager' default, so every string-path caller asking for sdpa/flex/FA2 was rejected with a spurious mismatch. Only compare when the caller passed an actual model, whose config carries the implementation resolved at load time. - UlyssesSPDataLoaderAdapter exchanged the local sequence length as a 0-dim tensor but allocated 1-element receive buffers. Backends that move raw bytes (nccl) tolerate that; gloo validates gather shapes strictly and rejects it. Send a 1-element tensor to match. With both fixed the sdpa Ulysses SP tests pass end to end on cpu/gloo (verified 2/2 locally, incl. the numerical-parity assertions). Signed-off-by: Guokai Ma <guokai.ma@intel.com>
With the engine's string-path registration fixed, the sdpa classes run on cpu/gloo. Guard the CUDA-only remainder (the disable-in-eval test hardcodes cuda:<rank> tensors, flex_attention needs CUDA) and make the mock PEFT model carry the load-resolved attention implementation a real PEFT wrapper would have, instead of the bare-config 'eager' default that trips the consistency check. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The multi-rank-only selection leaves enough of the 6h job budget to let the non-v1 half finish its last 5% instead of losing its failure report to the 150m kill. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The test pinned its tensors and models to cuda:<rank> although nothing in it is CUDA-specific (it runs sdpa); use the accelerator device like the rest of the file, so it runs on any backend instead of being skipped. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
TestMoECheckpoint hardcodes fp16 configs, so deepspeed.initialize's sanity check fails on accelerators without fp16 support — the same hardware lottery deepspeedai#8398 handled elsewhere. Add the class-level skipif. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
TestZeroPartialOffloadConfigSweep hardcodes fp16 and test_checkpoint_pipe_engine enables it for zero_stage > 0; both trip initialize's sanity check on accelerators without fp16 support (deepspeedai#8398's hardware lottery). Skip those parametrizations so the gap is a skip. With this, every fp16-config test in the suite guards on is_fp16_supported(). Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The 180m guard still cut the final 6% of the multi-rank non-v1 half on a slow runner (zero test failures up to the kill; the job only failed on the timeout's exit code). 210m keeps the whole job, worst case, inside the 6h limit while letting the half finish and report green. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
This is an on going work. Note I'll rebase when revealed bugs are fixed. The goal is fix all issues exposed by this CI, then turn on multi-rank test in CPU workflow.
Problem
On the
cpu-torch-latestrunner (single-socketubuntu-24.04), every unit test that needs more than one rank is skipped:Root cause:
CPU_Accelerator.device_count()reports the number of NUMA nodes (accelerator/cpu_accelerator.py), which is 1 on the single-socket runner, andDistributedExec._launch_procs()intests/unit/common.pygates process count ondevice_count(). Multi-rank CPU tests do not actually need one device per rank — ranks are ordinary processes communicating over gloo.Change
Set
LOCAL_SIZE=4for theunit-testsjob.device_count()readsLOCAL_SIZEfirst, so the launch gate now admitsworld_size<=4tests. This is the same signal the DeepSpeed launcher sets for spawned processes; the unit-test harness re-setsLOCAL_SIZEper worker in_dist_run(), so the CI-level value only affects the gate and cannot leak into test bodies.Evidence this is safe
TestDistIsendIrecv(tests/unit/comm/test_dist.py, world_size=2) and the autotp universal-checkpoint test (tests/unit/checkpoint/test_autotp_uc_checkpoint.py, world_size=4) already bypass the per-device gate for CPU and run green in this very CI job.LOCAL_SIZEhas a single reader in the codebase (CPU_Accelerator.device_count()); the launcher only writes it.Expected impact
~207 world_size=2 and ~85 world_size=4 tests that are currently skipped will now execute on CPU CI (world_size>=8 stays skipped). Some of them may have latent failures — this PR intentionally surfaces them so they can be triaged.