Skip to content

[DON'T MERGE] Run multi-rank CPU unit tests in CI via LOCAL_SIZE - #8381

Open
delock wants to merge 37 commits into
deepspeedai:masterfrom
delock:ci/cpu-multi-rank-local-size
Open

delock wants to merge 37 commits into
deepspeedai:masterfrom
delock:ci/cpu-multi-rank-local-size

Conversation

@delock

@delock delock commented Sep 1, 2026 •

Copy link
Copy Markdown
Collaborator

This is an on going work. Note I'll rebase when revealed bugs are fixed. The goal is fix all issues exposed by this CI, then turn on multi-rank test in CPU workflow.

Problem

On the cpu-torch-latest runner (single-socket ubuntu-24.04), every unit test that needs more than one rank is skipped:

SKIPPED tests/unit/common.py:275: Skipping test because not enough GPUs are available: 2 required, 1 available

Root cause: CPU_Accelerator.device_count() reports the number of NUMA nodes (accelerator/cpu_accelerator.py), which is 1 on the single-socket runner, and DistributedExec._launch_procs() in tests/unit/common.py gates process count on device_count(). Multi-rank CPU tests do not actually need one device per rank — ranks are ordinary processes communicating over gloo.

Change

Set LOCAL_SIZE=4 for the unit-tests job. device_count() reads LOCAL_SIZE first, so the launch gate now admits world_size<=4 tests. This is the same signal the DeepSpeed launcher sets for spawned processes; the unit-test harness re-sets LOCAL_SIZE per worker in _dist_run(), so the CI-level value only affects the gate and cannot leak into test bodies.

Evidence this is safe

  • TestDistIsendIrecv (tests/unit/comm/test_dist.py, world_size=2) and the autotp universal-checkpoint test (tests/unit/checkpoint/test_autotp_uc_checkpoint.py, world_size=4) already bypass the per-device gate for CPU and run green in this very CI job.
  • LOCAL_SIZE has a single reader in the codebase (CPU_Accelerator.device_count()); the launcher only writes it.

Expected impact

~207 world_size=2 and ~85 world_size=4 tests that are currently skipped will now execute on CPU CI (world_size>=8 stays skipped). Some of them may have latent failures — this PR intentionally surfaces them so they can be triaged.

@delock
delock requested a review from loadams as a code owner September 1, 2026 01:21

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: af3841378c

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

# multi-rank test. CPU ranks are plain processes over gloo, so advertise 4
# local devices to let world_size<=4 tests run. The test harness re-sets
# LOCAL_SIZE per worker, so this value only affects the launch gate.
LOCAL_SIZE: '4'

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep CUDA-only distributed tests out of the CPU run

In the cpu-torch-latest job, advertising four devices admits every distributed test with world_size <= 4, not only CPU-safe tests. For example, tests/unit/ulysses_alst/test_ulysses_sp_hf.py:240-263 defines an unguarded two-rank test that creates tensors on cuda:<rank>; because CPU_Accelerator.is_available() returns true, the harness does not skip it, and the all-unit pytest invocation at workflow line 283 will fail on the CPU-only PyTorch installation. Scope this override to an explicitly CPU-compatible subset or add CPU capability checks before enabling the previously skipped tests.

Useful? React with 👍 / 👎.

# multi-rank test. CPU ranks are plain processes over gloo, so advertise 4
# local devices to let world_size<=4 tests run. The test harness re-sets
# LOCAL_SIZE per worker, so this value only affects the launch gate.
LOCAL_SIZE: '4'

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Add the required sign-off trailer

This is a non-merge commit, but its commit message has no Signed-off-by trailer. Add the author sign-off so the commit satisfies the repository's commit and CI requirements.

AGENTS.md reference: AGENTS.md:L8-L8

Useful? React with 👍 / 👎.

@delock

delock commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator Author

This is an on going work. Note I'll rebase when revealed bugs are fixed. The goal is fix all issues exposed by this CI, then turn on multi-rank test in CPU workflow.

@delock

delock commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator Author

Experiment results: enabling multi-rank CPU tests via LOCAL_SIZE=4

Baseline comparison: same-commit master run (all multi-rank tests skipped) vs this branch. Zero regressions: every failure is a test that was previously skipped.

What now runs

~1800 previously-skipped multi-rank tests execute on the CPU runner (gloo). Final complete numbers from the last run (split invocations):

  • unit/v1 half (complete): 916 passed / 131 failed / 206 skipped
  • sequential (complete): 54 passed / 9 failed
  • non-v1 half: killed by the 150m guard (still too slow on 4 vCPU), 59 failures listed inline

Failure inventory & root causes (triaged)

Count Test Root cause
92 v1/zero/test_offload_states.py Asserts memory_allocated() drops after offload — a VRAM concept. CPU_Accelerator.memory_allocated() returns RSS, which never shrinks on free. Test/accelerator semantic gap; needs discussion (gate on device semantics or a better CPU metric).
24 v1/zero/test_zero_autocast.py Hardcoded dist_backend='nccl' + a CUDA/NCCL-oriented bf16 gate. Needs accelerator-derived backend + capability-based skip.
25 ulysses SP tests Numerical mismatches on CPU (real correctness questions, need deep dive).
16 onebit optimizer tests NoneType.size at onebit/adam.py:108 — runtime bug on CPU path.
~20 pipe / zeropp / moe-checkpoint / coalesce Smaller follow-ups, same GPU-assumption patterns.

Test-side fixes already in this branch (validated by CI)

  • All 5 DDP reference sites route through wrap_ddp_reference() (CPU modules must not pin device_ids) — dropped this class from 114 failures to 7.
  • reduce_boolean_flags / autotp tests use current_device_name() instead of current_device() (a LOCAL_RANK string on CPU, not a device).
  • fp16-config tests get a capability skipif (not get_accelerator().is_fp16_supported()) — GH ubuntu-24.04 runners are hardware-heterogeneous w.r.t. AVX512-FP16, so hardcoded fp16 was a runner lottery.
  • fork_rng probes the device module for a per-device RNG instead of matching accelerator names.

CI infrastructure findings (need maintainer decisions)

  1. --maxfail=100 in PYTEST_OPTS + no timeout in DistributedExec._close_pool = permanent wedge: when the 100th failure trips the interrupt, teardown (pool.starmap(_dist_destroy), close/join) can block forever on wedged gloo peers; the run then dies at the 6h job limit without printing any summary. This is what cancelled the first three attempts. (This branch works around it by splitting the suite and raising maxfail.)
  2. Suite size vs 4 vCPU: the full multi-rank suite does not fit one 6h job. Options: split invocations (done here), lower -n, or a dedicated larger runner.
  3. CPU_Accelerator.device_count()'s NUMA-node semantics remain the root cause of the original gap; the long-term fix could be a test-harness-level exemption for CPU (ranks are processes, not devices).

Happy to split the commits into separate PRs (workflow change / mechanical test fixes / triage follow-ups) per maintainer preference — the failure inventory above is intended as the working list.

@delock

delock commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator Author

Given the CPU UT would run ~hrs after enabling LOCAL_SIZE=4, I wouldn't suggest to turn on this in master. However it is beneficial to use this branch to reveal problems in CPU accelerator and fix them. I'm mark this PR as do not merge.

@delock delock changed the title Run multi-rank CPU unit tests in CI via LOCAL_SIZE [DON'T MERGE] Run multi-rank CPU unit tests in CI via LOCAL_SIZE Sep 3, 2026
banxingmjj pushed a commit to openanolis/DeepSpeed that referenced this pull request Sep 3, 2026
…dai#8398)

## Problem

`TestMultipleModels::test_zero_optimizer`, `TestSimpleMoE`, `TestMoE`,
`TestPRMoE`, and `TestMOETensorParallel` hardcode `"fp16": {"enabled":
True}` in their DeepSpeed configs. The engine's sanity check then
raises:

```
ValueError: Type fp16 is not supported on your device.
```

on any accelerator whose `is_fp16_supported()` is false. On CPU that
maps to the AVX512-FP16 capability of the host, and GitHub's
`ubuntu-24.04` runners are hardware-heterogeneous: **the same test
passes on one runner and fails on the next** (observed directly in deepspeedai#8381
— 146 failures appeared on one runner generation and none on another,
with identical code).

## Change

Skip these tests via a capability query:

```python
@pytest.mark.skipif(not get_accelerator().is_fp16_supported(), reason="fp16 is not supported on this accelerator")
```

- capability only, no accelerator-name matching;
- mirrors the existing bf16 skip precedent in
`tests/unit/v1/zero/test_zero_user_backward.py`;
- deliberately a **skip** rather than silently running bf16 — these
tests exist to cover the fp16 paths.

## Validation

Validated as part of the multi-rank CPU CI experiment in deepspeedai#8381: the 146
hardware-lottery failures became deterministic skips, zero regressions
on previously-passing tests.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
banxingmjj pushed a commit to openanolis/DeepSpeed that referenced this pull request Sep 3, 2026
## Problem

`tests/unit/v1/zero/test_zero_user_backward.py` builds torch DDP
**reference** models (the known-good baseline that ZeRO results are
compared against) at five sites:

```python
model_ddp = DDP(model_ddp, device_ids=[rank], output_device=rank)
```

`device_ids=[rank]` assumes rank ↔ GPU index. torch's DDP contract only
allows `device_ids`/`output_device` for single-device GPU modules; **CPU
modules live on one shared device and must omit them**, so multi-rank
CPU runs died inside the DDP constructor with:

```
ValueError: DistributedDataParallel device_ids and output_device arguments only work with single-device/multiple-device GPU modules or CPU modules, ...
```

## Change

Route all five sites through one helper:

```python
def wrap_ddp_reference(model, device, rank):
    # Only indexed devices take device_ids/output_device; CPU modules live on one shared device.
    if torch.device(device).type == 'cpu':
        return DDP(model)
    return DDP(model, device_ids=[rank], output_device=rank)
```

Design notes:

- the condition restates the exact precondition torch's own DDP
constructor enforces, in torch's device vocabulary — it follows the
model's actual device rather than the global accelerator configuration;
- a single helper means new reference-model sites cannot forget the
branch (the first fix round in deepspeedai#8381 missed 4 of the 5 sites for exactly
this reason);
- GPU behavior is unchanged.

## Validation

Validated as part of the multi-rank CPU CI experiment in deepspeedai#8381: all
DDP-constructor failures were eliminated (the file's few remaining
failures there are unrelated — see the triage table in that PR), zero
regressions vs the same-commit baseline.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
banxingmjj pushed a commit to openanolis/DeepSpeed that referenced this pull request Sep 3, 2026
…eepspeedai#8397)

## Problem

`get_accelerator().current_device()` returns a **device index** on GPU
backends (`torch.cuda.current_device()` → int), but on CPU it returns
the `LOCAL_RANK` environment value — a plain **string** like `'1'`. Two
test-side consumers fed that value straight into tensor/device
placement:

- `reduce_boolean_flags` in `tests/unit/common.py` (backbone of
`allclose_on_all_ranks`, the "all ranks succeed or fail together" check)
- 15 call sites in `tests/unit/v1/autotp/test_autotp_training.py`

On CPU this fails immediately with `RuntimeError: Invalid device string:
'1'` — before the first collective even runs.

## Change

- Use `current_device_name()`, which returns a full device string on
every backend (`'cpu'`, `'cuda:N'`, `'mps:0'`, …) and is equivalent to
the index on GPU backends.
- In `reduce_boolean_flags`, carry the flag in a 1-dim tensor: gloo
rejects 0-dim inputs to `all_gather_into_tensor` (NCCL tolerates them),
so the previous form would have failed on the very next line.

## Validation

Validated as part of the multi-rank CPU CI experiment in deepspeedai#8381
(same-commit baseline comparison): this failure class disappeared, zero
regressions on previously-passing tests.

---------

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
cpu-torch-latest runs on a single-socket runner where
CPU_Accelerator.device_count() reports 1 NUMA node, so the
per-device gate in tests/unit/common.py skips every test that
needs more than one rank. CPU ranks are plain processes over
gloo and need no per-rank hardware, so advertise 4 local
devices via LOCAL_SIZE, the env var device_count() reads first.
The test harness re-sets LOCAL_SIZE per worker, so this value
only affects the launch gate.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
With LOCAL_SIZE=4 the suite now runs to ~63% and then all xdist workers
go silent for hours until the 6h job limit cancels the run - the pool
worker cleanup hang that DS_DISABLE_REUSE_DIST_ENV was added for. Fresh
pools per test cost some wall time but let the run finish and print the
failure summary.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Both full-suite attempts wedge before printing a summary: once the
100th failure trips PYTEST_OPTS' --maxfail, pytest-xdist's interrupt
path stalls forever in mp pool teardown (no timeout guards _close_pool),
and the 6h job limit cancels the run. Run the suite as two fresh-worker
halves, override maxfail so all failures are listed, and cap each half
with timeout so the sequential tail always runs.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
@delock
delock force-pushed the ci/cpu-multi-rank-local-size branch from f46af32 to f2793e7 Compare September 3, 2026 14:01
banxingmjj pushed a commit to openanolis/DeepSpeed that referenced this pull request Sep 15, 2026
…epspeedai#8407)

## Description

`train_cifar` unconditionally passed
`devices=[get_accelerator().current_device_name()]` to `fork_rng`. On
CPU that becomes `devices=['cpu']`, and `torch.cpu` has no
`get_rng_state`, so the context manager raises `AttributeError` before
the test body starts — the CPU RNG lives in the global generator that
`fork_rng` already forks.

Probe the device module for a per-device RNG instead of matching
accelerator names:

- backends whose device module has `get_rng_state` (e.g. cuda) behave
exactly as before;
- backends without one (cpu) pass `devices=[]`, which changes nothing
beyond the global-RNG save/restore `fork_rng` always performs.

The crash is only reachable from the multi-rank tests that call
`train_cifar` (`test_onebit.py`, `test_pipe.py`), which the CPU runner
currently skips at the device gate; it was exposed by the `LOCAL_SIZE=4`
experiment in deepspeedai#8381. Sibling fixes from the same series landed as deepspeedai#8397,
deepspeedai#8398 and deepspeedai#8399.

## Test plan

- pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the
changed file;
- exercised under the `LOCAL_SIZE=4` multi-rank CPU run tracked in
deepspeedai#8381: the `test_onebit` / `test_pipe` callers got past `fork_rng` and
into their test bodies with this exact change.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
banxingmjj pushed a commit to openanolis/DeepSpeed that referenced this pull request Sep 15, 2026
…peedai#8409)

## Description

The dynamic offload-state tests assert strict allocated-memory deltas
around `offload_states()` / `reload_states()`:

- `alloc_after_offload < alloc_before_offload`
- `alloc_after_reload > alloc_after_offload`

That contract assumes `memory_allocated()` is allocator bookkeeping,
which holds on cuda (`torch.cuda.memory_allocated()`). On cpu,
`CPU_Accelerator.memory_allocated()` reports process RSS (psutil), and
RSS does not shrink when tensors are freed — so all 92 parameterized
cases fail even when the offload itself is correct (the device-placement
and data-integrity checks in the same tests pass).

Gate only the memory-delta asserts on whether the accelerator's torch
device module exposes `memory_allocated` (cuda does; `torch.cpu` does
not), mirroring the capability probe used for `fork_rng` in
`train_cifar` (deepspeedai#8407):

- cuda and other allocator-backed backends: behavior unchanged
- cpu: the unobservable deltas are skipped; all device-placement
validations still run

Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381 (92 of the 131
v1-half failures there).

## Validation (executed on real hardware)

- 20-core x86_64 CPU, torch 2.13.0+cpu, gloo backend, 2 ranks
(`LOCAL_SIZE=2`)
- Before:
`TestDynamicOffloadStatesZero12[False-1-False-False-optim_states]` fails
on the persistent-state delta assert
- After: 5 representative cases pass (persistent and grad paths, ZeRO
stage 1/2/3, `static_offload_optimizer=True` branch) — 5 passed in 55.6s
- pre-commit (yapf / flake8 / check-torchdist / codespell) passes on the
changed file

Sibling PRs from the same series: deepspeedai#8397, deepspeedai#8398, deepspeedai#8399, deepspeedai#8407.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
@delock

delock commented Sep 16, 2026

Copy link
Copy Markdown
Collaborator Author

Status update

Fixes extracted so far (spin-off PRs)

Investigating the failures exposed by this CI change surfaced three real bugs, each with a standalone fix:

  1. split_half_float_double() silently drops all gradient buckets on CPU — legacy type strings built from the accelerator device name never match CPU tensors (torch.FloatTensor has no device prefix), so with contiguous_gradients=False + reduce_scatter=False ZeRO-1/2 skips the gradient all-reduce entirely; mean configs mask it, sum is off by world_size. Fix in Fix silent gradient-bucket drop on CPU in ZeRO-1/2 non-contiguous reduction #8382 (also fixes the TestCoalesceCollectiveCount "baseline issued no collectives" failure).
  2. Test-harness comparison helper crashes on CPU — reduce_boolean_flags() passed current_device() (a bare rank string) to torch and used a 0-dim all-gather that gloo rejects. The current_device_name() part already landed on master via Use device names, not rank ids, for device placement in test helpers #8397; the rest is in Fix silent gradient-bucket drop on CPU in ZeRO-1/2 non-contiguous reduction #8382.
  3. ZeRO-3 coalesced parameter gather only waits for the last async handle — gloo handles are independent and CPU synchronize() is a no-op, so earlier gathers can still be in flight when param.data is read, exposing uninitialized memory as NaN. Reproduced flaky (9/20 → 0/20 with the fix); a separate PR for partition_parameters.py is being prepared.

Current failures on this branch (84 total: 75 in the unit/v1 half + 9 sequential)

Current focus: phase-1 timeout / worker wedge

The job takes 4h05m and the first half (unit/ non-v1, timeout 150m) never finishes — it reaches 94% and a worker wedges for ~40min until the timeout kills it, losing that half's results. From the job log:

  • gw1 claims TestDistInferenceAllReduce::test[128-dtype2] (fp16) at 09:54:17 and produces no output for the remaining 111 minutes;
  • gw0/gw2 wedge later, but the tests they were running were later passed by other workers — secondary casualties of state left by the first wedge;
  • gw3 runs alone until the timeout.

Prime suspect is dist.inference_all_reduce() on an fp16 tensor over CPU/gloo. Investigating next; fixing this should also bring the job back to ~2h45m and recover phase-1 results.

The shm-based inference_all_reduce keys its /dev/shm segments by
MASTER_ADDR/MASTER_PORT. The bind-and-release probe in get_master_port()
hands every pool the same first-free port when tests run one process each
(--forked), so concurrent pools share segments: ranks of different pools
cross-write the collective state and produce wrong results or spin
forever (the CI worker wedge). Also the op never unlinks its ~66MB-per-
rank segments, which exhausts /dev/shm over a full run.

Offset the port probe per process and per pool so each pool gets a
distinct key, and remove the segments once the pool (or spawned procs) is done.

Verified with two concurrent gloo pools calling inference_all_reduce:
shared key gives wrong results (rank sees 2.0 instead of 3.0), distinct
keys are correct; three rounds of -n 2 --forked TestDistInferenceAllReduce
pass with zero /dev/shm leftovers.

Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
Concurrent xdist workers each downloading bert-base-uncased race on the
transformers cache lock and fail with PermissionError. Populate a shared
HF_HOME once before the suite and point both unit-test halves at it (the
sequential phase already used it), so tests only read the cache.

Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
huggingface-cli is now a deprecation stub in the runner's
huggingface_hub: it prints a warning and exits 1, killing the job in
the pre-download step before any test runs. Use its replacement, hf
download, which is already installed on the runner.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
bert-base-uncased is stored on Xet, and the anonymous xet-read-token
request 404s, so hf download dies before populating the cache. Fall
back to the classic resolve endpoint, which works without a token.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
(cherry picked from commit de667d5)
Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
(cherry picked from commit 22b50e5)
Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
(cherry picked from commit 2e089a1)
pull Bot pushed a commit to fish003/DeepSpeed that referenced this pull request Sep 21, 2026
…eepspeedai#8539)

## Problem

While investigating multi-rank CPU CI (deepspeedai#8381),
`TestUnmanagedGradientAccumulation::test_unmanaged_varying_backward_count[3]`
failed intermittently with parameters turned to NaN, and
`test_zero_coalesce_grad_reduction` showed mismatches with denormal
garbage values (`4.8e+37`, `1e-38`) — the signature of uninitialized
memory.

## Root cause

`_allgather_params_coalesced()` launches one async all_gather per
parameter, but only waits on the **last** handle before pointing
`param.data` at the `torch.empty` flat buffers:

```python
launch_handles[-1].wait()
param.data = gathered_tensor.narrow(...)
if not get_accelerator().resolves_data_dependency():
    get_accelerator().synchronize()
```

On CUDA this is safe by construction: the ops are enqueued from a single
thread onto the same stream (FIFO), and `torch.cuda.synchronize()` waits
for everything anyway. On gloo each async handle runs on an independent
background thread with **no ordering between handles** — waiting for the
last one says nothing about the earlier ones — and the CPU accelerator's
`synchronize()` is a no-op. Reading `param.data` can therefore race the
gather and expose uninitialized memory.

The corruption is timing-dependent (locally 9/20 runs fail; deepspeedai#8382's
comparison-helper fix made the numeric assertions reachable), which is
why it was never caught on GPU CI.

## Fix

Wait on every handle. On CUDA the earlier handles are guaranteed
complete by stream ordering, so the extra waits are immediate status
checks with no real waiting; the change makes correctness independent of
that structural coincidence.

## Verification (CPU/gloo, world sizes 2–3)

| Test | Before | After |
|---|---|---|
| `test_unmanaged_varying_backward_count[3]` | 9/20 failed | 0/20 failed
|
| `TestZero3ParamPartitioningBase` / `TestGradientAllreduceOp` /
`TestUnmanagedGradientAccumulation` (regression, 57 cases) | — | 55 pass
+ 2 skip |

GPU behavior is unchanged by construction (same buckets, same order,
stream-FIFO + full synchronize); GPU CI will re-confirm.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
The baseline DistributedFixture builds an fp16 engine config for every
dtype=float16 parametrization, and deepspeed.initialize's sanity check
raises "Type fp16 is not supported on your device." on accelerators
without fp16 support. Because the failure happens inside the fixture's
distributed run, pytest reports a setup ERROR for every dependent test
(48 on the multi-rank CPU run) instead of a skip.

Skip inside the fixture like deepspeedai#8398 did for the fp16-config tests;
DistributedFixture propagates the skip to all dependent tests. Devices
with fp16 support are unchanged.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
pull Bot pushed a commit to bhardwajRahul/DeepSpeed that referenced this pull request Sep 22, 2026
…dai#8610)

Follow up deepspeedai#8381 @delock 

**Fix zeropp 9 (sequential)** in
deepspeedai#8381 (comment)
```
unit/runtime/zero/test_zeropp.py
```
<img width="624" height="349" alt="image"
src="https://github.com/user-attachments/assets/85750e15-91ca-4290-ae90-734302f20324"
/>

**passed in
https://github.com/deepspeedai/DeepSpeed/actions/runs/35571162595/job/106242965025**

This pull request updates the `tests/unit/runtime/zero/test_zeropp.py`
test suite to ensure that certain tests are only run on accelerators
that support FP16 precision. This prevents test failures on hardware
that does not support FP16.

Test robustness improvements:

* Added `@pytest.mark.skipif` decorators to the `test`, `test_eval`, and
`test_gradient_accumulation` methods in `TestZeroPPConfigSweep` and a
`test` method in another test class, so these tests are skipped if the
accelerator does not support FP16. This uses
`get_accelerator().is_fp16_supported()` to check hardware capability.
[[1]](diffhunk://#diff-e3c08df84c5c79ec0d57c31a7817cddb1ba2f31c620c482eaf4561698ffc3f27R173)
[[2]](diffhunk://#diff-e3c08df84c5c79ec0d57c31a7817cddb1ba2f31c620c482eaf4561698ffc3f27R216)
[[3]](diffhunk://#diff-e3c08df84c5c79ec0d57c31a7817cddb1ba2f31c620c482eaf4561698ffc3f27R258)
[[4]](diffhunk://#diff-e3c08df84c5c79ec0d57c31a7817cddb1ba2f31c620c482eaf4561698ffc3f27R384)
* Imported `get_accelerator` from `deepspeed.accelerator` to enable the
FP16 support check.

---------

Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
VenusTZZ pushed a commit to VenusTZZ/DeepSpeed that referenced this pull request Sep 22, 2026
Follow up deepspeedai#8381 @delock 

fix
```
unit/pipe/test_pipe_module.py
unit/runtime/pipe/test_pipe.py
```


This pull request refactors device handling in the activation
checkpointing code to simplify and standardize how device variables are
used. The main change is to consistently use a `device` variable instead
of `cuda_device` and to remove unused or redundant stream-related code.
This improves code clarity and maintainability without changing
functionality.

**Device handling simplification:**

* Replaced all occurrences of `cuda_device` with a unified `device`
variable obtained from `get_accelerator().current_device_name()` in
`checkpointing.py`.
[[1]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL604-R604)
[[2]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL838-R831)
[[3]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL941-R933)
* Updated all device argument usages in calls to `copy_to_device`,
`gather_partitioned_activations`, and `move_to_device` to use the new
`device` variable.
[[1]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL618-R617)
[[2]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL709-R714)
[[3]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL851-R843)
[[4]](diffhunk://#diff-a4333224075c38d4a6f6aa97123c10f93e519b39cce1f02ecaa55d881493bb2dL952-R950)

**Stream and code cleanup:**

* Removed creation and usage of the `transport_stream` variable, and
deleted commented-out or unused stream synchronization code.

These changes make the device management more straightforward and reduce
potential confusion around device and stream handling in activation
checkpointing.

Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
jinyouzhi added a commit to jinyouzhi/DeepSpeed that referenced this pull request Sep 23, 2026
Follow up deepspeedai#8381  @delock 
Fix CPU CI failures in
https://github.com/deepspeedai/DeepSpeed/actions/runs/35675316100/job/106580600974?pr=8381
```
FAILED unit/inference/quantization/test_intX_quantization.py::TestQuantizedInt::test_half_int8_quantization 
FAILED unit/checkpoint/test_convert_checkpoint.py::TestCheckpointConvert::test_convert_zero_checkpoint_to_fp32_state_dict 
```

---------

Signed-off-by: Jin, Youzhi <youzhi.jin@intel.com>
TestCoalesceFP16 forces an fp16 config, and deepspeed.initialize's
sanity check raises "Type fp16 is not supported on your device." on
accelerators without fp16 support — the same hardware lottery deepspeedai#8398
handled for the other fp16-config tests. Add the same class-level
skipif so the gap is a skip, not three failures.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The multiple-backward checkpointing test compares DDP and engine
gradients over three optimizer iterations at dtype-default tolerances.
The engine keeps fp32 accounting while the DDP reference steps the bf16
weights directly, so the two paths diverge by roughly one bf16 ulp of
weight per step even from exactly equal gradients; by iteration 2 that
crosses a reduction/relu rounding boundary and produces ~8e-3 absolute
gradient differences on O(1) gradients. Whether the difference survives
bf16 quantization depends on the CPU bf16 kernel's rounding, so the
test flips between pass and fail across torch versions (fails on
2.10/2.11, passes on 2.13) without any behavioral change.

Give iterations >= 2 an absolute tolerance floor (atol=1e-2 alongside
the bf16-default rtol): the test already documents that late iterations
drift, and the floor absorbs quantization-boundary chaos while still
catching real gradient errors. Verified on both kernel families: 6/6
fail -> 6/6 pass on torch 2.11, unchanged 6/6 pass on torch 2.13.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
pull Bot pushed a commit to AmirulAndalib/DeepSpeed that referenced this pull request Sep 24, 2026
deepspeedai#8648)

## Description


`TestZeroUserBackwardWithCheckpointing::test_checkpointed_multiple_backward`
compares DDP and engine gradients over three optimizer iterations at
dtype-default tolerances, and flips between pass and fail across torch
versions on cpu (fails on 2.10/2.11 kernels, passes on 2.13) without any
behavioral change.

A minimal two-rank repro isolates the mechanism:

```
                   torch 2.13          torch 2.11 (= CI family)
iter 0 grads       exact match         exact match
post-step          —                   linear2.weight differs by 1.7e-6   <- the seed
iter 1 grads       exact match         exact match (quantization absorbs it)
iter 2 grads       exact match         7.8e-3                              <- exceeds rtol
final weights      1 bf16 ulp          1 bf16 ulp
```

The engine keeps fp32 accounting while the DDP reference steps the bf16
weights directly, so the two mathematically-equivalent paths diverge by
~one bf16 ulp of weight per optimizer step even from bit-identical
gradients; by iteration 2 that crosses a reduction/relu rounding
boundary. Whether the difference survives bf16 quantization depends on
the CPU bf16 kernel rounding — hence the version dependence.

This PR gives iterations >= 2 an absolute tolerance floor (`atol=1e-2`
alongside the bf16-default `rtol=1.6e-2`). The test already documents
"small differences at later iterations are expected due to bfloat16
precision"; the floor absorbs quantization-boundary chaos while still
catching real gradient errors.

## Validation (executed on real hardware)

- 20-core x86_64 CPU, gloo, 2 ranks (`LOCAL_SIZE=2`)
- torch 2.11 (the failing kernel family): 6/6 parametrizations **fail ->
pass**
- torch 2.13 (the passing family): 6/6 pass before and after (no
regression)
- pre-commit passes on the changed file

Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in deepspeedai#8381.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
With torch_autocast the engine reduces gradients in the autocast dtype
while the fp32 DDP baseline reduces in fp32. On gloo an fp16 collective
round-trips through fp32, so after one optimizer step the step-1 loss
carries fp16-ulp relative differences (observed 2e-4..8e-4) that the
default tolerance rejects even though step-0 matches bit-exact. Allow a
low-precision tolerance for fp16 below ZeRO-3; the ZeRO-3 fp16 variants
keep the strict tolerance (their CPU divergence is much larger and under
investigation).

Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
Every rank worker JIT-loads deepspeed_shm_comm on first init_distributed.
Concurrent loads serialize on torch's FileBaton, which has no stale-lock
handling: a worker killed while holding the baton (exec timeout, forked
test teardown, the 150m phase timeout) leaves the lock file behind and
every later load blocks forever in file_baton.wait — the xdist workers
wedge in place and the phase times out without a summary. The wedge
followed whichever test first hit a leftover lock, which is why it
seemed to move between runs.

Warm the op in the launcher (and clear a stale baton when the op is
already built) so workers only hit the disk cache, and pre-build the hot
ops in CI before the suite starts.

Signed-off-by: Ma, Guokai <guokai.ma@intel.com>
The dtype=float16 parametrizations of TestNoSyncCtxt crash
deepspeed.initialize's sanity check ("Type fp16 is not supported on
your device.") on accelerators without fp16 support — the same
hardware lottery deepspeedai#8398 handled elsewhere. Guard the three
dtype-parametrized methods so the gap is a skip, not eight failures;
stages 2/3 additionally never reach their expected no_sync
AssertionError on such hosts.

Verified locally on an fp16-incapable CPU: 8 fail -> 8 skip, all 17
remaining parametrizations pass.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
ZeRO-3 also rounds the partitioned parameters through the autocast
dtype on every all-gather, so the engine's forward runs on fp16-rounded
weights while the fp32 baseline rounds per op; the loss carries
percent-level dtype noise (measured 0.4%-2.7% across runs) instead of
the sub-ulp noise covered below ZeRO-3. Extend the parity tolerance
accordingly (rtol=5e-2, atol=5e-2 with 2x margin over the measured
band); bf16 and fp32 paths are unchanged.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The default preset (2.10.0) predates four stable releases and carries a
known CPU SDPA integer-divide-by-zero on zero-head tensors, which kills
a rank with SIGFPE in the uneven-TP tests (trailing ranks hold size-0
attention shards; fixed in newer kernels, verified on 2.13). Add the
2.14.0-cpu preset case and make it the default for this branch so the
experiment runs against a current torch; the older presets stay
selectable for manual dispatch.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The plain cpu-torch-latest run already guards every single-rank test,
so this branch's budget should go only to the multi-rank tests its
LOCAL_SIZE gate admits. Gate a collection filter on DS_MULTIRANK_ONLY:
it deselects non-DistributedTest items and world_size<=1 classes
(honoring the per-test world_size mark with the launcher's precedence)
and the workflow sets it. This also gives the non-v1 half enough of its
time budget to expose the full multi-rank backlog instead of a slice.

Groundwork for a dedicated cpu-torch-multi-latest workflow that will
own exactly this selection.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Two latent bugs surfaced by the multi-rank CPU run (gloo):

- register_with_transformers validated core_attn_implementation against
  a bare AutoConfig's _attn_implementation when given a string model
  path. A config that has not gone through model loading still holds
  the unresolved 'eager' default, so every string-path caller asking
  for sdpa/flex/FA2 was rejected with a spurious mismatch. Only compare
  when the caller passed an actual model, whose config carries the
  implementation resolved at load time.

- UlyssesSPDataLoaderAdapter exchanged the local sequence length as a
  0-dim tensor but allocated 1-element receive buffers. Backends that
  move raw bytes (nccl) tolerate that; gloo validates gather shapes
  strictly and rejects it. Send a 1-element tensor to match.

With both fixed the sdpa Ulysses SP tests pass end to end on cpu/gloo
(verified 2/2 locally, incl. the numerical-parity assertions).

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
With the engine's string-path registration fixed, the sdpa classes run
on cpu/gloo. Guard the CUDA-only remainder (the disable-in-eval test
hardcodes cuda:<rank> tensors, flex_attention needs CUDA) and make the
mock PEFT model carry the load-resolved attention implementation a real
PEFT wrapper would have, instead of the bare-config 'eager' default
that trips the consistency check.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The multi-rank-only selection leaves enough of the 6h job budget to let
the non-v1 half finish its last 5% instead of losing its failure report
to the 150m kill.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The test pinned its tensors and models to cuda:<rank> although nothing in
it is CUDA-specific (it runs sdpa); use the accelerator device like the
rest of the file, so it runs on any backend instead of being skipped.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
TestMoECheckpoint hardcodes fp16 configs, so deepspeed.initialize's
sanity check fails on accelerators without fp16 support — the same
hardware lottery deepspeedai#8398 handled elsewhere. Add the class-level skipif.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
TestZeroPartialOffloadConfigSweep hardcodes fp16 and
test_checkpoint_pipe_engine enables it for zero_stage > 0; both trip
initialize's sanity check on accelerators without fp16 support (deepspeedai#8398's
hardware lottery). Skip those parametrizations so the gap is a skip.

With this, every fp16-config test in the suite guards on
is_fp16_supported().

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
The 180m guard still cut the final 6% of the multi-rank non-v1 half on
a slow runner (zero test failures up to the kill; the job only failed
on the timeout's exit code). 210m keeps the whole job, worst case,
inside the 6h limit while letting the half finish and report green.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants