-
Notifications
You must be signed in to change notification settings - Fork 5k
[DON'T MERGE] Run multi-rank CPU unit tests in CI via LOCAL_SIZE #8381
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
delock
wants to merge
37
commits into
deepspeedai:master
Choose a base branch
from
delock:ci/cpu-multi-rank-local-size
base: master
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
37 commits
Select commit
Hold shift + click to select a range
ec752ba
Run multi-rank CPU unit tests in CI via LOCAL_SIZE
delock f2b7d0b
Disable dist-env reuse for the full multi-rank CPU run
delock f2793e7
Split CPU unit tests into halves with per-half timeouts
delock 4d4019d
Merge branch 'master' into ci/cpu-multi-rank-local-size
delock 89ca163
Isolate shm-allreduce segments per test pool and clean them up
delock 453f60f
Pre-download HF fixtures and share one cache across all test phases
delock 6799643
Download HF fixtures with the hf CLI
delock 2654cac
Disable Xet transfers for the anonymous HF fixture download
delock d429f62
Merge remote-tracking branch 'origin/master' into pr-8381-new
delock 132298c
Merge remote-tracking branch 'origin/master' into pr-8381-new
delock dfb0c50
Scope the bf16 version floors to NCCL transports in the test check
delock 0194033
Merge remote-tracking branch 'origin/master' into pr-8381-new
delock 3fba8dc
Stop pinning DDP device_ids in the autocast baseline on CPU
delock 3b74fba
Merge remote-tracking branch 'origin/master' into pr-8381-new
delock 33e7dcc
Skip the dynamic offload-state tests on the cpu accelerator
delock 940d404
Merge remote-tracking branch 'origin/master' into ci/cpu-multi-rank-l…
delock 61abf25
Fix pipe tests failures on CPU
jinyouzhi 68fdd74
Skip fp16 ZeroPP tests on accelerators without fp16 support
jinyouzhi 2adfc0d
Skip ZERO++ Quantized in CPU
jinyouzhi 08c966f
Merge remote-tracking branch 'origin/master' into pr-8381-new
delock 09e3cba
Skip fp16 universal-checkpoint fixtures without fp16 support
delock 0cb5cdc
Merge remote-tracking branch 'origin/master' into pr-8381-new
delock 2b23590
Skip fp16 coalesce tests on accelerators without fp16 support
delock e4ec54f
Loosen the late-iteration gradient tolerance in the checkpointing test
delock 2bdbbdf
Allow fp16 ulp noise in autocast loss parity below ZeRO-3
delock 8fa6e22
Pre-build the shm comm op so rank workers never race the JIT lock
delock 18ee3ec
Skip fp16 no_sync tests on accelerators without fp16 support
delock ecc198e
Allow fp16 ulp noise in autocast loss parity at ZeRO-3
delock 5770829
Move the CPU workflow to torch 2.14.0 for this branch
delock a84e1b7
Run only the multi-rank tests in this workflow
delock 5914ddc
Fix Ulysses SP registration check and seqlen exchange on gloo
delock 86eeb05
Adapt the ulysses_sp_hf tests for non-CUDA accelerators
delock c81a577
Raise the non-v1 half timeout to 180m
delock a0680d6
De-hardcode the disable-in-eval test device
delock 1e6872a
Skip fp16 MoE checkpoint tests on accelerators without fp16 support
delock e6c3d95
Skip the remaining fp16-config tests without fp16 support
delock 30a96e4
Raise the non-v1 half timeout to 210m
delock File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -81,7 +81,20 @@ jobs: | |
| runs-on: ubuntu-24.04 | ||
|
|
||
| env: | ||
| DEFAULT_TORCH_PRESET: '2.10.0-cpu' | ||
| # The runner is single-socket, so CPU_Accelerator.device_count() reports 1 | ||
| # NUMA node and the per-device gate in tests/unit/common.py skips every | ||
| # multi-rank test. CPU ranks are plain processes over gloo, so advertise 4 | ||
| # local devices to let world_size<=4 tests run. The test harness re-sets | ||
| # LOCAL_SIZE per worker, so this value only affects the launch gate. | ||
| LOCAL_SIZE: '4' | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more.
This is a non-merge commit, but its commit message has no AGENTS.md reference: AGENTS.md:L8-L8 Useful? React with 👍 / 👎. |
||
| # Multi-rank tests churn mp pools for hours; reused pools eventually hang in | ||
| # cleanup and stall workers until the 6h job limit. Fresh pools per test are | ||
| # slower but let the suite finish (knob documented in tests/unit/common.py). | ||
| DS_DISABLE_REUSE_DIST_ENV: '1' | ||
| # Only the multi-rank tests this LOCAL_SIZE gate admits run here; single-rank | ||
| # tests are deselected because the plain cpu-torch-latest run guards them. | ||
| DS_MULTIRANK_ONLY: '1' | ||
| DEFAULT_TORCH_PRESET: '2.14.0-cpu' | ||
| DEFAULT_TRANSFORMERS_SOURCE: 'git' | ||
| # Manual PyPI fallback only; scheduled and default manual runs use Git. | ||
| DEFAULT_TRANSFORMERS_VERSION: '4.51.3' | ||
|
|
@@ -158,6 +171,11 @@ jobs: | |
| torchvision_install_version='0.25.0' | ||
| torch_test_version='2.10' | ||
| ;; | ||
| '2.14.0-cpu') | ||
| torch_install_version='2.14.0' | ||
| torchvision_install_version='0.29.0' | ||
| torch_test_version='2.14' | ||
| ;; | ||
| *) | ||
| echo "Unsupported torch_preset: $selected_preset" >&2 | ||
| exit 1 | ||
|
|
@@ -270,9 +288,37 @@ jobs: | |
| run: | | ||
| pip list | ||
|
|
||
| - name: Pre-build JIT ops | ||
| # Rank workers JIT-load ops concurrently on first use; torch's build | ||
| # lock has no stale handling, so a leftover baton from a killed build | ||
| # wedges every later load. Build the hot ops once up front so tests | ||
| # only hit the disk cache. | ||
| run: | | ||
| python -c " | ||
| from deepspeed.comm.torch import build_shm_op | ||
| build_shm_op() | ||
| import torch | ||
| from deepspeed.ops.adam import DeepSpeedCPUAdam | ||
| DeepSpeedCPUAdam([torch.nn.Parameter(torch.randn(8))]) | ||
| " | ||
|
|
||
| - name: Pre-download HF test fixtures | ||
| # Concurrent xdist workers downloading the same model race on the | ||
| # transformers cache lock and fail with PermissionError; populate the | ||
| # shared cache once so tests only hit it read-only. | ||
| run: | | ||
| HF_HUB_DISABLE_XET=1 HF_HOME=/tmp/hf_home hf download bert-base-uncased | ||
|
|
||
| - name: Unit tests | ||
| run: | | ||
| unset TORCH_CUDA_ARCH_LIST # only jit compile for current arch | ||
| cd tests | ||
| HF_HOME=/tmp/hf_home/ pytest $PYTEST_OPTS --forked -n 4 unit/ --torch_ver="$TORCH_TEST_VERSION" | ||
| HF_HOME=/tmp/hf_home/ pytest $PYTEST_OPTS --forked -m 'sequential' unit/ --torch_ver="$TORCH_TEST_VERSION" | ||
| # The suite is split so each half gets fresh xdist workers: multi-rank pool | ||
| # teardown eventually wedges a worker, and one 3400-test process never | ||
| # reaches its summary inside the 6h job limit. maxfail is raised so every | ||
| # failure is listed, and timeout caps each half. | ||
| overall=0 | ||
| HF_HOME=/tmp/hf_home timeout 210m pytest $PYTEST_OPTS --maxfail=100000 --forked -n 4 unit/ --ignore=unit/v1 --torch_ver="$TORCH_TEST_VERSION" || overall=$? | ||
| HF_HOME=/tmp/hf_home timeout 150m pytest $PYTEST_OPTS --maxfail=100000 --forked -n 4 unit/v1 --torch_ver="$TORCH_TEST_VERSION" || overall=$? | ||
| HF_HOME=/tmp/hf_home/ pytest $PYTEST_OPTS --forked -m 'sequential' unit/ --torch_ver="$TORCH_TEST_VERSION" || overall=$? | ||
| exit $overall | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
In the
cpu-torch-latestjob, advertising four devices admits every distributed test withworld_size <= 4, not only CPU-safe tests. For example,tests/unit/ulysses_alst/test_ulysses_sp_hf.py:240-263defines an unguarded two-rank test that creates tensors oncuda:<rank>; becauseCPU_Accelerator.is_available()returns true, the harness does not skip it, and the all-unit pytest invocation at workflow line 283 will fail on the CPU-only PyTorch installation. Scope this override to an explicitly CPU-compatible subset or add CPU capability checks before enabling the previously skipped tests.Useful? React with 👍 / 👎.