Skip to content

Skip fp16-config tests on accelerators without fp16 support - #8684

Merged
delock merged 1 commit into
deepspeedai:masterfrom
delock:pr-o-fp16-lottery-skips
Sep 28, 2026
Merged

delock merged 1 commit into
deepspeedai:masterfrom
delock:pr-o-fp16-lottery-skips

Conversation

@delock

@delock delock commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

Description

Every fp16-config test that reaches deepspeed.initialize crashes its sanity check (Type fp16 is not supported on your device.) on accelerators whose is_fp16_supported() is false. On CPU that maps to the AVX512-FP16 capability of the host, and GitHub's ubuntu-24.04 runners are hardware-heterogeneous, so these tests flip between failure and skip depending on which runner they land on. #8398 added the first skipifs; the multi-rank CPU run in #8381 flushed out six more files:

File Shape of the gap
checkpoint/test_universal_checkpoint.py fp16 parametrizations fail inside the baseline DistributedFixture's distributed run — pytest reports a setup ERROR for every dependent test (48 on the multi-rank run) instead of a skip
v1/zero/test_zero_coalesce_grad_reduction.py TestCoalesceFP16 forces an fp16 config
runtime/test_no_sync_ctxt.py dtype=float16 parametrizations of three methods; stages 2/3 additionally never reach their expected no_sync AssertionError on such hosts
checkpoint/test_moe_checkpoint.py whole class hardcodes fp16
runtime/zero/test_zero_offloadpp.py TestZeroPartialOffloadConfigSweep hardcodes fp16
checkpoint/test_pipeline.py fp16 enabled for zero_stage > 0; only that parametrization skips, zero_stage=0 keeps running

With this PR, every fp16-config test in the suite guards on is_fp16_supported() — the capability gap is a skip, not a failure, on any accelerator.

Validation (executed on real hardware)

  • 20-core x86_64 CPU without AVX512-FP16, gloo, 2-4 ranks (LOCAL_SIZE=2/4)
  • Before: 18 failures + 48 setup ERRORs across these files on the multi-rank CPU run
  • After: every affected parametrization skips; adjacent non-fp16 parametrizations keep passing (e.g. pipeline zero_stage=0 runs to completion)
  • Full-suite evidence in [DON'T MERGE] Run multi-rank CPU unit tests in CI via LOCAL_SIZE #8381: the multi-rank CPU run went 8 failures -> 0 with these guards

Sibling PRs from the same series: #8397, #8398, #8399, #8407, #8409, #8559, #8648.

Every fp16-config test that reaches deepspeed.initialize crashes its
sanity check ("Type fp16 is not supported on your device.") on
accelerators whose is_fp16_supported() is false — on CPU that maps to
the AVX512-FP16 capability of the host, and GitHub's ubuntu-24.04
runners are hardware-heterogeneous. deepspeedai#8398 added skipifs for the first
three files; the multi-rank CPU run in deepspeedai#8381 flushed out six more:

- checkpoint/test_universal_checkpoint.py: the fp16 parametrizations
  fail inside the baseline DistributedFixture's distributed run, which
  reports a setup ERROR for every dependent test instead of a skip.
- v1/zero/test_zero_coalesce_grad_reduction.py: TestCoalesceFP16
  forces an fp16 config.
- runtime/test_no_sync_ctxt.py: the dtype=float16 parametrizations of
  the three dtype-parametrized methods; stages 2/3 additionally never
  reach their expected no_sync AssertionError on such hosts.
- checkpoint/test_moe_checkpoint.py: the whole class hardcodes fp16.
- runtime/zero/test_zero_offloadpp.py: TestZeroPartialOffloadConfigSweep
  hardcodes fp16.
- checkpoint/test_pipeline.py: fp16 is enabled for zero_stage > 0; skip
  only that parametrization so zero_stage=0 keeps running.

With this, every fp16-config test in the suite guards on
is_fp16_supported(): the gap is a skip, not a failure, on any
accelerator. Verified on an fp16-incapable CPU (the affected
parametrizations skip; adjacent non-fp16 ones keep passing) and by the
multi-rank CPU run in deepspeedai#8381.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
def test_zero_stage(self, zero_stage, dtype):
# The fp16 parametrization crashes initialize's sanity check on accelerators
# without fp16 support (#8398's hardware lottery); skip instead of failing.
if dtype == torch.float16 and not get_accelerator().is_fp16_supported():

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we have a check for bf16 here as well?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the comments. Currently CPU with accelerated instruction set (AVX or AMX) has BF16 support. We can add bf16 skip check when we need to run validation on accelerators without bf16 support.

@delock
delock added this pull request to the merge queue Sep 28, 2026
Merged via the queue into deepspeedai:master with commit df00e0b Sep 28, 2026
13 of 15 checks passed
@delock
delock deleted the pr-o-fp16-lottery-skips branch September 28, 2026 07:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants