Skip to content

Loosen the late-iteration gradient tolerance in the checkpointing test - #8648

Merged
pengdurice merged 1 commit into
deepspeedai:masterfrom
delock:pr-l-ub-late-iter-tol
Sep 24, 2026
Merged

pengdurice merged 1 commit into
deepspeedai:masterfrom
delock:pr-l-ub-late-iter-tol

Conversation

@delock

@delock delock commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator

Description

TestZeroUserBackwardWithCheckpointing::test_checkpointed_multiple_backward compares DDP and engine gradients over three optimizer iterations at dtype-default tolerances, and flips between pass and fail across torch versions on cpu (fails on 2.10/2.11 kernels, passes on 2.13) without any behavioral change.

A minimal two-rank repro isolates the mechanism:

                   torch 2.13          torch 2.11 (= CI family)
iter 0 grads       exact match         exact match
post-step          —                   linear2.weight differs by 1.7e-6   <- the seed
iter 1 grads       exact match         exact match (quantization absorbs it)
iter 2 grads       exact match         7.8e-3                              <- exceeds rtol
final weights      1 bf16 ulp          1 bf16 ulp

The engine keeps fp32 accounting while the DDP reference steps the bf16 weights directly, so the two mathematically-equivalent paths diverge by ~one bf16 ulp of weight per optimizer step even from bit-identical gradients; by iteration 2 that crosses a reduction/relu rounding boundary. Whether the difference survives bf16 quantization depends on the CPU bf16 kernel rounding — hence the version dependence.

This PR gives iterations >= 2 an absolute tolerance floor (atol=1e-2 alongside the bf16-default rtol=1.6e-2). The test already documents "small differences at later iterations are expected due to bfloat16 precision"; the floor absorbs quantization-boundary chaos while still catching real gradient errors.

Validation (executed on real hardware)

  • 20-core x86_64 CPU, gloo, 2 ranks (LOCAL_SIZE=2)
  • torch 2.11 (the failing kernel family): 6/6 parametrizations fail -> pass
  • torch 2.13 (the passing family): 6/6 pass before and after (no regression)
  • pre-commit passes on the changed file

Exposed by the LOCAL_SIZE=4 multi-rank CPU run in #8381.

The multiple-backward checkpointing test compares DDP and engine
gradients over three optimizer iterations at dtype-default tolerances.
The engine keeps fp32 accounting while the DDP reference steps the bf16
weights directly, so the two paths diverge by roughly one bf16 ulp of
weight per step even from exactly equal gradients; by iteration 2 that
crosses a reduction/relu rounding boundary and produces ~8e-3 absolute
gradient differences on O(1) gradients. Whether the difference survives
bf16 quantization depends on the CPU bf16 kernel's rounding, so the
test flips between pass and fail across torch versions (fails on
2.10/2.11, passes on 2.13) without any behavioral change.

Give iterations >= 2 an absolute tolerance floor (atol=1e-2 alongside
the bf16-default rtol): the test already documents that late iterations
drift, and the floor absorbs quantization-boundary chaos while still
catching real gradient errors. Verified on both kernel families: 6/6
fail -> 6/6 pass on torch 2.11, unchanged 6/6 pass on torch 2.13.

Signed-off-by: Guokai Ma <guokai.ma@intel.com>
@pengdurice
pengdurice added this pull request to the merge queue Sep 24, 2026
Merged via the queue into deepspeedai:master with commit a4490b2 Sep 24, 2026
13 of 15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants