Loosen the late-iteration gradient tolerance in the checkpointing test - #8648
Merged
Merged
Conversation
The multiple-backward checkpointing test compares DDP and engine gradients over three optimizer iterations at dtype-default tolerances. The engine keeps fp32 accounting while the DDP reference steps the bf16 weights directly, so the two paths diverge by roughly one bf16 ulp of weight per step even from exactly equal gradients; by iteration 2 that crosses a reduction/relu rounding boundary and produces ~8e-3 absolute gradient differences on O(1) gradients. Whether the difference survives bf16 quantization depends on the CPU bf16 kernel's rounding, so the test flips between pass and fail across torch versions (fails on 2.10/2.11, passes on 2.13) without any behavioral change. Give iterations >= 2 an absolute tolerance floor (atol=1e-2 alongside the bf16-default rtol): the test already documents that late iterations drift, and the floor absorbs quantization-boundary chaos while still catching real gradient errors. Verified on both kernel families: 6/6 fail -> 6/6 pass on torch 2.11, unchanged 6/6 pass on torch 2.13. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
pengdurice
approved these changes
Sep 24, 2026
Merged
via the queue into
deepspeedai:master
with commit Sep 24, 2026
a4490b2
13 of 15 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
TestZeroUserBackwardWithCheckpointing::test_checkpointed_multiple_backwardcompares DDP and engine gradients over three optimizer iterations at dtype-default tolerances, and flips between pass and fail across torch versions on cpu (fails on 2.10/2.11 kernels, passes on 2.13) without any behavioral change.A minimal two-rank repro isolates the mechanism:
The engine keeps fp32 accounting while the DDP reference steps the bf16 weights directly, so the two mathematically-equivalent paths diverge by ~one bf16 ulp of weight per optimizer step even from bit-identical gradients; by iteration 2 that crosses a reduction/relu rounding boundary. Whether the difference survives bf16 quantization depends on the CPU bf16 kernel rounding — hence the version dependence.
This PR gives iterations >= 2 an absolute tolerance floor (
atol=1e-2alongside the bf16-defaultrtol=1.6e-2). The test already documents "small differences at later iterations are expected due to bfloat16 precision"; the floor absorbs quantization-boundary chaos while still catching real gradient errors.Validation (executed on real hardware)
LOCAL_SIZE=2)Exposed by the
LOCAL_SIZE=4multi-rank CPU run in #8381.