feat(dpmodel): graph-native se_atten attention (NeighborGraph PR-D) - #5715
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughDPA1 graph-native lowering now supports attention with static neighbor-pair enumeration, new segment reduction helpers, and expanded graph/export validation. Eligibility text and graph-path documentation were updated to match the new attention-capable graph lower. ChangesGraph-native attention support
Estimated code review effort: 4 (Complex) | ~60 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (4)
source/tests/common/dpmodel/test_segment_softmax.py (1)
55-65: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick winAdd a regression test for masked-entry-larger-than-max.
None of the mask tests here cover a masked entry whose value exceeds the unmasked max in the same segment — the scenario that triggers the NaN-propagation issue flagged in
segment.py. Once that's fixed, a test like the one below would guard the regression:def test_masked_entry_extreme_value_no_nan(self) -> None: logits = np.array([1.0, 1e30, 2.0]) # masked entry (idx 1) dwarfs the max ids = np.array([0, 0, 0], dtype=np.int64) mask = np.array([True, False, True]) w = segment_softmax(logits, ids, 1, mask=mask) assert not np.any(np.isnan(w)) assert w[1] == 0.0🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@source/tests/common/dpmodel/test_segment_softmax.py` around lines 55 - 65, Add a regression test in test_segment_softmax for the masked-entry-larger-than-max case that currently leads to NaN propagation in segment_softmax. Extend the existing mask coverage by creating a segment where the masked element has an extreme value above the unmasked max, then assert the result contains no NaNs, the masked position is exactly zero, and the unmasked weights still normalize correctly. Use the existing segment_softmax test pattern in test_masked_entries_zero to keep the new case consistent.deepmd/dpmodel/utils/neighbor_graph/pairs.py (1)
92-117: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value
dstvalues are unused in the shape-static path (only its shape matters).
_pairs_shape_staticderives query/key edges purely from index arithmetic assuming the center-major layout documented in the module docstring; the actualdstvalues are never consulted to validate that assumption. This matches the documented contract, but if a caller ever passes adst/static_nneicombination that doesn't match the assumed layout, this silently produces wrong pairs with no diagnostic. Consider a lightweight assertion (e.g.,e_tot % nn == 0) or a debug-mode check thatdstis actually constant within each block, to fail fast on a layout mismatch instead of silently mis-pairing.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@deepmd/dpmodel/utils/neighbor_graph/pairs.py` around lines 92 - 117, The shape-static path in `_pairs_shape_static` relies on center-major block layout but never validates that `dst` actually matches that assumption, so a mismatched `static_nnei`/layout can silently produce wrong pairs. Add a lightweight guard in `_pairs_shape_static` to fail fast on layout mismatches, such as verifying `e_tot % nn == 0` and/or checking that `dst` is constant within each `nn` block in a debug-friendly way. Keep the existing index-arithmetic logic for `query_edge`, `key_edge`, and `pair_mask`, but ensure the contract is enforced before returning.deepmd/dpmodel/descriptor/dpa1.py (2)
1671-1684: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win"Bit-exact" claim needs a caveat for the default (smooth + compact) configuration.
The docstring states this is a "Bit-exact analogue of
call" and the "Known limitations" section only liststebd_input_modeandexclude_types. But pertest_block_compact_graph_smooth_clean_divergenceintest_dpa1_graph_attention_parity.py, whenstatic_nnei is None(the default, compact/carry-all form) andsmooth=True(also the class default), the output deliberately diverges from dense (up to ~1e-4) by design — the carry-all graph drops phantom sel-padding softmax terms that dense keeps. A reader of this docstring/API surface would not learn about this without digging into the test suite. Sincesmooth_type_embeddingdefaults toTrueandstatic_nneidefaults toNone, the "bit-exact" claim is misleading for the descriptor's own default configuration.Suggest adding a short caveat to the "Known limitations" (or a new "Notes") section referencing this divergence, mirroring what's already documented in the test docstring.
📝 Suggested docstring addition
Notes ----- Known limitations: - ``tebd_input_mode == "concat"`` only (strip mode lands later); - ``exclude_types`` is not yet supported and raises (lands in a later PR). + - When ``attn_layer > 0``, ``smooth_type_embedding=True`` (the class + default) combined with the compact/carry-all form (``static_nnei=None``, + also the default) intentionally diverges from the dense reference + (up to ~1e-4): the carry-all graph has no sel-padding slots, so it + drops the phantom denominator terms the dense smooth branch keeps. + Bit-exact parity (1e-12) only holds on the shape-static form + (``static_nnei`` set, as used by the dense ``call`` adapter) or when + ``smooth_type_embedding=False``. """Also applies to: 1712-1717
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@deepmd/dpmodel/descriptor/dpa1.py` around lines 1671 - 1684, Update the call_graph docstring in dpa1.py to add a caveat that the “bit-exact” claim does not hold for the default smooth + compact/carry-all configuration: when static_nnei is None and smooth=True, the graph path can intentionally diverge slightly from dense because it omits phantom sel-padding softmax terms. Add this to the existing “Known limitations” or a new “Notes” section, and keep the wording consistent with the behavior exercised by test_block_compact_graph_smooth_clean_divergence and the related call_graph documentation block.
1856-1932: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winExtract the shared
attnw_shiftdefault.GatedAttentionLayer.callalso uses20.0, so pulling this into a shared constant would keep the dense and graph paths aligned if that default ever changes.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@deepmd/dpmodel/descriptor/dpa1.py` around lines 1856 - 1932, The hardcoded attention shift value is duplicated in _graph_attention_one_layer and GatedAttentionLayer.call, so pull the 20.0 default into a shared constant or class attribute used by both paths. Update the graph attention logic to reference that shared symbol so the dense and graph implementations stay aligned if the default changes.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@deepmd/dpmodel/utils/neighbor_graph/segment.py`:
- Around line 59-89: The masked path in segment_softmax is using raw data for
the exponent shift, which can turn masked large values into inf and then nan
after multiplying by the mask. Update segment_softmax to compute shifted from
data_for_max (the same masked-safe values used for seg_max), and keep the
existing empty/fully-masked guards so exp and denom stay finite. Check the
segment_max/segment_sum flow and the _graph_attention_one_layer caller to ensure
masked attention logits cannot leak into the denominator.
---
Nitpick comments:
In `@deepmd/dpmodel/descriptor/dpa1.py`:
- Around line 1671-1684: Update the call_graph docstring in dpa1.py to add a
caveat that the “bit-exact” claim does not hold for the default smooth +
compact/carry-all configuration: when static_nnei is None and smooth=True, the
graph path can intentionally diverge slightly from dense because it omits
phantom sel-padding softmax terms. Add this to the existing “Known limitations”
or a new “Notes” section, and keep the wording consistent with the behavior
exercised by test_block_compact_graph_smooth_clean_divergence and the related
call_graph documentation block.
- Around line 1856-1932: The hardcoded attention shift value is duplicated in
_graph_attention_one_layer and GatedAttentionLayer.call, so pull the 20.0
default into a shared constant or class attribute used by both paths. Update the
graph attention logic to reference that shared symbol so the dense and graph
implementations stay aligned if the default changes.
In `@deepmd/dpmodel/utils/neighbor_graph/pairs.py`:
- Around line 92-117: The shape-static path in `_pairs_shape_static` relies on
center-major block layout but never validates that `dst` actually matches that
assumption, so a mismatched `static_nnei`/layout can silently produce wrong
pairs. Add a lightweight guard in `_pairs_shape_static` to fail fast on layout
mismatches, such as verifying `e_tot % nn == 0` and/or checking that `dst` is
constant within each `nn` block in a debug-friendly way. Keep the existing
index-arithmetic logic for `query_edge`, `key_edge`, and `pair_mask`, but ensure
the contract is enforced before returning.
In `@source/tests/common/dpmodel/test_segment_softmax.py`:
- Around line 55-65: Add a regression test in test_segment_softmax for the
masked-entry-larger-than-max case that currently leads to NaN propagation in
segment_softmax. Extend the existing mask coverage by creating a segment where
the masked element has an extreme value above the unmasked max, then assert the
result contains no NaNs, the masked position is exactly zero, and the unmasked
weights still normalize correctly. Use the existing segment_softmax test pattern
in test_masked_entries_zero to keep the new case consistent.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: 86eede99-fa33-4044-b859-5fe1eb620896
📒 Files selected for processing (13)
deepmd/dpmodel/descriptor/dpa1.pydeepmd/dpmodel/utils/neighbor_graph/__init__.pydeepmd/dpmodel/utils/neighbor_graph/env.pydeepmd/dpmodel/utils/neighbor_graph/pairs.pydeepmd/dpmodel/utils/neighbor_graph/segment.pysource/tests/common/dpmodel/test_center_edge_pairs.pysource/tests/common/dpmodel/test_dpa1_call_graph_block.pysource/tests/common/dpmodel/test_dpa1_graph_attention_parity.pysource/tests/common/dpmodel/test_segment_softmax.pysource/tests/pt_expt/descriptor/test_dpa1.pysource/tests/pt_expt/model/test_dpa1_graph_lower.pysource/tests/pt_expt/model/test_linear_model.pysource/tests/pt_expt/utils/test_neighbor_list.py
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## master #5715 +/- ##
==========================================
- Coverage 81.29% 81.19% -0.10%
==========================================
Files 990 991 +1
Lines 111019 111182 +163
Branches 4235 4232 -3
==========================================
+ Hits 90252 90275 +23
- Misses 19243 19381 +138
- Partials 1524 1526 +2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
Pushed two additional commits that remove the "attention graph-form export deferred" limitation:
AOTI parity vs eager carry-all measured at ≤5e-18 across system sizes. Benchmark on a Tesla T4 (fp64, diamond C at experimental density, rcut 6 / sel 180, eager): the graph path is flat ~100 µs/atom (O(N)) and still runs at 4096 atoms where the dense path OOMs; with attention the graph is consistently faster at every size that fits. Known limitations: relies on torch unbacked-SymInt maturity (validated on 2.10; CPU AOTI); jax.jit of the compact path still needs a static realization (PR-F); C++ gtest of an attention graph |
There was a problem hiding this comment.
🧹 Nitpick comments (1)
source/tests/pt_expt/model/test_dpa1_graph_lower.py (1)
240-292: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick winAlso assert
atom_virialparity.
do_atomic_virial=Trueis passed on both sides but the resultingatom_virialtensor is never compared, leaving a gap in parity coverage specifically for the newattn_layer=2graph-attention path this test targets.✅ Proposed addition
torch.testing.assert_close( out["virial"], ref["energy_derv_c_redu"].reshape(out["virial"].shape), **tol ) + torch.testing.assert_close( + out["atom_virial"], + ref["energy_derv_c"].reshape(out["atom_virial"].shape), + **tol, + )🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@source/tests/pt_expt/model/test_dpa1_graph_lower.py` around lines 240 - 292, The symbolic-trace parity test in test_graph_lower_symbolic_trace already compares energy, force, and virial, but it omits the atom-level virial output even though do_atomic_virial=True is used. Update the assertions in test_graph_lower_symbolic_trace to also compare traced versus reference atom_virial from forward_lower_graph_exportable and forward_common_lower_graph, using the same tolerance and reshaping pattern as the other tensor checks if needed.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@source/tests/pt_expt/model/test_dpa1_graph_lower.py`:
- Around line 240-292: The symbolic-trace parity test in
test_graph_lower_symbolic_trace already compares energy, force, and virial, but
it omits the atom-level virial output even though do_atomic_virial=True is used.
Update the assertions in test_graph_lower_symbolic_trace to also compare traced
versus reference atom_virial from forward_lower_graph_exportable and
forward_common_lower_graph, using the same tolerance and reshaping pattern as
the other tensor checks if needed.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: 7850ae92-4cd9-427e-b8b6-a37842539f6d
📒 Files selected for processing (5)
deepmd/dpmodel/array_api.pydeepmd/dpmodel/utils/neighbor_graph/pairs.pydeepmd/pt_expt/entrypoints/main.pysource/tests/pt_expt/infer/test_graph_deepeval.pysource/tests/pt_expt/model/test_dpa1_graph_lower.py
✅ Files skipped from review due to trivial changes (1)
- deepmd/pt_expt/entrypoints/main.py
🚧 Files skipped from review as they are similar to previous changes (1)
- deepmd/dpmodel/utils/neighbor_graph/pairs.py
…sh eligibility wording Address OutisLi review on deepmodeling#5715: - Notes on DescrptDPA1.call_graph (+ pointers on uses_graph_lower and the freeze lower_kind docstring): for smooth_type_embedding=True the carry-all graph attention intentionally drops the dense layout's sel-padding terms from the softmax denominator - sel-independent semantics that differ from the legacy dense lower by up to ~1e-4; bit-tight dense parity holds for the non-smooth branch and for the static_nnei dense-adapter realization. - Update the stale 'dpa1 attn_layer == 0' eligibility wording (freeze() docstring, training-path predicate/comment, dpmodel/pt_expt graph-lower docstrings) to the actual contract: mixed-types dpa1/se_atten with concat type embedding and no exclude_types, attention layers included; call_common docs no longer imply unconditional dense parity.
iProzd
left a comment
There was a problem hiding this comment.
Two follow-ups now that we are shipping the default flip:
-
Since
smooth_type_embedding=True+attn_layer=2are the dpa1 constructor defaults, existing checkpoints trained under the dense semantics will shift up to ~1e-4 when evaluated on pt_expt after this PR. Could we record this as an explicit user-facing behavior change (changelog / release notes / migration note), including the escape hatch (neighbor_graph_method="legacy"restores the old numbers)? Users hitting reference-value regressions should be able to find the explanation without readingcall_graphdocstrings. -
The pt_expt-graph vs dense end-to-end divergence is currently invisible to the consistent suites (they exercise the bit-exact shape-static adapter). Could we add one end-to-end test pinning the expected divergence — nonzero and bounded by the documented ~1e-4 magnitude — so a future refactor cannot silently change the carry-all smooth semantics?
|
Both review-body items addressed in fc30ee4 (the inline torch<2.6 guard was addressed earlier in da46ed2/60c5e8377):
|
…ftmax Built on the existing xp_maximum_at (no new array_api helper needed). Part of NeighborGraph PR-D (graph-native attention).
Segment-based (global (E,E) boolean deliberately avoided): compact eager form for carry-all graphs + shape-static nonzero-free form for the center-major static layout (jit/export/make_fx traceable). Part of NeighborGraph PR-D; PR-E angles reuse (unordered, no-self).
…r > 0) DescrptBlockSeAtten.call_graph grows _graph_attention: the dense per-center (nnei, nnei) attention square becomes the edge-pair axis (center_edge_pairs, ordered + self-included), softmax over keys becomes segment_softmax grouped by the query edge. Op-for-op mirror of GatedAttentionLayer.call (head_dim QKV slicing, normalize q/k/v, temperature/scaling, smooth shift trick, post-softmax sw and dotr weighting, residual + LayerNorm per layer). - shape-static adapter path (static_nnei threaded from the dense call adapter): bit-exact vs the dense body, rtol 1e-12, full flag matrix (attn_layer 1/2 x dotr x smooth x normalize x temperature, binding and non-binding sel). - carry-all (compact) graphs: exact for non-smooth; for smooth the dense branch keeps sel-padding slots in the softmax denominator (dense output is sel-DEPENDENT, up to ~1e-4) — the carry-all form drops those phantom terms by design (user decision 2026-07-03), pinned by a clean-divergence test. - edge_env_mat(return_sw=True) exposes the per-edge switch (zeroed on padding) for the smooth branch. - uses_graph_lower: attention configs are now graph-eligible (concat tebd, no exclude_types still required).
…ial parity
- test_make_fx_graph_attn: graph forward + autograd.grad at attn_layer=2
traces under make_fx for BOTH smooth branches (the shape-static
center_edge_pairs form is nonzero-free) — required since pt_expt compiled
training routes eligible models through the graph lower.
- model-level graph-vs-legacy lower parity now parametrized over
attn_layer {0, 2} (energy/force/virial/atom_virial, 1e-12 CPU).
- eligibility pins: attention+concat is graph-eligible; se_atten_v2
(tebd_input_mode='strip') correctly stays dense (strip = later PR;
the plan's 'se_atten_v2 inherits for free' did not hold).
- linear-model weight tests: pin smooth_type_embedding=False — the standard (graph-routed, carry-all) and linear (graph-ineligible, dense) submodels otherwise differ by the accepted smooth-attention denominator divergence (~1e-6), which is a route artifact, not a weight-combination bug. - new binding-sel sanity: carry-all graph attention diverges from the sel-truncated dense path when sel binds (spec decision deepmodeling#17).
…rity) neighbor_list=None now takes the carry-all graph default for eligible attention models; explicit World-1 builders take the legacy dense route. With smooth attention the two routes differ by design (PR-D), so the route-equivalence tests pin smooth_type_embedding=False.
for more information, see https://pre-commit.ci
…SymInts The compact (carry-all) pair enumeration used nonzero + tensor-repeat with Python control flow on their data-dependent sizes, so the attention graph lower failed torch.export with GuardOnDataDependentSymNode. Register those sizes as unbacked SymInt sizes (new torch-free xp_hint_dynamic_size shim, no-op for numpy/jax), take the empty-input fast paths only on concrete int shapes, build iotas via cumsum(ones)-1 (the array_api_compat arange wrapper branches on the length in Python), and skip the policy-compression nonzero when no filter applies (include_self and ordered - the attention default). Eager numpy/torch results are unchanged.
…> 0)
With the compact pair enumeration unbacked-SymInt-traceable, the carry-all
attention graph lower now exports to a graph-form .pt2 unchanged in ABI
(same 5-tensor NeighborGraph schema, dynamic edge axis) and with carry-all
semantics preserved (no sel truncation, unlike the dense-adapter nlist-form
export). Update the stale freeze-gate message (attention is eligible), add a
symbolic-trace merge gate at attn_layer in {0,2}, parametrize the DeepEval
graph .pt2 fixture over attn_layer (both artifacts: dynamic sizes, PBC and
non-PBC, 1e-10 vs the sel-capped dense reference at non-binding sel), and
add a single-atom zero-real-edge runtime test (the R==0 extreme of the
unbacked sizes).
The dense call is wrapped in @cast_precision, but the graph route's only float input (edge_vec) lives inside the NeighborGraph dataclass where the decorator cannot see it, so non-global-precision models (e.g. float32) crashed with a double-vs-float matmul on the graph route while the dense route worked. Cast edge_vec down to the descriptor precision on entry and the outputs back to the caller's dtype on exit (differentiable, so the model-level force autograd is unaffected). Add an fp32 graph-vs-dense route parity test at attn_layer 0 and 2.
…nt-wide NaN A masked entry whose raw logit exceeds the unmasked per-segment max by more than the exp overflow threshold (~709 fp64 / ~88 fp32) overflowed exp() to inf, and the post-hoc inf * 0 mask multiply produced nan, which the denominator sum then spread across the entire segment. Shift data_for_max (masked entries already -inf, exp(-inf) == 0 exactly) instead of the raw data; the mask multiply stays as a defensive no-op. Regression test with a masked logit 1e5 above the unmasked max. Addresses CodeRabbit review.
…sh eligibility wording Address OutisLi review on deepmodeling#5715: - Notes on DescrptDPA1.call_graph (+ pointers on uses_graph_lower and the freeze lower_kind docstring): for smooth_type_embedding=True the carry-all graph attention intentionally drops the dense layout's sel-padding terms from the softmax denominator - sel-independent semantics that differ from the legacy dense lower by up to ~1e-4; bit-tight dense parity holds for the non-smooth branch and for the static_nnei dense-adapter realization. - Update the stale 'dpa1 attn_layer == 0' eligibility wording (freeze() docstring, training-path predicate/comment, dpmodel/pt_expt graph-lower docstrings) to the actual contract: mixed-types dpa1/se_atten with concat type embedding and no exclude_types, attention layers included; call_common docs no longer imply unconditional dense parity.
…he pt_expt graph default and legacy escape hatch
fc30ee4 to
6fc45bd
Compare
The carry-all graph default routes DPA1 through segment_sum -> torch.index_add,
which is bit-exact on CPU but non-deterministic (atomicAdd) on CUDA. Surfaced
only in the merge-queue CUDA run:
- test_graph_lower_symbolic_trace: trace on CPU (model.to('cpu')) mirroring the
real .pt2 export, so CUDA params don't meet CPU graph tensors (FakeTensor
device-propagation error on aten.index_select).
- test_{descriptor,fitting_ll}_deterministic_dpa1: device-conditional assert --
exact on CPU, 1e-10 on CUDA (1-2 ULP index_add atomics).
- test_finetune_change_type: prec 1e-10 on CPU, 1e-5 on CUDA (two remapped
models accumulate via atomicAdd in different orders, ~1e-7).
…ing#5717) ## Summary NeighborGraph PR-E: the optional 3-body **angle** extension of the edge-graph neighbor-list contract (design discussion wanghan-iapcm#4). An angle is a pair of edges sharing a center (`dst(edge_a) == dst(edge_b)`), stored as `angle_index (2, A)` into `[0, E)` + `angle_mask (A,)` — `edge_vec` stays the ONLY geometry leaf, so force/virial assembly is untouched (proven by an invariance test). > **Stacked on deepmodeling#5715 (PR-D)** — reuses its `center_edge_pairs`. Please merge deepmodeling#5715 first; this branch then rebases clean onto master. Only the last 10 commits belong to this PR. ### What's added (all in `deepmd/dpmodel/utils/neighbor_graph/`) - `pad_and_guard_angles` (graph.py) — angle-axis padder mirroring `pad_and_guard_edges` (dynamic guard append / static `angle_capacity` with overflow ValueError). - `angles.py` (new): - `build_angle_index(edge_index, edge_vec, edge_mask, n_total, a_rcut, *, ordered=False, include_self=False, layout=None)` — unordered, no-self pairs of edges sharing a center where BOTH edges are within `a_rcut`; built on PR-D's `center_edge_pairs`; `pair_mask` folded into `angle_mask` (never discarded). - `attach_angles(graph, a_rcut, ...)` — post-hoc: edge graph in, graph with angle fields out (`dataclasses.replace`); default builders keep angles `None`. - `angle_to_edge_sum` / `angle_to_node_sum` — segment-sum aggregation to the query edge / shared center. - `graph_angle_cos(angle_index, edge_vec, eps=1e-6)` — per-angle cos θ mirroring dpa3 repflows `cosine_ij` eps placement exactly (`+eps` in norm denominators, `*(1-eps)` on the product). - `angle_padding_fraction(graph)` — mask-derived padding-waste report for static capacities. ### Semantics / decisions - **Unordered, no-self by default**: dpa3's dense angle tensor is the redundant ordered `a_sel x a_sel` square including the `j==k` diagonal; the graph set keeps one entry per unordered `{j,k}` pair and moves the degenerate diagonal to the (a_rcut-filtered) edge channel. Dense parity is therefore asserted against the OFF-DIAGONAL `cosine_ij[j,k]` (j != k) at rtol/atol **1e-12** (same-math fp64), at non-binding `a_sel`. The ordered+self full square stays available via flags. - **`a_sel` = normalization-only** (carry-all within `a_rcut`), consistent with the edge-`sel` decision. - Second oracle: se_t dot-product convention cross-checked from coordinates in the `sw == 1` regime (rtol 1e-12). ### Tests `source/tests/common/dpmodel/test_angle_builder.py` (21) + `test_graph_angle_cos_parity.py` (6): brute-force triplet oracle (all flag combinations, multi-center, static layout, `node_capacity` branch), dpa3 dense-parity + no-self-angle assertion, se_t coordinate oracle, force/virial bit-exact invariance with/without angles, padding-fraction (incl. `total==0`), torch-namespace smoke tests for every new function. Full neighbor-graph suite: 54 passed. ### Known limitations - **Machinery + angle-channel math only** — no dpa3 graph descriptor here (dpa3 is message-passing; wiring = PR-G). se_t/se_t_tebd are not migrated (`mixed_types=False`), used as oracle only. - Angle enumeration is the compact **eager** form (`nonzero` in `center_edge_pairs`) even when a static `layout` is passed — `angle_capacity` fixes the output shape only; shape-static enumeration for export is deferred to PR-G. - Not bit-parity with dense dpa3 **by construction** (unordered/no-self reformulation; recoverable in PR-G via symmetric angle→edge 2x + edge-channel diagonal). - Aggregation helpers follow the mask-then-reduce convention: callers mask per-angle data by `angle_mask` before summing (padding angles point at edge 0). - numpy/torch validated; jax rides the array-API surface (no jax-specific test here); `A ~ sum(deg^2)` capacity overhead mitigated by `a_rcut < rcut` and reported by `angle_padding_fraction`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added angle-graph utilities to build and attach 3-body angle relationships on top of neighbor graphs. * Introduced angle padding/guarding for fixed-capacity layouts. * Added helpers to compute angle cosine values, reduce angle data back to edges/nodes, and report angle padding coverage. * Expanded publicly available exports for angle/edge-pair utilities and additional segment reductions (max/softmax). * **Tests** * Added extensive unit tests covering index building, attachment, masking/aggregation correctness, padding behavior, and cosine parity across NumPy/Torch. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Han Wang <wang_han@iapcm.ac.cn> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…a1 graph path supports exclude_types (decision #18) (deepmodeling#5733) ## Pair `exclude_types` as a canonical NeighborGraph transform > **Stacked on deepmodeling#5715 (NeighborGraph PR-D).** Only the commits after `6fc45bd26` belong to this PR; will rebase onto master once deepmodeling#5715 merges. Makes pair-type exclusion a single canonical transform applied once at the neighbor-list/graph build seam, and uses it to add descriptor-level `exclude_types` support to the dpa1 graph path (removing that eligibility gate), consistently across dpmodel, pt_expt, jax, and the C++ inference path. ### What changed - **Canonical graph transform** `apply_pair_exclusion(graph, atype, pair_excl, *, compact=False)` in `deepmd/dpmodel/utils/neighbor_graph/`: ANDs `PairExcludeMask.build_edge_exclude_mask` into `graph.edge_mask`. `compact=False` is mask-only (shape-static, export/AOTI-safe); `compact=True` drops masked edges (eager-only; raises on angle-carrying graphs). Idempotent. - **Atomic-model seam** (`base_atomic_model`): model-level `pair_exclude_types` is applied at BUILD time, so `forward_common_atomic{,_graph}` no longer re-apply it — they consume a pre-excluded nlist/graph (single-owner, **not** a backstop). To keep that from being fail-open, an eager (numpy) fail-safe assertion (`_assert_nlist_pair_excluded` / `_assert_graph_pair_excluded`) rejects a non-excluded input at the seam with a clear error; it is a no-op under `torch.export` / jax `jit` and in compiled production (the exported-`.pt2` / C++ paths are covered by ingestion-site tests). See the ingestion-path inventory below. - **dpa1 graph path supports descriptor-level `exclude_types`**: the `NotImplementedError` and the `uses_graph_lower()` exclude condition are removed; exclusion applied inside `DescrptBlockSeAtten.call_graph` before the segment sums. Graph-vs-dense parity at non-binding sel is exact (rtol=atol=1e-12, attn_layer 0 and 2, type_one_side both). - **Build-time exclusion (dispatcher)**: `build_neighbor_graph` and the pt_expt graph builders (dense/ase/vesin/nv) gain optional `pair_excl`/`compact` with default post-search application; `_call_common_graph` passes model-level excludes at build time; oracle set-equality tests per available builder. - **Dense-nlist port**: `apply_pair_exclusion_nlist(nlist, atype_ext, pair_excl)` extracted from the inline seam code; `build_neighbor_list` + Vesin/Nv/Default strategies gain `pair_excl`; `return_mode='edges'` + `pair_excl` fails fast. - **C++ twin**: `buildPairExcludeTable` / `applyPairExclusion` / `applyPairExclusionNlist` in `source/api_cc/include/commonPT.h`, mirroring the Python transforms (same arg order/variable names, cross-referenced docs); `pair_exclude_types` serialized into `.pt2` `metadata.json` and rebuilt once in `DeepPotPTExpt::init` (device-resident table, uploaded once — no per-step H2D). Exclusion is applied at the C++ ingestion seam — the **single owner** on every run path; the exported lower consumes a pre-excluded input and never re-applies it. New gtest (8 tests) vs Python DeepEval reference at 1e-10, plus multi-rank LAMMPS exclusion tests (below). - **Fix**: `apply_pair_exclusion` uses `logical_and` + bool cast (array_api_strict rejected `bool*bool`), caught by the jax/strict consistency rows now traversing the graph path. ### Ingestion-path inventory (exclusion coverage) Because the seam is now the single owner (fail-open if skipped), here is every entry point that builds a nlist/graph reaching the exclusion-owning lower, and how each applies exclusion: | Entry point | Applies exclusion at | Coverage | |---|---|---| | dpmodel `_call_common` (dense) | `build_neighbor_list(pair_excl=)` | eager guard + consistency rows | | dpmodel `_call_common_graph` | `build_neighbor_graph(pair_excl=)` | eager guard + graph/dense parity | | dpmodel descriptor-level `exclude_types` | `apply_pair_exclusion` in `call_graph` | dpa1 graph parity | | pt_expt DeepEval — nlist | `apply_pair_exclusion_nlist` / vesin `pair_excl` | parity vs dpmodel | | pt_expt DeepEval — graph | `_build_eval_graph(pair_excl=)` (dense/ase/vesin/nv) | parity vs dpmodel | | pt_expt/jax/pd training | `build_neighbor_graph(pair_excl=)` in stat/forward | training e2e | | input statistics | `build_neighbor_list(pair_excl=)` in `EnvMatStatSe.iter` | stat tests (hash key = follow-up, pre-existing) | | **jax-md `call_lower`** | `apply_pair_exclusion_nlist` at the seam (**fixed here**) | `test_dense_neighbor_applies_model_pair_exclusion` | | C++ SP dense | `applyPairExclusionNlist` | gtest | | C++ MP dense (with-comm) | `applyPairExclusionNlist` (**fixed here**) | DPA3 MP≡SP + active-vs-baseline | | C++ SP graph | `applyPairExclusion` | gtest | | C++ MP graph (non-MP, extended) | `applyPairExclusion` | dpa1 MP≡SP + active-vs-baseline | | C++ MP graph (message-passing) | — | **fail-fast** (PR-G) | | C++ / DeepEval edge lower (SeZM/DPA4) | baked into the exported graph by the **pt backend** (unchanged) | out of scope — pt-backend export | Guard: `forward_common_atomic{,_graph}` additionally carry an eager fail-safe assertion (numpy only) that rejects a non-excluded input, so any *future* dpmodel/jax ingestion miss fails loudly instead of silently including excluded pairs. ### Known limitations - nv builder's `pair_excl` path has no local oracle test (CUDA-only); to be validated on a GPU box. - Input statistics remain on the dense path (graph-native stats is a separate follow-on); the stat-cache **hash key** does not yet include `pair_exclude_types` — pre-existing (predates this PR) and tracked as a separate fix. - smooth_type_embedding + exclude parity untestable at 1e-12 (pre-existing dense sel-padding divergence, deepmodeling#5715). - `build_edge_exclude_mask` still returns int32 (bool cast at call sites; follow-up). ### Spin routing Spin models auto-inject `exclude_types` (virtual/placeholder types) into their backbone descriptor; before this PR that condition *accidentally* kept spin on the dense path. With exclude_types now graph-eligible, spin backbones flipped onto the carry-all graph route, which (a) diverges from the sel-capped reference on sel-binding spin systems and (b) trips a torch-inductor scatter codegen assertion during spin `.pt2` export. Fixed explicitly: `DescrptDPA1.disable_graph_lower()` (not serialized; re-derived structurally) is set in `SpinModel.__init__` — the single choke point covering `get_spin_model`, `SpinModel.deserialize`, and the pt_expt spin classes — plus a belt-and-braces `neighbor_graph_method="legacy"` at `SpinModel.call_common`. Regression tests pin the routing and its serialize→deserialize survival; the full spin export suite (23) and spin checkpoint-interop suite (12) are green. ### Verification Full pt_expt suite: 1196 passed / 39 skipped / 3 failed — the 3 failures (`test_dpa4_freeze_to_pt2`, `test_dpa4_deep_eval_*`) are byte-identical on the base commit (pre-existing torch-inductor dpa4 export issue on this box, unrelated). dpmodel exclusion suites 69 passed; consistency dpa1 99 passed/63 skipped (incl. jax + array_api_strict exclude rows); C++ `Dpa1PairExcl` gtest 8/8. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added model-level pair-type exclusion across neighbor lists and neighbor graphs, including propagation into exported `.pt2` metadata and enforcement at inference ingestion. * Expanded DPA1 graph-native attention support (including higher attention layers) and introduced stable segmented reductions for graph attention. * Added export-time guards to block unsupported graph tracing on older torch versions. * **Bug Fixes** * Improved dense vs graph consistency for exclusion/masking behavior. * Preserved spin model legacy neighbor routing for sel-binding cases. * **Documentation** * Refreshed graph-export eligibility and backend behavior notes for attention/smooth differences. * **Tests** * Added/expanded parity, exclusion, compaction, and tracing regression coverage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Han Wang <wang_han@iapcm.ac.cn> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
…model on the NeighborGraph lower (PR-G-dpa2) (deepmodeling#5779) DPA-2 becomes graph-native full-stack — the first **message-passing** model on the NeighborGraph lower (PR-G-dpa2 of the [NeighborGraph series](wanghan-iapcm#4); follows dpa1's route from deepmodeling#5581/deepmodeling#5583/deepmodeling#5604/deepmodeling#5714/deepmodeling#5715/deepmodeling#5717/deepmodeling#5733). The dense path coexists untouched; the graph route is opt-in for inference (`--lower-kind graph`) and the default for pt_expt training/eager on graph-eligible configs, exactly like dpa1. ## What's new **dpmodel (backend-agnostic math)** - `DescrptBlockRepformers.call_graph`: every repformer op ported to the flat edge list — conv/grrg/drrd/g1g1/`LocalAtten` via `segment_*` over `dst`; the dense `nnei×nnei` gated attention (`Atten2Map`/`Atten2MultiHeadApply`/`Atten2EquiVarApply`) via `center_edge_pairs(ordered, include_self)` + per-head `segment_softmax` (extended to trailing feature dims). - `DescrptDPA2.call_graph` + `uses_graph_lower` + dense-call adapter. The multi-level nlist (`build_multiple_neighbor_list`) becomes per-block edge masks; in the shape-static adapter layout the per-block **slice** replicates the dense `nlist[:, :, :ns]` truncation, making the adapter **bit-exact vs the dense path at ANY sel** (verified to 5e-16 at deliberately binding sel with attention enabled) — this is what keeps the cross-backend consistency suites green after the routing flip. Ineligible configs (`use_three_body`, compression, spin) keep the legacy dense route. - Owned-node energy mask: `fit_output_to_model_output_graph` consumes `n_local`, excluding halo rows from the differentiated energy (prevents cross-rank double counting; force stays full-N, halo partials reverse-commed by LAMMPS). **pt_expt / export** - Routing, autograd force/virial, compiled training, and freeze eligibility are inherited generically from the dpa1 machinery. - Per-layer MPI halo refresh on the graph path: `_exchange_ghosts_graph` — identity on ghost-free single-rank graphs (`src` IS the owner), `deepmd_export::border_op` overwrite of halo rows on extended-region multi-rank graphs. - Message-passing graph `.pt2` archives embed a **with-comm AOTInductor artifact** (`model/extra/forward_lower_with_comm.pt2`), traced with the 8-tensor comm ABI shared with the dense flow. **C++ / LAMMPS** - `DeepPotPTExpt::run_model_graph_with_comm` + dispatch replacing the PR-G fail-fast: MP message-passing graph models run multi-rank on the extended-region graph + with-comm artifact; non-MP (dpa1) multi-rank unchanged; old graph archives without the artifact get a clear re-freeze error. - Fixtures (`gen_dpa2.py` section B), `dpa2_graph_ptexpt` universal gtest row, and `test_lammps_dpa2_graph_pt2.py` (single-rank vs reference, per-atom virial, `mpirun -n 2` ≡ `-n 1`, graph-vs-nlist cross-artifact, bounded empty-rank). ## Validation - GPU (Tesla T4, CUDA): dpa2-graph LAMMPS **5/5** incl. the MP==SP gate; dpa1-graph LAMMPS 6/6; dense dpa2 MPI 1/1; every `dpa2_graph` C++ gtest variant passed; python GPU suites green. - CPU: full dpmodel+pt_expt sweep 2112 passed; dpa2 consistency suite (pt/dp/jax/array-api-strict) green via the bit-exact adapter. - Real bugs found & fixed during GPU validation: CUDA device placement of the `nlocal`/`nghost` comm scalars (the graph route consumes them in on-device owned-mask index math); `_trace_and_compile_graph` hardcoded the global device for trace samples (latent since deepmodeling#5604). ## Known limitations / deliberate divergences (documented in `doc/model/dpa2.md` + code Notes) - Carry-all graph attention is sel-independent by design: at binding sel — and for smooth attention generally — the graph route diverges from dense (dpa1 precedent, `KNOWN_GRAPH_DENSE_DIVERGENT`). - Per-atom virial attribution differs elementwise from the dense decomposition for message-passing models (full-to-src vs autograd spread); per-frame totals agree (cross-checked at 1e-8 in fixture generation). - A truly empty MPI rank (zero owned+ghost atoms) fails loudly rather than running; the thrown error cannot propagate through peers blocked in `border_op` collectives, so the job stalls until MPI timeout — phantom-node support is a follow-up. The LAMMPS test bounds this with a hard timeout and asserts it never silently succeeds. - Pre-existing dense bug surfaced (NOT introduced here): `se_atten` at `attn_layer=0` leaks a deterministic `-davg/dstd` padding residual when `set_davg_zero=False` and `exclude_types == []` (`PairExcludeMask` all-ones short-circuit is the only padding mask on that path). The graph path masks padding correctly, so graph and dense deliberately differ in that regime (pinned by test + docstring). A dense-side fix will be proposed separately. ## Follow-ups (separate issues/PRs) - Dense `se_atten` padding-residual fix (above). - Wire the `BUILD_PT_EXPT` C++ gtest suite into CI (pre-existing gap — the universal gtest rows, including the new one, never run in CI today). - `TestDeepPotPTExptWithCommLoadFailure.multi_rank_compute_throws` is vacuous (sets `nswap` but not `nprocs`; pre-existing since deepmodeling#5450). - Empty-rank phantom-node support / clean collective abort; nested with-comm artifact schema assertion in the export test. PR-G-dpa3 (repflows + angle channel + `charge_spin`, reusing this PR's MP machinery) comes next. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Enabled graph-native DPA2 descriptor/repformer execution for eligible energy models, including multi-rank “with-comm” ghost exchange inside graph artifacts. * Added `comm_dict` plumbing and persistent graph-lower enable/disable controls across save/restart. * Improved multi-rank owned/halo reduction using `n_local`, including flat node-axis `aparam` handling for graph ABI and exports. * **Bug Fixes** * Refined multi-rank reduction correctness by excluding ghost contributions from differentiated reductions. * Improved graph attention/softmax stability with phantom-aware behavior for smoother continuity. * **Documentation** * Added pt_expt graph-native route documentation, eligibility rules, and graph-freeze with-comm requirements. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Han Wang <wang_han@iapcm.ac.cn> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: OutisLi <137472077+OutisLi@users.noreply.github.com>
Implements NeighborGraph PR-D: the graph path now supports
attn_layer > 0for dpa1/se_atten, removing the attn_layer=0-only restriction shipped in #5583.What
segment_max+ numerically-stable, mask-awaresegment_softmax(deepmd/dpmodel/utils/neighbor_graph/segment.py), built on the existingxp_maximum_at.center_edge_pairs(neighbor_graph/pairs.py): pairs of edges sharing a center — the edge-pair axis shared with the upcoming angle machinery (PR-E). Segment-based enumeration (a global(E,E)boolean is deliberately avoided:O(N²·nnei²)memory). Two forms: compact eager (dynamicP, carry-all graphs) and shape-static (P = n_center·nnei², pure arange/reshape arithmetic, nononzero) for the center-major static layout — this keeps the traced/compiled/export path traceable.DescrptBlockSeAtten._graph_attention: op-for-op ragged mirror ofGatedAttentionLayer/NeighborGatedAttention— per-centerq@kᵀbecomes per-pairq_m·k_n, softmax over keys becomessegment_softmaxgrouped by the query edge; head_dim QKV slicing, q/k/v normalize, temperature/scaling, smooth shift trick, post-softmaxswanddotrweighting, residual + LayerNorm per layer.edge_env_mat(return_sw=True)exposes the per-edge switch (zeroed on padding) for the smooth branch.uses_graph_lowerwidened: attention configs (concat tebd, no exclude_types) are now graph-eligible — pt_expt eager/compiled/exported paths route them through the graph lower by default.Numerical semantics (reviewed decision)
calladapter,from_dense_quartet(compact=False)+static_nnei): bit-exact vs the dense body, rtol 1e-12, full flag matrix (attn_layer 1/2 × dotr × smooth × normalize × temperature, binding AND non-binding sel).smooth_type_embedding=True, the dense branch keeps sel-padding slots in the attention softmax denominator (weightexp(-attnw_shift)), which makes the dense output depend on sel itself (measured up to ~1e-4 with an identical physical neighbor set). The carry-all form drops those phantom terms by design — the sel-independent math. Pinned by a clean-divergence test; route-equivalence fixtures pinsmooth_type_embedding=False.tebd_input_mode="strip") remains graph-ineligible (strip mode is a later PR) — pinned by test.Testing
test_make_fx_graph_attn(graph forward + autograd at attn_layer=2 traces under make_fx, both smooth branches — required since compiled training uses the graph lower); model-level graph-vs-legacy force/virial/atom-virial parity parametrized over attn_layer {0,2}.Known limitations
neighbor_graph_method="legacy"/ explicit World-1 builders.num_heads == 1assumed (dpa1 never exposes num_heads); fail-fast otherwise.center_edge_pairsis eager-only (nonzero); traced paths use the shape-static form.Summary by CodeRabbit