Skip to content

feat(xla): define operator-level MLX/IREE numeric contracts #932

Description

@inureyes

Parent and affected work

Problem

The remaining OpenXLA multimodal family ports now fail at a shared boundary rather than at missing model topology: MLX CUDA and StableHLO/IREE can execute mathematically equivalent operators with different operand materialization, accumulator precision, reassociation, reduction trees, fused-kernel ordering, or backend-selected convolution plans. The current family implementations encode some of these choices locally, so each full-model oracle run is rediscovering the same backend contract one architecture at a time.

This is not a request to relax tolerances. The existing family gates must remain unchanged, and final greedy output remains token-exact. If an MLX kernel has no deterministic, representable execution contract, the affected XLA capability must remain fail-closed until a reproducible contract is implemented or the reference path itself is made deterministic.

The reference baseline also changed after the current seven branches were cut: #923 fixed a shared-memory race in MLX CUDA qmm_sm80 and changed the kernel artifact identity. Any Q4 evidence collected before f7a877e7 must therefore be treated as historical until it is reproduced on the deterministic kernel.

Current evidence

Family Draft PR First unresolved boundary
Molmo2 #916 at 684c6ee1 Current-head actual local-task gate reaches projector.output_all; first failure is flat index 358403 with max_abs=0.1640625, rms=0.0026185848. The immediately preceding projector.product passes (max_abs=0.037597656, rms=0.0002196383), isolating final dense-projection accumulation/association.
Gemma3 VLM #917 at 9ab13992 Current-head actual run was bounded before model start by a one-time pinned-MLX cache rebuild. The last valid actual evidence remains siglip.hidden.block1.output at 002bb77e: flat index 658517, 8 failures, max_abs=0.6672516; current head passes the exact example check and focused Gemma3 tests but makes no new parity claim.
Molmo #918 at a4047b22 Exact assets, hashes, compiler, and target were verified, but the release standalone runner exceeded the bounded five-minute GNU final-link window and produced no executable. No current-head checkpoint comparison ran; intermediate/KV-logit/token/lifecycle gates remain unqualified.
Youtu-VL #919 at 20c2bffd Pinned HF and local-task IREE actuals agree through preprocessing, patch projection, and row-wise normalization. First unresolved stage is layer.7.full, flat index 310: actual/reference 0.3125/0.28515625, max_abs=0.203125, 2,999 unchanged-threshold mismatches.
Qwen2.5-VL #920 at 02d41a18 Rebased on the shared foundation; focused Qwen2.5 tests pass. This host exposes no NVIDIA device, so post-#923 production MLX CUDA Q4 qmv/qmm evidence has not been reproduced and the family remains fail-closed.
Qwen3-VL #921 at a4b641e6 Rebased on the shared foundation; Qwen3 and DeepStack focused tests pass. This host exposes no NVIDIA device, so the stale pre-#923 Q4 comparison is not promoted and current production Q4 parity remains unqualified.
Gemma3n audio #922 at 0a5aa90d Shared CUDA-prefix contract is adopted and current diagnostics check plus 17 audio unit tests pass. The production first SSCP convolution still has 6 BF16-ULP residual failures with max_abs=0.25; no reproducible cuDNN-plan-equivalent IREE schedule is represented, so the capability remains disabled.

Goal

Define and enforce one versioned, operator-level numeric contract for MLX-to-StableHLO/IREE qualification, then use bounded micro-oracles to resolve shared drift before repeating full-model gates.

Required implementation

  • Add a versioned numeric-contract descriptor to the XLA auxiliary artifact identity. It must cover source checkpoint dtype, input/output materialization dtype, contraction accumulator/result dtype, explicit rounding boundaries, reduction or scan association policy where specified, quantized dequantization order, and any backend kernel or algorithm identity required for reproducibility.
  • Represent the descriptor in shared XLA configuration and validation code rather than as family-specific strings. Artifact loading must reject a missing or mismatched numeric-contract fingerprint before native invocation.
  • Add explicit emitter helpers for the contract-sensitive operations already used by the affected families: affine Q4 dequantization, dense/Q4 matrix multiplication, LayerNorm/RMSNorm reduction and restoration, residual addition, SiLU/GELU materialization, softmax/attention scaling, prefix reductions/scans, and convolution result typing.
  • Keep the helpers semantically narrow. A helper may specify operand/result dtype and an explicit association schedule, but it must not silently apply graph-wide F16/BF16 casts or rewrite unrelated operations.
  • Build deterministic micro-oracles that run identical operands through the MLX reference operation and the emitted IREE operation, record the first divergent output, and include exact backend/build/kernel identity. At minimum cover Q4 qmv/qmm on the post-fix(cuda): drain cp.async before qmm_sm80 epilogue smem store (#910) #923 kernel, dense matmul, normalization, activation plus projection, residual add, attention/softmax, cumulative normalization, and the Gemma3n SSCP convolution shape.
  • Separate three reference claims in reports: exact production-MLX equivalence, exact canonical decomposed-operation equivalence, and mathematical closeness. Only the first claim can satisfy a family production gate unless that family issue explicitly defines the canonical decomposed operation as its production reference.
  • For unspecified backend algorithms such as cuDNN convolution plan selection, either import and pin a reproducible engine configuration, implement a proven matching schedule, or keep the affected capability disabled. Do not substitute a tolerance increase or a mathematical proxy for a missing execution contract.
  • Rebase the seven affected family branches on the deterministic f7a877e7 baseline before using their evidence. Re-run only the bounded operator probes first; repeat a heavyweight full-model oracle only after the first divergent operator boundary passes.
  • Replace ad hoc family-local precision workarounds with the shared helpers when doing so preserves the family graph contract. Keep family-specific schedules only when the checkpoint architecture genuinely specifies a distinct operation.
  • Update diagnostics so every actual report includes the numeric-contract version, MLX revision and patch identity, IREE/compiler version, target/device, selected kernel or algorithm identity when observable, and the exact first divergent tensor coordinate.
  • Keep CLI/server capability predicates fail-closed while a required family contract or oracle is unresolved. Merging shared infrastructure must not advertise an unqualified family.

Merge strategy

Implement the shared descriptor, fingerprint validation, explicit helpers, and micro-oracle harness in a narrow foundation PR first. After that PR lands, rebase the seven draft family PRs sequentially, resolve their shared-file overlap against the foundation, and run the affected bounded probes. A family PR becomes ready only after its original intermediate and token-exact gates pass; diagnostic or fail-closed infrastructure may be split into a separate PR when it is independently complete and does not expose the unfinished family.

Non-goals

  • Raising or removing the existing family-specific absolute, RMS, ULP, or token-exact acceptance thresholds.
  • Treating generic StableHLO mathematical equivalence as proof of production MLX equivalence.
  • Encoding one GPU model's undocumented kernel schedule as a repository-wide default.
  • Re-running all heavyweight model gates before a bounded operator probe shows that the first divergent boundary is fixed.
  • Advertising Molmo, Molmo2, Youtu-VL, Qwen2.5-VL, Qwen3-VL, Gemma3 VLM, or Gemma3n audio before its existing issue acceptance criteria pass.

Validation

  • Add unit tests for numeric-contract serialization, fingerprint stability, mismatch rejection, and every explicit materialization boundary.
  • Add mutation-sensitive emitter tests that fail when a specified cast, accumulator/result dtype, operation order, or association schedule is removed or reassociated.
  • Run deterministic MLX/IREE micro-oracles repeatedly on the production-relevant CUDA target and at least one local CPU target where the operation is supported.
  • Reproduce the Qwen2.5-VL and Qwen3-VL Q4 probes after fix(cuda): drain cp.async before qmm_sm80 epilogue smem store (#910) #923 and distinguish a corrected stale reference from a remaining representable contract mismatch.
  • Re-run the first-divergence probes for Gemma3 VLM, Molmo2, and Gemma3n audio before any full-model rerun.
  • Run the existing mlxcel-xla structural/golden tests, native IREE compile/load/invoke checks, text-only XLA architecture oracle, and continuous-batch regression suites.
  • Record commands, target/device, compile time, transfers, memory, and first-divergence output without committing generated model artifacts or incidental CSV files.

Acceptance criteria

  • A versioned operator-level numeric contract is part of every affected auxiliary artifact fingerprint and is rejected on mismatch before invocation.
  • Contract-sensitive emitter operations use shared explicit helpers with mutation-sensitive tests for dtype, rounding, operation order, and required association schedules.
  • Deterministic MLX/IREE micro-oracles cover the listed operation classes and report reproducible backend identity plus the first divergent coordinate.
  • Q4 evidence for feat: add Qwen2.5-VL OpenXLA vision path #920 and feat(xla): add Qwen3-VL DeepStack vision path #921 is refreshed on or after fix(cuda): drain cp.async before qmm_sm80 epilogue smem store (#910) #923, with stale pre-fix evidence clearly separated.
  • Each affected family either passes its unchanged original intermediate/token-exact gates or remains fail-closed with a precise unrepresentable backend contract recorded.
  • No tolerance is relaxed and no mathematical proxy is reported as production-backend parity.
  • Existing text-only XLA and continuous-batch behavior remains unchanged and qualified.

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:architectureArchitecture and code structure changesarea:coremlxcel-core: MLX FFI, primitives, KV cache, layersarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)priority:highHigh prioritystatus:blockedBlocked by dependencies or other issuestype:enhancementNew features, capabilities, or significant additions

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions