Skip to content

[Fix] Match GLM5.3 mHC pre-mixing precision with HF - #2154

Closed
bingshao333 wants to merge 1 commit into
InternLM:feat/glm53flash-f4-mhcfrom
bingshao333:fix/glm53flash-f4-mhc-correctness-f41
Closed

bingshao333 wants to merge 1 commit into
InternLM:feat/glm53flash-f4-mhcfrom
bingshao333:fix/glm53flash-f4-mhc-correctness-f41

Conversation

@bingshao333

Copy link
Copy Markdown

Summary

The BF16 hc_pre path casts FP32-normalized residual streams back to BF16 and performs the mixing projection in BF16. This introduces rounding before the Sinkhorn split, so the collapsed streams and input/parameter gradients can diverge from Hugging Face's FP32 reference.

Keep both the unweighted RMS normalization and mixing projection in FP32, then cast the collapsed stream back to the activation dtype. This is an F4 follow-up to #2107, targeting feat/glm53flash-f4-mhc for merge before the upper stack layers are restacked.

Changes

  • Preserve FP32 normalized activations and project with hc_fn.float() inside hc_pre.
  • Add HF forward/backward regression coverage for BF16 residual streams with FP32 and BF16 projection weights, checking the input and all trainable HC parameters. Tighten the existing BF16 pre-mix assertions.
  • Port the two F4 files from JT-Ushio's a8d798a, preserving Tao Ji as the commit author and adding type annotations to the new test.

Validation

  • pytest tests/model/test_glm53_mhc.py -q: 10 passed, 2 skipped. Both skipped cases require CUDA.
  • On the unchanged F4 base (6841448), both newly added CPU precision regressions fail; both pass with this fix.
  • Ruff lint and format checks pass for both changed files; git diff --check passes.
  • Full pre-commit: mypy passes for 385 source files; Pydantic, YAML/conflict/line-ending checks, codespell, and pyupgrade pass. The aggregate command fails on docformatter/Ruff formatting in six existing autotest/ files; the identical formatting changes were reproduced on the unchanged F4 base.

Local validation uses PyTorch 2.9.1 and Transformers 5.17.0 on CPU. This restores reference arithmetic in a performance-sensitive projection; GPU throughput and complete-model training were not measured in this environment.

Select the F4 mHC primitive and regression-test changes from
JT-Ushio/xtuner-muonsplit-verify commit
a8d798a.

Preserve the original author and add type annotations to the new test.
Copilot AI balanced review requested due to automatic review settings October 10, 2026 12:50

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants