Skip to content

[Test] Run text and image cases through shared GLM parity checks - #2140

Closed
xs1997zju wants to merge 1 commit into
feat/glm53flash-f7-logits-parityfrom
feat/glm53flash-f8-multimodal-parity
Closed

xs1997zju wants to merge 1 commit into
feat/glm53flash-f7-logits-parityfrom
feat/glm53flash-f8-multimodal-parity

Conversation

@xs1997zju

@xs1997zju xs1997zju commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Stack created with GitHub Stacks CLI • Give Feedback 💬

Stack #2112 (bottom to top):

  1. [Feature] Add GLM-5.3-Flash F0: 25B cropped reference checkpoint builder #2105 — F0: reduced checkpoint
  2. [Feature] Add GLM-5.3-Flash F3: Kimi Delta Attention (KDA) #2106 — F3: KDA
  3. [Feature] Add GLM-5.3-Flash F4: mHC four-stream residual #2107 — F4: mHC
  4. [Feature] Add GLM-5.3-Flash F5: NoPE DSA + KPool indexer + clamped SwiGLU #2108 — F5: DSA
  5. [Feature] Add GLM-5.3-Flash F1: VL data preprocessing pipeline #2109 — F1: VL data
  6. [Feature] Add GLM-5.3-Flash F2: vision tower + projector (eager) #2110 — F2: vision tower
  7. [Feature] Add GLM-5.3-Flash F6 core: text model + MTP + compose model #2111 — F6: text model and compose model
  8. [Test] Check GLM-5.3 end-to-end logits alongside loss #2139 — F7: logits assertions in the existing accuracy test
  9. [Test] Run text and image cases through shared GLM parity checks #2140 — F8: shared text/image accuracy cases ← you are here

Base: feat/glm53flash-f7-logits-parity (#2139). Review this layer against its base.

The existing HF↔XTuner real-checkpoint accuracy test now runs its four original text cases plus one single-image and one two-image case through the same forward/loss/logits comparison loop. Only tests/model/test_glm53_text_moe.py changes; no separate replay entry point or dataset pipeline is added.

Use the full compose model and the existing repository cat image. The image cases share processor outputs between HF and XTuner, mask image placeholders from labels, and compare full-vocabulary logits at the final eight supervised next-token positions. Text and image loss checks are evaluated separately with the existing thresholds; per-sample logits thresholds remain relative L2 <5% and cosine >0.998.

XT vision uses FlashAttention; native HF uses eager because the installed Transformers vision implementation does not support FlashAttention 2. This does not claim identical attention kernels.

Validation:

  • Single-node 8×H200, local reduced GLM-5.3-Flash checkpoint: 3 passed, original EP1/4/8 parameterization. Each configuration runs all six cases.
  • Maximum sampled logits relative L2 across the tested configurations/ranks: single image 1.382283%, two images 1.245808%. All text/image loss and logits checks passed.
  • Ruff and diff whitespace checks passed.

Only eval/no-grad forward was tested. No backward, optimizer updates, packing, MTP, video, or SP2 validation is claimed for this layer.

@xs1997zju
xs1997zju added this pull request to stack #2112 October 8, 2026 13:11
@xs1997zju
xs1997zju marked this pull request as ready for review October 8, 2026 13:13
@xs1997zju

Copy link
Copy Markdown
Collaborator Author

Superseded by #2141, an ordinary PR targeting the existing F6 layer (#2111), combining the logits and image parity tests without adding a stack layer.

@xs1997zju xs1997zju closed this Oct 9, 2026
@jayhenry
jayhenry deleted the feat/glm53flash-f8-multimodal-parity branch October 9, 2026 06:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant