MegatronBridgePosttrainTrainer.train passes megatron.bridge.training.vlm_step.forward_step for all models, including text-only ones. Without pipeline parallelism, vlm_step.get_batch pads tokens/labels/loss_mask/position_ids to a multiple of 128 but leaves cu_seqlens unchanged.
Repro: train posttrain with packed_sequence: true (e.g. the Qwen3-32B SFT posttrain config) and a packed row whose length is not a multiple of 128 (e.g. 30,992 tokens). The first forward fails in _apply_rotary_pos_emb_thd:
split_with_sizes expects split_sizes to sum exactly to 31104 ... but got split_sizes=[30992]
Only rows padded all the way to seq_length (pad_to_max_length=True) avoid it.
Suggested fix: use gpt_step.forward_step for text-only models, or append the padded length to cu_seqlens whenever get_batch pads. Swapping in gpt_step.forward_step works for us: losses match HF on the same packed rows.
Environment: rocm/primus:v26.7 (Primus release/v26.7, Megatron-Bridge 9577b12), MI350X.
MegatronBridgePosttrainTrainer.trainpassesmegatron.bridge.training.vlm_step.forward_stepfor all models, including text-only ones. Without pipeline parallelism,vlm_step.get_batchpads tokens/labels/loss_mask/position_ids to a multiple of 128 but leavescu_seqlensunchanged.Repro:
train posttrainwithpacked_sequence: true(e.g. the Qwen3-32B SFT posttrain config) and a packed row whose length is not a multiple of 128 (e.g. 30,992 tokens). The first forward fails in_apply_rotary_pos_emb_thd:Only rows padded all the way to
seq_length(pad_to_max_length=True) avoid it.Suggested fix: use
gpt_step.forward_stepfor text-only models, or append the padded length tocu_seqlenswheneverget_batchpads. Swapping ingpt_step.forward_stepworks for us: losses match HF on the same packed rows.Environment:
rocm/primus:v26.7(Primusrelease/v26.7, Megatron-Bridge9577b12), MI350X.