Skip to content

train posttrain with packed sequences crashes: vlm_step pads tokens but not cu_seqlens for text-only models #1194

Description

@Arist12

MegatronBridgePosttrainTrainer.train passes megatron.bridge.training.vlm_step.forward_step for all models, including text-only ones. Without pipeline parallelism, vlm_step.get_batch pads tokens/labels/loss_mask/position_ids to a multiple of 128 but leaves cu_seqlens unchanged.

Repro: train posttrain with packed_sequence: true (e.g. the Qwen3-32B SFT posttrain config) and a packed row whose length is not a multiple of 128 (e.g. 30,992 tokens). The first forward fails in _apply_rotary_pos_emb_thd:

split_with_sizes expects split_sizes to sum exactly to 31104 ... but got split_sizes=[30992]

Only rows padded all the way to seq_length (pad_to_max_length=True) avoid it.

Suggested fix: use gpt_step.forward_step for text-only models, or append the padded length to cu_seqlens whenever get_batch pads. Swapping in gpt_step.forward_step works for us: losses match HF on the same packed rows.

Environment: rocm/primus:v26.7 (Primus release/v26.7, Megatron-Bridge 9577b12), MI350X.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions