Skip to content

[Feature] Add constrained DeepSeek V4 tool generation to DSpark K7 #265

Description

@sjduan

Summary

Add generation-time, schema-aware tool-call constraints for DeepSeek V4 Flash DSpark K7. The current
local implementation uses a request-scoped XGrammar structural-tag matcher on the Worker host and a
fixed packed mask ABI in PyPTO-Lib. It retains one fused K7 device dispatch per decode step; the
existing parser still turns generated DSML into OpenAI-compatible tool-call responses. This is an
implementation progress report, not a claim that the feature has merged upstream.

Area

Model integration

Motivation / Use Case

Tool definitions supplied by different agent clients can require different function names and JSON
fields. Prompting and post-generation parsing alone cannot prevent missing required fields or unknown
fields such as cmd instead of a declared command. The server should enforce the request's tool
contract while sampling, without hard-coded client-specific field rewrites.

Proposed API / Behavior

  • Keep the DeepSeek V4 tool-choice trigger semantics aligned with the inspected vLLM implementation:
    required and named choices, or auto with at least one strict:true tool, activate structural-tag
    generation. Ordinary non-strict auto remains free generation; this parity does not guarantee
    schema-valid arguments for that mode.
  • Preflight active schemas before an SSE response starts and return HTTP 400 for unsupported schemas
    or a missing provider. Send a serializable constraint specification over Engine→Worker IPC once per
    request; keep matcher state on the Worker host and advance it only with actually accepted tokens.
  • Validate the longest legal K7 draft prefix, stage per-target-row masks plus that prefix length, and
    apply the mask before device greedy selection. Cap device prepare and accept to the validated
    prefix. Unconstrained requests use the same ABI with an all-allowed mask. PyPTO-Lib has no
    XGrammar, JSON Schema, DSML, or client-specific dependency.

Current local progress: Serving 36279eb and PyPTO-Lib 60ecef8 on matching feature branches;
their K7 positional ABI must be deployed together. Host/ABI tests (225 total; 108 focused rerun),
masked-greedy and draft-cap device goldens, and 16-device a2a3 K7 HTTP smoke passed. Smoke coverage
includes streaming/non-streaming required, named choice, strict auto, unconstrained requests,
mixed concurrency, and a reasoning-to-tool request. An adversarial prompt requesting cmd yielded
the schema-required command under the constrained path. Neither repository change has an upstream
PR yet.

Alternatives Considered

  • Prompt-only instruction or output-time field renaming: cannot guarantee schema correctness and
    would bind Serving to particular agent clients.
  • Round-trip logits to the Host for masking, or add a separate sampling graph: does not preserve the
    current fused K7 path. The fixed device mask ABI keeps the model dispatch boundary unchanged.
  • Force every ordinary auto request into strict generation by default: differs from inspected vLLM
    behavior and needs a separate, explicit policy decision.

Additional Context

Known gaps and deferred work relative to a complete vLLM-like implementation:

  1. Mid-decode preemption/recompute: constrained requests currently cannot be selected as
    preemption victims. Serving's current recompute resets num_computed_tokens and re-Prefills only
    the original prompt, while retaining already generated output IDs; it does not rebuild KV from the
    entire committed token history. Recreating a matcher alone would put KV, output and grammar at
    different positions. vLLM retains the full prompt-plus-output token sequence for recompute. Fix
    generic Serving replay, in-flight boundaries, request-local drafter/grammar restoration, and
    prefix-cache interactions before enabling constrained-request preemption. This is deferred from
    the first version.
  2. Ordinary non-strict auto: free generation matches the inspected vLLM trigger rule, but does
    not prevent malformed tool arguments. A possible opt-in “constrain all auto tools” policy is
    deferred and must be evaluated separately; do not silently change default semantics.
  3. Breadth: unlike vLLM's more general structured-output integration, this first version only
    covers DeepSeek V4 Flash DSpark K7 with XGrammar and device greedy sampling. Other model formats,
    providers and non-greedy sampling have no constraint implementation yet.
  4. First-version closure: verify stream cancellation/abort and recovery of subsequent requests,
    streaming reasoning-to-tool transitions, parallel/multiple tool calls, truncation, and explicit
    early rejection of unsupported non-greedy sampling. Document or package the XGrammar dependency
    and the paired Serving/Lib ABI. Current non-deterministic device failures use Serving's existing
    step-level error handling, which can stop other requests in the same batch; request-level fault
    isolation is not claimed. Quantitative mask-staging/performance work remains exploratory.

References in the local source trees: pypto_serving/serving/constraints/,
pypto_serving/model/deepseek_dspark/npu_runner.py,
models/deepseek_v4_flash_dspark/lm_head.py, and
models/deepseek_v4_flash_dspark/decode_prepare.py.
Related general capability inventory: pypto-serving #7.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions