Summary
Add generation-time, schema-aware tool-call constraints for DeepSeek V4 Flash DSpark K7. The current
local implementation uses a request-scoped XGrammar structural-tag matcher on the Worker host and a
fixed packed mask ABI in PyPTO-Lib. It retains one fused K7 device dispatch per decode step; the
existing parser still turns generated DSML into OpenAI-compatible tool-call responses. This is an
implementation progress report, not a claim that the feature has merged upstream.
Area
Model integration
Motivation / Use Case
Tool definitions supplied by different agent clients can require different function names and JSON
fields. Prompting and post-generation parsing alone cannot prevent missing required fields or unknown
fields such as cmd instead of a declared command. The server should enforce the request's tool
contract while sampling, without hard-coded client-specific field rewrites.
Proposed API / Behavior
- Keep the DeepSeek V4 tool-choice trigger semantics aligned with the inspected vLLM implementation:
required and named choices, or auto with at least one strict:true tool, activate structural-tag
generation. Ordinary non-strict auto remains free generation; this parity does not guarantee
schema-valid arguments for that mode.
- Preflight active schemas before an SSE response starts and return HTTP 400 for unsupported schemas
or a missing provider. Send a serializable constraint specification over Engine→Worker IPC once per
request; keep matcher state on the Worker host and advance it only with actually accepted tokens.
- Validate the longest legal K7 draft prefix, stage per-target-row masks plus that prefix length, and
apply the mask before device greedy selection. Cap device prepare and accept to the validated
prefix. Unconstrained requests use the same ABI with an all-allowed mask. PyPTO-Lib has no
XGrammar, JSON Schema, DSML, or client-specific dependency.
Current local progress: Serving 36279eb and PyPTO-Lib 60ecef8 on matching feature branches;
their K7 positional ABI must be deployed together. Host/ABI tests (225 total; 108 focused rerun),
masked-greedy and draft-cap device goldens, and 16-device a2a3 K7 HTTP smoke passed. Smoke coverage
includes streaming/non-streaming required, named choice, strict auto, unconstrained requests,
mixed concurrency, and a reasoning-to-tool request. An adversarial prompt requesting cmd yielded
the schema-required command under the constrained path. Neither repository change has an upstream
PR yet.
Alternatives Considered
- Prompt-only instruction or output-time field renaming: cannot guarantee schema correctness and
would bind Serving to particular agent clients.
- Round-trip logits to the Host for masking, or add a separate sampling graph: does not preserve the
current fused K7 path. The fixed device mask ABI keeps the model dispatch boundary unchanged.
- Force every ordinary
auto request into strict generation by default: differs from inspected vLLM
behavior and needs a separate, explicit policy decision.
Additional Context
Known gaps and deferred work relative to a complete vLLM-like implementation:
- Mid-decode preemption/recompute: constrained requests currently cannot be selected as
preemption victims. Serving's current recompute resets num_computed_tokens and re-Prefills only
the original prompt, while retaining already generated output IDs; it does not rebuild KV from the
entire committed token history. Recreating a matcher alone would put KV, output and grammar at
different positions. vLLM retains the full prompt-plus-output token sequence for recompute. Fix
generic Serving replay, in-flight boundaries, request-local drafter/grammar restoration, and
prefix-cache interactions before enabling constrained-request preemption. This is deferred from
the first version.
- Ordinary non-strict
auto: free generation matches the inspected vLLM trigger rule, but does
not prevent malformed tool arguments. A possible opt-in “constrain all auto tools” policy is
deferred and must be evaluated separately; do not silently change default semantics.
- Breadth: unlike vLLM's more general structured-output integration, this first version only
covers DeepSeek V4 Flash DSpark K7 with XGrammar and device greedy sampling. Other model formats,
providers and non-greedy sampling have no constraint implementation yet.
- First-version closure: verify stream cancellation/abort and recovery of subsequent requests,
streaming reasoning-to-tool transitions, parallel/multiple tool calls, truncation, and explicit
early rejection of unsupported non-greedy sampling. Document or package the XGrammar dependency
and the paired Serving/Lib ABI. Current non-deterministic device failures use Serving's existing
step-level error handling, which can stop other requests in the same batch; request-level fault
isolation is not claimed. Quantitative mask-staging/performance work remains exploratory.
References in the local source trees: pypto_serving/serving/constraints/,
pypto_serving/model/deepseek_dspark/npu_runner.py,
models/deepseek_v4_flash_dspark/lm_head.py, and
models/deepseek_v4_flash_dspark/decode_prepare.py.
Related general capability inventory: pypto-serving #7.
Summary
Add generation-time, schema-aware tool-call constraints for DeepSeek V4 Flash DSpark K7. The current
local implementation uses a request-scoped XGrammar structural-tag matcher on the Worker host and a
fixed packed mask ABI in PyPTO-Lib. It retains one fused K7 device dispatch per decode step; the
existing parser still turns generated DSML into OpenAI-compatible tool-call responses. This is an
implementation progress report, not a claim that the feature has merged upstream.
Area
Model integration
Motivation / Use Case
Tool definitions supplied by different agent clients can require different function names and JSON
fields. Prompting and post-generation parsing alone cannot prevent missing required fields or unknown
fields such as
cmdinstead of a declaredcommand. The server should enforce the request's toolcontract while sampling, without hard-coded client-specific field rewrites.
Proposed API / Behavior
requiredand named choices, orautowith at least onestrict:truetool, activate structural-taggeneration. Ordinary non-strict
autoremains free generation; this parity does not guaranteeschema-valid arguments for that mode.
or a missing provider. Send a serializable constraint specification over Engine→Worker IPC once per
request; keep matcher state on the Worker host and advance it only with actually accepted tokens.
apply the mask before device greedy selection. Cap device prepare and accept to the validated
prefix. Unconstrained requests use the same ABI with an all-allowed mask. PyPTO-Lib has no
XGrammar, JSON Schema, DSML, or client-specific dependency.
Current local progress: Serving
36279eband PyPTO-Lib60ecef8on matching feature branches;their K7 positional ABI must be deployed together. Host/ABI tests (225 total; 108 focused rerun),
masked-greedy and draft-cap device goldens, and 16-device a2a3 K7 HTTP smoke passed. Smoke coverage
includes streaming/non-streaming
required, named choice, strictauto, unconstrained requests,mixed concurrency, and a reasoning-to-tool request. An adversarial prompt requesting
cmdyieldedthe schema-required
commandunder the constrained path. Neither repository change has an upstreamPR yet.
Alternatives Considered
would bind Serving to particular agent clients.
current fused K7 path. The fixed device mask ABI keeps the model dispatch boundary unchanged.
autorequest into strict generation by default: differs from inspected vLLMbehavior and needs a separate, explicit policy decision.
Additional Context
Known gaps and deferred work relative to a complete vLLM-like implementation:
preemption victims. Serving's current recompute resets
num_computed_tokensand re-Prefills onlythe original prompt, while retaining already generated output IDs; it does not rebuild KV from the
entire committed token history. Recreating a matcher alone would put KV, output and grammar at
different positions. vLLM retains the full prompt-plus-output token sequence for recompute. Fix
generic Serving replay, in-flight boundaries, request-local drafter/grammar restoration, and
prefix-cache interactions before enabling constrained-request preemption. This is deferred from
the first version.
auto: free generation matches the inspected vLLM trigger rule, but doesnot prevent malformed tool arguments. A possible opt-in “constrain all auto tools” policy is
deferred and must be evaluated separately; do not silently change default semantics.
covers DeepSeek V4 Flash DSpark K7 with XGrammar and device greedy sampling. Other model formats,
providers and non-greedy sampling have no constraint implementation yet.
streaming reasoning-to-tool transitions, parallel/multiple tool calls, truncation, and explicit
early rejection of unsupported non-greedy sampling. Document or package the XGrammar dependency
and the paired Serving/Lib ABI. Current non-deterministic device failures use Serving's existing
step-level error handling, which can stop other requests in the same batch; request-level fault
isolation is not claimed. Quantitative mask-staging/performance work remains exploratory.
References in the local source trees:
pypto_serving/serving/constraints/,pypto_serving/model/deepseek_dspark/npu_runner.py,models/deepseek_v4_flash_dspark/lm_head.py, andmodels/deepseek_v4_flash_dspark/decode_prepare.py.Related general capability inventory: pypto-serving #7.