Summary
A model's textual tool call is never offered to the parser, so it renders to the user as raw markup and the action never happens. The cause is not the tool-call grammar — the grammar parses the payload correctly when asked. It is that TextDialectRecovery::Auto, the shipped default, is unreachable for every cloud-slug model, because OpenHuman asserts native_tool_calling: true on every construction path.
Observed on openrouter/deepseek/deepseek-v4-flash: 5 of 100 turn-state snapshots in one evening carry leaked markup, across 5 distinct threads.
The visible symptom
A turn made 46 web calls and ended with this as its visible assistant reply. No directory was created:
<|DSML|tool_calls>
<|DSML|invoke name="shell">
<|DSML|parameter name="command" string="true">mkdir -p ~/OpenHuman/projects/life-scenarios/out</|DSML|parameter>
</|DSML|invoke>
</|DSML|tool_calls>
The grammar is not implicated — measured
Feeding that byte-exact 195-character payload from the transcript through tinytools-agent's parse():
input_chars=195 dsml_occurrences=6
calls=1 residual_text=""
name="shell" source=InvokeXml args={"command":"mkdir -p ~/OpenHuman/projects/life-scenarios/out"}
One call, correct name, correct arguments, and residual_text="" — had the parser been asked, the user would have seen no markup at all and the mkdir would have run.
Two theories to retire before anyone re-treads them:
- Not a fullwidth-character problem.
parse/grammar/invoke_xml.rs:34's PREFIX is (?:[||]{1,2}\s*DSML\s*[||]{1,2}\s*|[A-Za-z_][\w.-]*:)? — U+FF5C is explicitly handled, and parse/grammar/tagged.rs's module doc names "the fullwidth | those templates actually emit".
- Not a malformed payload. It is syntactically complete: 2 matched
tool_calls> markers, 1 invoke name=, 1 parameter name=, 3 closers, nothing truncated.
Root cause: the gate decides before the parser is consulted
vendor/tinyagents/crates/tinyagents-harness/src/agent_loop/run_loop.rs:1112:
if forced_text_dialect || text_dialect_recovery_enabled {
recover_text_dialect_calls(ctx, &mut response, &call_id, &recovery);
}
and :1046-1054:
let text_dialect_recovery_enabled = match self.policy.text_dialect_recovery {
TextDialectRecovery::Off => false,
TextDialectRecovery::On => true,
TextDialectRecovery::Auto => !binding.model.profile()
.map(|profile| profile.tool_calling).unwrap_or(false),
};
The rationale is stated at :1043-1045: "a model that still answered in prose was explaining or quoting the format, not making a call (I-2)."
That assumption is falsified by the transcript
thread-01035d47, record 42, iteration 15 — the same assistant message carries:
- 23 native
tool_calls (web_search_tool / web_fetch), all correctly dispatched; and
- the DSML
shell call in content, never dispatched.
A response concurrently making 23 correct native calls and emitting one well-formed invoke is not explaining the format. Mixed-mode responses are the case I-2 does not cover, and the heuristic cannot see them because it keys on the model's advertised capability, not on the response.
Why Auto can never fire here — and the blast radius is the whole fleet
Auto asks whether the model reports tool_calling. For an OpenAI-compatible provider, derive_profile sets it unconditionally:
vendor/tinyagents/vendor/tinyinference/crates/tinyinference-llm/src/providers/openai/transport.rs:285
tool_calling: true,
and OpenHuman then actively asserts it. Three entry points pass a literal true:
crates/openhuman-core/src/inference/provider/factory/cloud_slug.rs:218 (..., config, true)
crates/openhuman-core/src/inference/provider/factory/cloud_slug.rs:245 (..., config, true)
crates/openhuman-core/src/inference/provider/factory/turn_model.rs:33 (..., temperature, true)
threaded to cloud_slug.rs:413 as native_tool_calling: Some(native_tool_calling), with our own backend model defaulting the same way at inference/provider/openhuman_backend_model.rs:101 (native_tool_calling: true).
So profile.tool_calling is always true on a cloud-slug path, Auto always evaluates to false, and recover_text_dialect_calls is never reached — for every BYOK/cloud-slug model, not just this one.
Auto is the documented default. The shipped default recovery policy is therefore effectively Off across the fleet — not by decision, but as a side effect of three hardcoded booleans. This leak is one visible instance of a policy that cannot fire.
Scale, measured
Scanning the turn-state store (memory/conversations/turn_states/), 5 of 100 snapshots contain leaked DSML in streamingText — the field the UI renders — across 5 distinct threads on one model in one evening:
| thread |
lifecycle |
iteration |
leaked tail |
thread-01035d47 |
completed |
16 / 15 |
shell — mkdir -p …/life-scenarios/out |
thread-1964c9a4 |
completed |
15 / 15 |
shell — echo "checking for notion result" |
thread-00197ae0 |
interrupted |
14 / 15 |
parameter name="path" — …/out/dist/post.md |
thread-a4eb680d |
completed |
9 / 15 |
apply_patch (see below) |
thread-10a09b16 |
completed |
7 / 50 |
markup present, not at the tail |
Three of the five are at, over, or one short of the iteration ceiling, which is consistent with the leak correlating with long, high-volume turns.
Why this model, and why it will recur
<|DSML|tool_calls> is DeepSeek's own native chat-template format. OpenRouter serves the model behind an OpenAI-style function-calling API, so the model is speaking a translated protocol while its template natively emits DSML, and it falls back to its own format mid-response. That predicts recurrence on this family and on any model whose native template differs from the API surface it is served behind.
A separate case, not a grammar gap either
thread-a4eb680d's leak is a different shape: 244,052 characters carrying 447 DSML markers, 149 invoke name= openers and 148 tool_calls> markers in one assistant message. It opens legitimately with an apply_patch whose edits JSON value is never terminated, then restarts <|DSML|tool_calls><|DSML|invoke name="apply_patch"> inside its own value ~148 times, escaping depth compounding each round.
That is a model repetition loop, not a payload the grammar declines. A 148-deep self-nesting value is arguably unparseable by design, and the owner is whatever should have stopped the repetition — run_loop's repeat_progress middleware appears in that thread's journal, so a breaker may exist and not have fired. Filed here as context; it needs its own analysis and is not claimed to share this root cause.
The constraint any fix has to respect
Recovery parses calls out of model-visible content. Today's gate means a native-capable model's content is never scanned, which is incidentally a safety property: content cannot manufacture tool calls. I-2 justifies the gate on capability grounds, not security grounds, so that property is currently unclaimed and a future change could trade it away without noticing it existed.
Loosening the gate to "scan when the response contains both native calls and text-dialect markup" creates a path where text that reaches the model — injected tool output, a fetched page, a file it read — can induce it to echo markup that then executes. This is not hypothetical on this path: the same run fetched 8 live third-party pages and the log shows detector [prompt_injection] … source="agent.tool_output" firing throughout.
Two narrower rules were considered and both fail:
- "Recover only calls naming tools already on the belt" — no protection. The leaked call names
shell, which is on the belt.
- "Recover only when the native channel returned nothing" — does not fit the observed case. The native channel returned 23 calls; the DSML one was additional.
Suggested direction: detect and surface, never dispatch
When the gate has skipped recovery and the content still contains a well-formed text-dialect invoke: strip it from the visible reply, and emit a turn-level signal — an audit event plus an instruction to the model that its textual call was not dispatched and must be re-issued natively.
That fixes both harms (markup shown as prose; the silently dropped action) while keeping content non-executable, so an echoed marker from a fetched page costs a wasted turn and a log line rather than a shell run. If the model was genuinely quoting the format — I-2's case — the cost is one redundant nudge.
Existing workaround
OPENHUMAN_TOOL_DISPATCHER=xml (config/schema/agent.rs:259, overlaid at config/load/env_overlay.rs:94-104, documented auto | native | xml | pformat | python | typescript) forces the text dialect today and would make these calls dispatch, with no code change. Not recommended as the fix — it switches the whole protocol to text mode for every call and re-opens the content-as-calls surface globally, which is the surface this issue argues for not widening.
Suggested split
- The three hardcoded
trues defeat a policy we ship as Auto. Either they belong per-model, or Auto should key on evidence in the response rather than a capability we assert on the model's behalf. This stands on its own merits independently of the leak.
- Detect-and-surface in the harness, with the content-never-executes property claimed explicitly rather than left incidental.
A human should pick between them; this issue deliberately proposes no patch.
Verification notes
Every code citation above was read by extracting file bytes through a script rather than from rendered terminal output, after this session's terminal was observed substituting substrings — a search for the literal tool_calling printed lines reading pub n: bool,. Counts: 49 occurrences of tool_calling across 13 files against 872 of the two-character n:. Line numbers were re-derived with full paths after an earlier probe printed basenames only, which made mod.rs:377 ambiguous between several files (it is pub fn permissive(), a mocks-and-tests profile, not the production path).
Summary
A model's textual tool call is never offered to the parser, so it renders to the user as raw markup and the action never happens. The cause is not the tool-call grammar — the grammar parses the payload correctly when asked. It is that
TextDialectRecovery::Auto, the shipped default, is unreachable for every cloud-slug model, because OpenHuman assertsnative_tool_calling: trueon every construction path.Observed on
openrouter/deepseek/deepseek-v4-flash: 5 of 100 turn-state snapshots in one evening carry leaked markup, across 5 distinct threads.The visible symptom
A turn made 46 web calls and ended with this as its visible assistant reply. No directory was created:
The grammar is not implicated — measured
Feeding that byte-exact 195-character payload from the transcript through
tinytools-agent'sparse():One call, correct name, correct arguments, and
residual_text=""— had the parser been asked, the user would have seen no markup at all and themkdirwould have run.Two theories to retire before anyone re-treads them:
parse/grammar/invoke_xml.rs:34'sPREFIXis(?:[||]{1,2}\s*DSML\s*[||]{1,2}\s*|[A-Za-z_][\w.-]*:)?— U+FF5C is explicitly handled, andparse/grammar/tagged.rs's module doc names "the fullwidth|those templates actually emit".tool_calls>markers, 1invoke name=, 1parameter name=, 3 closers, nothing truncated.Root cause: the gate decides before the parser is consulted
vendor/tinyagents/crates/tinyagents-harness/src/agent_loop/run_loop.rs:1112:and
:1046-1054:The rationale is stated at
:1043-1045: "a model that still answered in prose was explaining or quoting the format, not making a call (I-2)."That assumption is falsified by the transcript
thread-01035d47, record 42, iteration 15 — the same assistant message carries:tool_calls(web_search_tool/web_fetch), all correctly dispatched; andshellcall incontent, never dispatched.A response concurrently making 23 correct native calls and emitting one well-formed invoke is not explaining the format. Mixed-mode responses are the case I-2 does not cover, and the heuristic cannot see them because it keys on the model's advertised capability, not on the response.
Why
Autocan never fire here — and the blast radius is the whole fleetAutoasks whether the model reportstool_calling. For an OpenAI-compatible provider,derive_profilesets it unconditionally:and OpenHuman then actively asserts it. Three entry points pass a literal
true:threaded to
cloud_slug.rs:413asnative_tool_calling: Some(native_tool_calling), with our own backend model defaulting the same way atinference/provider/openhuman_backend_model.rs:101(native_tool_calling: true).So
profile.tool_callingis always true on a cloud-slug path,Autoalways evaluates tofalse, andrecover_text_dialect_callsis never reached — for every BYOK/cloud-slug model, not just this one.Autois the documented default. The shipped default recovery policy is therefore effectivelyOffacross the fleet — not by decision, but as a side effect of three hardcoded booleans. This leak is one visible instance of a policy that cannot fire.Scale, measured
Scanning the turn-state store (
memory/conversations/turn_states/), 5 of 100 snapshots contain leaked DSML instreamingText— the field the UI renders — across 5 distinct threads on one model in one evening:thread-01035d47shell—mkdir -p …/life-scenarios/outthread-1964c9a4shell—echo "checking for notion result"thread-00197ae0parameter name="path"—…/out/dist/post.mdthread-a4eb680dapply_patch(see below)thread-10a09b16Three of the five are at, over, or one short of the iteration ceiling, which is consistent with the leak correlating with long, high-volume turns.
Why this model, and why it will recur
<|DSML|tool_calls>is DeepSeek's own native chat-template format. OpenRouter serves the model behind an OpenAI-style function-calling API, so the model is speaking a translated protocol while its template natively emits DSML, and it falls back to its own format mid-response. That predicts recurrence on this family and on any model whose native template differs from the API surface it is served behind.A separate case, not a grammar gap either
thread-a4eb680d's leak is a different shape: 244,052 characters carrying 447DSMLmarkers, 149invoke name=openers and 148tool_calls>markers in one assistant message. It opens legitimately with anapply_patchwhoseeditsJSON value is never terminated, then restarts<|DSML|tool_calls><|DSML|invoke name="apply_patch">inside its own value ~148 times, escaping depth compounding each round.That is a model repetition loop, not a payload the grammar declines. A 148-deep self-nesting value is arguably unparseable by design, and the owner is whatever should have stopped the repetition —
run_loop'srepeat_progressmiddleware appears in that thread's journal, so a breaker may exist and not have fired. Filed here as context; it needs its own analysis and is not claimed to share this root cause.The constraint any fix has to respect
Recovery parses calls out of model-visible content. Today's gate means a native-capable model's content is never scanned, which is incidentally a safety property: content cannot manufacture tool calls. I-2 justifies the gate on capability grounds, not security grounds, so that property is currently unclaimed and a future change could trade it away without noticing it existed.
Loosening the gate to "scan when the response contains both native calls and text-dialect markup" creates a path where text that reaches the model — injected tool output, a fetched page, a file it read — can induce it to echo markup that then executes. This is not hypothetical on this path: the same run fetched 8 live third-party pages and the log shows
detector [prompt_injection] … source="agent.tool_output"firing throughout.Two narrower rules were considered and both fail:
shell, which is on the belt.Suggested direction: detect and surface, never dispatch
When the gate has skipped recovery and the content still contains a well-formed text-dialect invoke: strip it from the visible reply, and emit a turn-level signal — an audit event plus an instruction to the model that its textual call was not dispatched and must be re-issued natively.
That fixes both harms (markup shown as prose; the silently dropped action) while keeping content non-executable, so an echoed marker from a fetched page costs a wasted turn and a log line rather than a
shellrun. If the model was genuinely quoting the format — I-2's case — the cost is one redundant nudge.Existing workaround
OPENHUMAN_TOOL_DISPATCHER=xml(config/schema/agent.rs:259, overlaid atconfig/load/env_overlay.rs:94-104, documentedauto | native | xml | pformat | python | typescript) forces the text dialect today and would make these calls dispatch, with no code change. Not recommended as the fix — it switches the whole protocol to text mode for every call and re-opens the content-as-calls surface globally, which is the surface this issue argues for not widening.Suggested split
trues defeat a policy we ship asAuto. Either they belong per-model, orAutoshould key on evidence in the response rather than a capability we assert on the model's behalf. This stands on its own merits independently of the leak.A human should pick between them; this issue deliberately proposes no patch.
Verification notes
Every code citation above was read by extracting file bytes through a script rather than from rendered terminal output, after this session's terminal was observed substituting substrings — a search for the literal
tool_callingprinted lines readingpub n: bool,. Counts: 49 occurrences oftool_callingacross 13 files against 872 of the two-charactern:. Line numbers were re-derived with full paths after an earlier probe printed basenames only, which mademod.rs:377ambiguous between several files (it ispub fn permissive(), a mocks-and-tests profile, not the production path).