Skip to content

Textual tool calls are never parsed on cloud-slug models: TextDialectRecovery::Auto is unreachable because we assert native_tool_calling: true #6562

Description

@M3gA-Mind

Summary

A model's textual tool call is never offered to the parser, so it renders to the user as raw markup and the action never happens. The cause is not the tool-call grammar — the grammar parses the payload correctly when asked. It is that TextDialectRecovery::Auto, the shipped default, is unreachable for every cloud-slug model, because OpenHuman asserts native_tool_calling: true on every construction path.

Observed on openrouter/deepseek/deepseek-v4-flash: 5 of 100 turn-state snapshots in one evening carry leaked markup, across 5 distinct threads.

The visible symptom

A turn made 46 web calls and ended with this as its visible assistant reply. No directory was created:

<|DSML|tool_calls>
<|DSML|invoke name="shell">
<|DSML|parameter name="command" string="true">mkdir -p ~/OpenHuman/projects/life-scenarios/out</|DSML|parameter>
</|DSML|invoke>
</|DSML|tool_calls>

The grammar is not implicated — measured

Feeding that byte-exact 195-character payload from the transcript through tinytools-agent's parse():

input_chars=195 dsml_occurrences=6
calls=1 residual_text=""
  name="shell" source=InvokeXml args={"command":"mkdir -p ~/OpenHuman/projects/life-scenarios/out"}

One call, correct name, correct arguments, and residual_text="" — had the parser been asked, the user would have seen no markup at all and the mkdir would have run.

Two theories to retire before anyone re-treads them:

  • Not a fullwidth-character problem. parse/grammar/invoke_xml.rs:34's PREFIX is (?:[||]{1,2}\s*DSML\s*[||]{1,2}\s*|[A-Za-z_][\w.-]*:)? — U+FF5C is explicitly handled, and parse/grammar/tagged.rs's module doc names "the fullwidth | those templates actually emit".
  • Not a malformed payload. It is syntactically complete: 2 matched tool_calls> markers, 1 invoke name=, 1 parameter name=, 3 closers, nothing truncated.

Root cause: the gate decides before the parser is consulted

vendor/tinyagents/crates/tinyagents-harness/src/agent_loop/run_loop.rs:1112:

if forced_text_dialect || text_dialect_recovery_enabled {
    recover_text_dialect_calls(ctx, &mut response, &call_id, &recovery);
}

and :1046-1054:

let text_dialect_recovery_enabled = match self.policy.text_dialect_recovery {
    TextDialectRecovery::Off  => false,
    TextDialectRecovery::On   => true,
    TextDialectRecovery::Auto => !binding.model.profile()
        .map(|profile| profile.tool_calling).unwrap_or(false),
};

The rationale is stated at :1043-1045: "a model that still answered in prose was explaining or quoting the format, not making a call (I-2)."

That assumption is falsified by the transcript

thread-01035d47, record 42, iteration 15 — the same assistant message carries:

  • 23 native tool_calls (web_search_tool / web_fetch), all correctly dispatched; and
  • the DSML shell call in content, never dispatched.

A response concurrently making 23 correct native calls and emitting one well-formed invoke is not explaining the format. Mixed-mode responses are the case I-2 does not cover, and the heuristic cannot see them because it keys on the model's advertised capability, not on the response.

Why Auto can never fire here — and the blast radius is the whole fleet

Auto asks whether the model reports tool_calling. For an OpenAI-compatible provider, derive_profile sets it unconditionally:

vendor/tinyagents/vendor/tinyinference/crates/tinyinference-llm/src/providers/openai/transport.rs:285
    tool_calling: true,

and OpenHuman then actively asserts it. Three entry points pass a literal true:

crates/openhuman-core/src/inference/provider/factory/cloud_slug.rs:218   (..., config, true)
crates/openhuman-core/src/inference/provider/factory/cloud_slug.rs:245   (..., config, true)
crates/openhuman-core/src/inference/provider/factory/turn_model.rs:33    (..., temperature, true)

threaded to cloud_slug.rs:413 as native_tool_calling: Some(native_tool_calling), with our own backend model defaulting the same way at inference/provider/openhuman_backend_model.rs:101 (native_tool_calling: true).

So profile.tool_calling is always true on a cloud-slug path, Auto always evaluates to false, and recover_text_dialect_calls is never reached — for every BYOK/cloud-slug model, not just this one.

Auto is the documented default. The shipped default recovery policy is therefore effectively Off across the fleet — not by decision, but as a side effect of three hardcoded booleans. This leak is one visible instance of a policy that cannot fire.

Scale, measured

Scanning the turn-state store (memory/conversations/turn_states/), 5 of 100 snapshots contain leaked DSML in streamingText — the field the UI renders — across 5 distinct threads on one model in one evening:

thread lifecycle iteration leaked tail
thread-01035d47 completed 16 / 15 shell — mkdir -p …/life-scenarios/out
thread-1964c9a4 completed 15 / 15 shell — echo "checking for notion result"
thread-00197ae0 interrupted 14 / 15 parameter name="path" — …/out/dist/post.md
thread-a4eb680d completed 9 / 15 apply_patch (see below)
thread-10a09b16 completed 7 / 50 markup present, not at the tail

Three of the five are at, over, or one short of the iteration ceiling, which is consistent with the leak correlating with long, high-volume turns.

Why this model, and why it will recur

<|DSML|tool_calls> is DeepSeek's own native chat-template format. OpenRouter serves the model behind an OpenAI-style function-calling API, so the model is speaking a translated protocol while its template natively emits DSML, and it falls back to its own format mid-response. That predicts recurrence on this family and on any model whose native template differs from the API surface it is served behind.

A separate case, not a grammar gap either

thread-a4eb680d's leak is a different shape: 244,052 characters carrying 447 DSML markers, 149 invoke name= openers and 148 tool_calls> markers in one assistant message. It opens legitimately with an apply_patch whose edits JSON value is never terminated, then restarts <|DSML|tool_calls><|DSML|invoke name="apply_patch"> inside its own value ~148 times, escaping depth compounding each round.

That is a model repetition loop, not a payload the grammar declines. A 148-deep self-nesting value is arguably unparseable by design, and the owner is whatever should have stopped the repetition — run_loop's repeat_progress middleware appears in that thread's journal, so a breaker may exist and not have fired. Filed here as context; it needs its own analysis and is not claimed to share this root cause.

The constraint any fix has to respect

Recovery parses calls out of model-visible content. Today's gate means a native-capable model's content is never scanned, which is incidentally a safety property: content cannot manufacture tool calls. I-2 justifies the gate on capability grounds, not security grounds, so that property is currently unclaimed and a future change could trade it away without noticing it existed.

Loosening the gate to "scan when the response contains both native calls and text-dialect markup" creates a path where text that reaches the model — injected tool output, a fetched page, a file it read — can induce it to echo markup that then executes. This is not hypothetical on this path: the same run fetched 8 live third-party pages and the log shows detector [prompt_injection] … source="agent.tool_output" firing throughout.

Two narrower rules were considered and both fail:

  • "Recover only calls naming tools already on the belt" — no protection. The leaked call names shell, which is on the belt.
  • "Recover only when the native channel returned nothing" — does not fit the observed case. The native channel returned 23 calls; the DSML one was additional.

Suggested direction: detect and surface, never dispatch

When the gate has skipped recovery and the content still contains a well-formed text-dialect invoke: strip it from the visible reply, and emit a turn-level signal — an audit event plus an instruction to the model that its textual call was not dispatched and must be re-issued natively.

That fixes both harms (markup shown as prose; the silently dropped action) while keeping content non-executable, so an echoed marker from a fetched page costs a wasted turn and a log line rather than a shell run. If the model was genuinely quoting the format — I-2's case — the cost is one redundant nudge.

Existing workaround

OPENHUMAN_TOOL_DISPATCHER=xml (config/schema/agent.rs:259, overlaid at config/load/env_overlay.rs:94-104, documented auto | native | xml | pformat | python | typescript) forces the text dialect today and would make these calls dispatch, with no code change. Not recommended as the fix — it switches the whole protocol to text mode for every call and re-opens the content-as-calls surface globally, which is the surface this issue argues for not widening.

Suggested split

  1. The three hardcoded trues defeat a policy we ship as Auto. Either they belong per-model, or Auto should key on evidence in the response rather than a capability we assert on the model's behalf. This stands on its own merits independently of the leak.
  2. Detect-and-surface in the harness, with the content-never-executes property claimed explicitly rather than left incidental.

A human should pick between them; this issue deliberately proposes no patch.

Verification notes

Every code citation above was read by extracting file bytes through a script rather than from rendered terminal output, after this session's terminal was observed substituting substrings — a search for the literal tool_calling printed lines reading pub n: bool,. Counts: 49 occurrences of tool_calling across 13 files against 872 of the two-character n:. Line numbers were re-derived with full paths after an earlier probe printed basenames only, which made mod.rs:377 ambiguous between several files (it is pub fn permissive(), a mocks-and-tests profile, not the production path).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority: p1Next. Wrong behaviour a user will hit, or a security weakness behind a condition.

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions