Skip to content

Rust router: support the NATS request transport (router side) #88

Description

@jiejingzhangamd

Today

The Rust router speaks HTTP only. lib.rs has said so since v0.1.0:

Configs outside this set (NATS transport, other PD connectors, edge endpoints)
are served by the Python backend.

That was a reasonable scope cut, but it has become a real fork in the road: the
Rust data plane and the NATS request transport are currently either/or, and
the two differ in ways that matter rather than in ways that are cosmetic.

Why it matters

Measured while validating graceful scale-down (see #83):

NATS HTTP
who knows what is in flight infera — it owns the request path and holds the in-flight set only the engine; the router dials it directly and never sees the request
drain exact — draining 1 in-flight NATS request(s), no polling poll the engine's /metrics behind a settle window
drain with nothing in flight 3 ms ≥ 6 s (a single zero reading cannot be told from a stale gauge)
request cancellation infera.cancel.<worker> tears down the engine connection none — a timeout or client disconnect leaves the engine generating
admission control JetStream backlog per worker, steers away from a saturated one none

The admission control result is the one worth repeating: with one deliberately
saturated worker (concurrency 1) and one fast one at limit 3, twenty requests
sent under backlog went +0 / +20 where round-robin would have been +10 / +10.
That covers the window scaling cannot — a burst shorter than a 140 s cold start
cannot be answered by adding workers.

None of this is available to a deployment that chooses the Rust data plane.

Proposed scope: router side only

The worker side stays Python (NatsRequestServer). The Rust router only needs
to be a client:

  • publish to infera.req.<token(worker_id)> with a reply inbox
  • consume the framed reply stream — rs-type ∈ {data, done, error},
    rs-status on done — and relay it as the HTTP response / SSE body
  • both request timeouts, and publish to infera.cancel.<worker> on timeout
    or client disconnect
  • route by the worker's registered request_transport, so a mixed fleet
    works and this is not a global mode switch
  • drain: unsubscribe, then wait on the in-flight set — the whole point

Deliberately not in the first pass:

  • the worker side (Rust workers are not a thing)
  • JetStream admission control — it depends on the consumer API and is cleanly
    additive once the core path works

Risks worth naming up front

Two implementations of one wire protocol. infera/common/nats_request.py is
644 lines and 17 async methods; the framing, subject naming and cancel semantics
would exist twice, and a change to either drifts silently. Worth deciding
whether the protocol gets a written spec and a cross-implementation test before
the second implementation exists, rather than after.

New dependency. async-nats in rust/router, which currently has no
messaging dependency at all.

Alternative

Leave the Rust router HTTP-only and treat "Rust data plane" and "NATS transport"
as an explicit, documented either/or. That is the status quo; this issue exists
because the tradeoff is currently discovered rather than chosen.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions