Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Fault Injection

The serve mode can kill a streaming chat completion mid-generation in every way a real downstream does — on demand, deterministically, per request. It exists as the test double for gateway resilience work (mid-stream continuation/resume middleware): if the sim can produce every death signature, that middleware can be e2e-tested without waiting for real incidents.

Every faulting stream first emits real partial output — the initial role frame plus after_chunks content-bearing delta frames of deterministic placeholder text — so there is always something to resume. The fault path bypasses the simulation engine (like echo-directives) and contains no randomness: the same trigger produces the same frames and the same death at the same byte, every time.

Scope: streaming POST /v1/chat/completions and streaming POST /v1/completions. A fault header on a non-streaming request is rejected with a 400 (invalid_fault_directive) — never silently ignored. With no trigger present, behavior is completely unchanged.

Both endpoints matter because a mid-stream-continuation resume leg is a streaming /v1/completions request, so chain-resume tests need to kill one. The frames follow the endpoint’s own wire shape: /v1/completions emits text_completion chunks carrying text and has no role frame, so after_chunks=N puts exactly N frames on the wire before the death (one fewer than the chat flavor, which leads with its role frame). mid_reasoning and mid_tool_call have no delta object to use there, so they stream the same partial payloads as raw text — an unterminated <think> block and a tool call the model never finished writing — which is how both actually appear on a real base-model completions stream.

Trigger: the x-inference-lab-fault header

The header is only honored when the server runs with --enable-directives — it is client-controlled and can stall or abort connections at will, so it sits behind the same “untrusted clients must not reach this server” gate as echo-directives. Without the flag the header is a 400 (invalid_fault_directive), never silently ignored. The staging deployment already sets the flag.

x-inference-lab-fault: <mode>[;after_chunks=<u32>][;delay_ms=<u64>][;utf8=<bool>]
ParameterDefaultMeaning
<mode>requiredone of the eleven mode names below
after_chunks3content-bearing delta frames emitted before the fault fires (the initial role-only frame is always sent and not counted)
delay_ms10fixed pacing between frames, milliseconds
utf8falsecut_mid_frame only: cut inside a multi-byte UTF-8 character

Parts are ;-separated; whitespace around parts and = values is ignored. Unknown modes, unknown keys, or malformed values are a 400 listing the valid modes.

A header was chosen over a body extension because the platform’s proxy layer runs strict request sanitization that strips unknown body fields; headers pass through. Verify this empirically through the full stack (client → dwctl → onwards → sim) when the branch reaches a preview environment — full-platform verification is out of scope here and happens after deploy.

Fallback trigger: static per-model config

For clients that cannot set a header, a model’s TOML config can apply one fault to every streaming chat completion on that model (non-streaming requests are served normally; an explicit header on the request still wins). Unlike the header, this is operator input validated at server boot, so it does not require --enable-directives:

[fault]
mode = "cut_mid_frame"   # same names as the header
after_chunks = 5         # optional, default 3
delay_ms = 10            # optional, default 10
utf8 = true              # optional, cut_mid_frame only

An invalid [fault] block fails server startup, not individual requests.

Precedence per request: header > model [fault] config > (echo-directives >) normal path.

Modes

ModeAfter the N content frames…Client observes (curl)
cut_between_framesconnection closes on a frame boundary, without the chunked-encoding terminator (FIN)exit 18, transfer closed with outstanding read data
cut_mid_framehalf of the next frame’s bytes, then close — torn JSON. utf8=true cuts one byte into a 2-byte UTF-8 character (é) in the delta textexit 18, partial data: line
resetabortive close: SO_LINGER=0 then drop → TCP RST, not FINexit 56, connection reset by peer
stallnothing, forever; connection stays open until the client gives upexit 28 (client timeout)
error_envelope_200OpenRouter-style error envelope as an SSE data frame, then [DONE]; HTTP status stays 200exit 0, {"error":{"message":…,"code":502,"metadata":{…}}}
error_400_in_ssevLLM-style 400 object in-stream (the nemotron-incident signature), then clean close, no [DONE]exit 0, {"object":"error",…,"code":400}
no_donefinish_reason frame, then the separate choices: [] usage frame if requested, then clean close — [DONE] never comesexit 0, stream just ends
no_usagefinish_reason frame then [DONE], but the usage frame never arrives (even when stream_options.include_usage was set)exit 0, usage missing
cancelled_499the exact dynamo frontend-cancellation body, then clean closeexit 0, {"error":{"code":499,"message":"CancelledError: ","type":"request_cancelled"}}
mid_reasoningframes carry reasoning_content deltas instead of content; dies cut_between_frames-styleexit 18, last deltas are reasoning
mid_tool_callframe 1 announces a tool call (id + name), later frames stream arguments fragments that never terminate; dies cut_between_frames-styleexit 18, partial tool call

Notes:

  • The error body shapes for error_envelope_200 and error_400_in_sse are representative; exact shapes sync with the death-taxonomy workstream as it lands. The cancelled_499 body is exact.
  • delay_ms=0 is fine for the graceful modes, but the abrupt modes always wait a short flush grace (~25 ms) before killing the connection so the partial output reliably reaches the wire first.
  • reset needs the raw socket, which the server threads through per-connection; on non-unix platforms (or when handlers are driven outside the real server, as in unit tests) it degrades to a FIN with a warning log.

Examples

All against a local sim (inference-lab serve --config configs/ --hardware b200 --port 8080); $BODY is any streaming chat request:

BODY='{"model":"DeepSeek-V4-Flash","stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hello"}],"max_tokens":16}'
URL=http://localhost:8080/v1/chat/completions
# 1. cut_between_frames — 5 frames then FIN (curl exit 18)
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cut_between_frames;after_chunks=5' -d "$BODY"

# 2a. cut_mid_frame — torn JSON frame
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cut_mid_frame' -d "$BODY"

# 2b. cut_mid_frame, cut inside a multi-byte UTF-8 character (pipe through xxd to see it)
curl -sN $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cut_mid_frame;utf8=true' -d "$BODY" | xxd | tail

# 3. reset — TCP RST (curl exit 56, "connection reset by peer")
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: reset' -d "$BODY"

# 4. stall — 3 frames then silence; bound the wait client-side
curl -N --max-time 10 $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: stall' -d "$BODY"

# 5. error_envelope_200 — OpenRouter-style envelope then [DONE]
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: error_envelope_200' -d "$BODY"

# 6. error_400_in_sse — 400-shaped object inside the 200 stream
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: error_400_in_sse' -d "$BODY"

# 7. no_done — finish_reason + usage frame, then the stream ends without [DONE]
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: no_done' -d "$BODY"

# 8. no_usage — [DONE] arrives but the requested usage never does
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: no_usage' -d "$BODY"

# 9. cancelled_499 — exact dynamo cancellation body
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cancelled_499' -d "$BODY"

# 10. mid_reasoning — dies while streaming reasoning_content deltas
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: mid_reasoning' -d "$BODY"

# 11. mid_tool_call — dies with a tool call's arguments unterminated
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: mid_tool_call;after_chunks=4' -d "$BODY"

Killing a resume leg

A resume leg is a streaming /v1/completions request whose prompt is token ids. Every mode above works against it; only the frame shape differs (no role frame, text instead of a delta):

RESUME='{"model":"DeepSeek-V4-Flash","prompt":[1,2,3,4],"stream":true,"priority":0,"stream_options":{"include_usage":true},"max_tokens":64}'

# exactly 2 text frames, then FIN (curl exit 18)
curl -N http://localhost:8080/v1/completions -H 'content-type: application/json' \
  -H 'x-inference-lab-fault: cut_between_frames;after_chunks=2' -d "$RESUME"