Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Introduction

Inference Lab is a simulation framework designed to evaluate and analyze LLM workloads.

It uses discrete-event simulation to model the behavior of a multi-GPU node serving LLM inference requests with the vLLM library. It contains a facsimile of the vLLM queueing, scheduling, and execution logic, with only the actual model inference replaced by a performance model based on the supplied GPU specs and model architecture.

Within each simulation step, the simulator:

  • Processes any newly arrived requests, adding them to the scheduling queue.
  • Schedules requests to serve based on the selected scheduling policy.
  • Calculates the compute and memory bandwidth usage for the workload that the scheduled requests represent, and the theoretical time required to execute the workload on the specified hardware.
  • Increments the simulation time by the calculated execution time, updating the state of all requests accordingly.

Caveats:

  • Step times are a datasheet roofline: peak FLOP rate and HBM bandwidth per precision stream, collectives on the fabric preset added serially, and no kernel-efficiency or fixed per-step overhead term. Every latency and throughput is an upper bound; the optional time_correction = { alpha, beta } on a hardware entry calibrates the step (alpha × roofline + beta) against a measured engine.

Features

  • Roofline Performance Modeling: compute (FLOPS) and memory bandwidth constraints per precision stream
  • Multiple Scheduling Policies: FCFS, Priority, SJF, and more
  • Chunked Prefill: Simulates realistic request interleaving
  • KV Cache Management: Models GPU memory and KV cache utilization
  • Workload Generation: Supports Poisson, Gamma, and closed-loop patterns
  • WebAssembly Support: Run simulations in the browser via WASM

Quick Start

See the Getting Started guide to begin using Inference Lab.

Getting Started

This guide will help you get started with Inference Lab.

Installation

Install from crates.io:

cargo install --locked inference-lab

Or build from source:

cargo build --release

Running Your First Simulation

From a checkout of the repository:

inference-lab --config configs/llama-3-70b.toml --workload workloads/quick.toml

configs/ holds one file per model (each with its hardware entries) and workloads/ the arrival patterns and request shapes; pass --hardware <name> when a model config has more than one entry.

Next Steps

Running Simulations

This guide covers how to run simulations and interpret results.

Basic Usage

Run a model config on one of its hardware entries against a workload:

inference-lab -c configs/gpt-oss-120b.toml --hardware b200 -w workloads/chat-closed-256.toml

For dataset mode, add tokenizer and chat template:

inference-lab -c configs/llama-3-70b.toml -w workloads/dataset-poisson.toml \
  --tokenizer tokenizer.json \
  --chat-template None

See Configuration for details on configuring workloads, policies, and hardware.

Output Modes

Console Output (Default)

By default, the simulator displays:

  • Real-time progress bar
  • Current simulation time
  • Queue status (running/waiting requests)
  • KV cache utilization

Final output includes:

  • Latency metrics (TTFT, E2E, per-token)
  • Throughput metrics (tokens/sec, requests/sec)
  • Utilization statistics (KV cache, FLOPS, bandwidth)
  • Preemption statistics

JSON Output

Save results to a file:

inference-lab -c configs/gpt-oss-120b.toml --hardware b200 -w workloads/chat-closed-256.toml -o results.json

Combine with -q for batch processing:

inference-lab -c configs/gpt-oss-120b.toml --hardware b200 -w workloads/chat-closed-256.toml -q -o results.json

Running Multiple Experiments

Comparing Hardware

for hw in b200 b300 gh200; do
  inference-lab -c configs/gpt-oss-120b.toml --hardware $hw \
    -w workloads/chat-closed-256.toml -q -o results_$hw.json
done

Sweeping Engine Args

for batch_size in 4096 8192 16384; do
  sed "s/max_num_batched_tokens = .*/max_num_batched_tokens = $batch_size/" \
    configs/gpt-oss-120b.toml > /tmp/gpt-oss-120b_$batch_size.toml
  inference-lab -c /tmp/gpt-oss-120b_$batch_size.toml --hardware b200 \
    -w workloads/chat-closed-256.toml -o results_$batch_size.json
done

Multiple Seeds

Override the workload’s seed:

for seed in {1..10}; do
  inference-lab -c configs/gpt-oss-120b.toml --hardware b200 \
    -w workloads/chat-closed-256.toml --seed $seed -q -o results_$seed.json
done

Understanding Results

Latency Metrics

Time to First Token (TTFT)

  • Time from request arrival to first token generation
  • Lower is better for interactive applications
  • Affected by: queue wait time, prefill computation

End-to-End (E2E) Latency

  • Total time from request arrival to completion
  • Includes prefill and all decode steps
  • Key metric for overall user experience

Per-Token Latency

  • Average time between consecutive output tokens
  • Lower is better for streaming applications
  • Primarily affected by batch size and model size

Throughput Metrics

Input Tokens/sec

  • Rate of processing prompt tokens
  • Indicates prefill throughput

Output Tokens/sec

  • Rate of generating output tokens
  • Indicates decode throughput

Requests/sec

  • Overall request completion rate
  • Key metric for capacity planning

Utilization Metrics

KV Cache

  • Percentage of KV cache memory in use
  • High utilization may lead to preemptions

FLOPS

  • Percentage of compute capacity utilized
  • Low FLOPS may indicate memory bottleneck

Bandwidth

  • Percentage of memory bandwidth utilized
  • High bandwidth utilization indicates memory-bound workload

Preemption Statistics

Preemptions occur when new requests need memory but the KV cache is full:

  • Total number of preemptions
  • Average preemptions per request
  • Can significantly impact TTFT for preempted requests

Troubleshooting

Simulation running slowly?

  • Reduce num_requests or use -q flag
  • Build with --release

Too many preemptions?

  • Raise gpu_memory_utilization or set kv_cache_capacity in [scheduler]
  • Reduce max_num_seqs or max_num_batched_tokens in scheduler config

Dataset loading errors?

  • Verify --tokenizer and --chat-template flags are provided
  • Check JSONL format matches OpenAI batch API format

For more details, see CLI Reference and Configuration.

Configuration

A simulation is a model config × one of its hardware entries × a workload. Model configs live in configs/, one file per model deployment; workloads in workloads/. Unknown fields in either are rejected.

inference-lab --config configs/qwen3.6-35b-a3b-fp8.toml --hardware b200 \
              --workload workloads/chat-closed-256.toml

Model config

model = "qwen3.6-35b-a3b-fp8"   # catalog preset, or an inline [model] table

[scheduler]                     # engine args shared by every hardware entry
max_num_batched_tokens = 16384
max_num_seqs = 4096
policy = "priority"
enable_chunked_prefill = true
block_size = 64

[hardware.b200]                 # one entry per hardware this model runs on
tp = 1

[hardware.b300]
tp = 1

[hardware.gh200]
tp = 1
scheduler = { max_num_batched_tokens = 8192 }   # per-entry override
  • model — a catalog preset name or an inline [model] table: weight streams and token-mixing layer classes.
  • [scheduler] — engine arguments: batching, KV blocks, memory utilisation, scheduling policy.
  • [hardware.<name>] — one per hardware the model is deployed on. The name is a hardware preset (b200, b300, gh200, h100) unless the entry sets spec. Each entry gives the parallel layout — tp (replica world size, weights sharded, per-layer all-reduces), ep (experts sharded over ep ranks), dp_attention (data-parallel attention: replicated attention weights, all-gather/reduce-scatter around the FFN, or dispatch/combine all-to-alls when ep > 1) — and may override scheduler keys or carry its own speculative block. The shipped configs carry the layouts production runs (tp8 + dp_attention for DeepSeek-V4-Pro / Kimi / GLM-5, tp = ep for the Qwen3.5/VL MoEs and Nemotron Ultra, plain tp elsewhere).
  • [speculative] — speculative decoding, optional; a shared default that an entry’s speculative replaces (acceptance traces and measured step costs are per hardware).
  • [memory] — KV tiers beyond HBM, picked from the stores the hardware offers (host DRAM over PCIe, Grace memory over NVLink-C2C, NVMe): evicted blocks fall through them and are promoted back instead of recomputed. A per = "node" store is shared by the workers on a node. How KV moves — fetch or recompute, what HBM and each tier evict, when blocks are written, whether a re-entry’s prefix is prefetched — is a set of policies with two presets, reactive (decides from the past, like shipped stacks) and oracle (knows every session’s next re-entry); see the reference.
  • replicas / [router] / [decode_router] — an entry’s replicas (default 1) runs that many identical workers, each with its own scheduler and KV cache; the shared [router] (or an entry’s router) picks which one each request enters: round_robin, least_loaded, prefix_affinity, or kv_aware. On a disaggregated topology [decode_router] (default: [router]) picks the decode worker each hand-off goes to; kv_aware_decode prices the transfer and the decode batch (see the reference).

--hardware picks the entry; it can be omitted when a file has one. inference-lab serve --config configs/ --hardware b200 serves every model with a b200 entry.

Hardware

Name a shipped preset as the entry:

[hardware.b200]
tp = 2

or point an entry at another preset, or at an inline per-GPU spec (a FLOP rate for every precision the model uses, bandwidth, capacity, optional [fabric] and [memory]):

[hardware.isambard]
spec = "gh200"
tp = 4

[hardware.custom]
tp = 1
[hardware.custom.spec]
name = "H100"
flops_fp8 = 1.979e15            # dense FLOP/s at fp8
flops_bf16 = 9.895e14           # dense FLOP/s at bf16
memory_bandwidth = 3.35e12      # bytes/sec
memory_capacity = 85899345920   # 80 GB

How much of that memory the engine may use, and how much goes to KV, are deployment settings and live in [scheduler] (gpu_memory_utilization, kv_cache_capacity).

A preset also carries its node’s collective fabric — gpus_per_node, scale_up (NVLink: bandwidth, latency, in-network reduction) and scale_out (per-GPU NIC across nodes) — which prices the TP all-reduces and EP all-to-alls of any entry with tp > 1 or ep > 1. An inline spec that omits [fabric] can only be used with tp = 1, ep = 1.

Model

Name a shipped preset:

model = "gemma-4-31b-it"

or describe the architecture inline as weight streams plus layer classes:

[model]
name = "Llama-3-70B"
hidden_dim = 8192
max_seq_len = 8192
attention_precision = "fp8"

[[model.weights]]               # one per-token GEMM stream per precision
precision = "fp8"
active_params = 70000000000
resident_params = 70000000000

[[model.layers]]                # token-mixing layer classes
kind = "attention"              # or "mla", "linear"
count = 80
heads = 64
head_dim = 128
kv_heads = 8                    # GQA
kv_precision = "fp8"

MoE adds routing = { routed_experts, experts_per_tok, moe_layers } to the expert stream; sliding-window layers are an attention class with window; MLA / DeepSeek sparse attention is the mla kind; GatedDeltaNet or Mamba layers are linear with their per-sequence state_bytes. See the Configuration Reference for every field.

Scheduler

Control request scheduling and batching:

[scheduler]
max_num_batched_tokens = 8192
max_num_seqs = 256
policy = "fcfs"
enable_chunked_prefill = true
block_size = 16

Scheduling Policies

Available policies:

  • fcfs - First-Come-First-Served (default)
  • sof - Shortest Output First
  • sif - Shortest Input First
  • stf - Shortest Total First
  • lif - Longest Input First
  • lof - Longest Output First
  • ltf - Longest Total First

Chunked Prefill

Enable chunked prefill to allow interleaving prompt processing with generation:

enable_chunked_prefill = true
long_prefill_token_threshold = 512  # Optional: chunk size limit
max_num_partial_prefills = 1        # Max concurrent partial prefills

Preemption-Free Mode

Enable conservative admission control to guarantee zero preemptions:

enable_preemption_free = true

Workload

A workload file is the workload table at top level: how requests arrive and their shapes.

Synthetic Workload

# workloads/chat-poisson-5rps.toml
arrival_pattern = "poisson"
arrival_rate = 5.0
num_requests = 100
seed = 42

[input_len_dist]
type = "lognormal"
mean = 6.9
std_dev = 0.7

[output_len_dist]
type = "lognormal"
mean = 5.3
std_dev = 0.8

Arrival Patterns

  • poisson - Poisson process with exponential inter-arrival times
  • uniform - Uniform random inter-arrival times
  • burst - Bursty traffic
  • fixed_rate - Fixed interval between requests
  • closed_loop - Fixed number of concurrent users
  • batched - Requests arrive in batches

Length Distributions

Four distribution types are supported:

Fixed:

[input_len_dist]
type = "fixed"
value = 1000

Uniform:

[input_len_dist]
type = "uniform"
min = 100
max = 2000

Normal:

[input_len_dist]
type = "normal"
mean = 1000.0
std_dev = 200.0

LogNormal:

[input_len_dist]
type = "lognormal"
mean = 6.9      # ln(1000)
std_dev = 0.7

Dataset Mode

Use real request traces instead of synthetic workloads:

dataset_path = "path/to/dataset.jsonl"
arrival_pattern = "poisson"
arrival_rate = 1.0

# These are used for sampling actual generation length
input_len_dist = { type = "fixed", value = 100 }  # Ignored
output_len_dist = { type = "fixed", value = 50 }  # Samples EOS

Dataset Format: JSONL file in OpenAI batch API format. Each line may target either /v1/chat/completions with a messages array or /v1/completions with a string prompt.

Example:

{"custom_id": "req-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-3.5-turbo", "messages": [{"role": "user", "content": "Hello!"}], "max_tokens": 100}}
{"custom_id": "req-2", "method": "POST", "url": "/v1/completions", "body": {"model": "gpt-3.5-turbo-instruct", "prompt": "Write a haiku about Rust.", "max_tokens": 80}}

Tokenizer: Dataset mode requires a tokenizer file to convert text to tokens. You’ll need to provide this via the --tokenizer flag:

inference-lab -c configs/llama-3-70b.toml -w workloads/dataset-poisson.toml --tokenizer tokenizer.json

The tokenizer should be a HuggingFace tokenizers JSON file (typically tokenizer.json from the model repository).

Chat Template: You’ll also need to specify how to format chat-style requests via --chat-template:

  • Use "None" for simple concatenation of messages
  • Use a Jinja2 template string for custom formatting (e.g., "{{user}}\n{{assistant}}")
  • Most models have their own chat template format
  • Plain /v1/completions prompts are tokenized directly and do not use the chat template

Example with no template:

inference-lab -c configs/llama-3-70b.toml -w workloads/dataset-poisson.toml \
  --tokenizer tokenizer.json \
  --chat-template None

Sessions

Agentic traffic: chains of requests where each step re-enters with its parent’s whole context as prefix (prompt and the parent’s output) plus some novel tokens, after the gap the harness spent between the parent’s completion and this arrival (a tool call running, a user typing).

sessions_path = "data/sessions/tracelab.jsonl"
arrival_pattern = "poisson"   # governs session starts
arrival_rate = 0.05           # sessions/s
num_sessions = 200            # stop starting sessions after this many
# Optional: seed the first live population part-way through its traces.
stationary_start_sessions = 128
# Optional: draw sessions uniformly with replacement instead of in file order.
resample_sessions = true
seed = 42

input_len_dist = { type = "fixed", value = 1 }   # ignored in session mode
output_len_dist = { type = "fixed", value = 1 }  # ignored in session mode

The arrival pattern decides when sessions start: poisson / uniform / burst at arrival_rate sessions per second, closed_loop keeps num_concurrent_users sessions in flight (a slot starts a fresh session when its session’s last step completes), batched starts every session at t=0. Every later step of a session arrives at its parent’s completion plus the step’s gap, so the simulated latency feeds back into the arrival process and long gaps are preserved. By default sessions are taken from the file in order and the file is cycled. num_sessions bounds session starts; num_requests separately bounds total emitted request steps across all sessions.

stationary_start_sessions avoids waiting a session-lifetime tail for an open-loop live population to reach stationarity. All N seeded sessions enter at t=0; the arrival clock is held until they have started, then resumes normally. Each starts at a trace step sampled in proportion to that step’s gap (the time the session waits before issuing it). Its inherited context receives fresh block hashes, so the first emitted request reports its shared prefix but must prefill that context once; later sessions start at step 0 normally.

resample_sessions = true draws every new session uniformly from the file with replacement. It preserves each sampled session’s within-trace structure while avoiding file-order effects. Repeated instances receive fresh block hashes and therefore do not share cache state.

Session file: JSONL, one session per line:

{"id": "claude:000adcd5", "steps": [
  {"input": 15524, "new": 15524, "output": 111, "gap": 0.0, "kind": "user"},
  {"input": 17079, "new": 1444, "output": 96, "gap": 0.104, "kind": "tool"}
]}

input is the step’s prompt length, new the tokens of it that are not the parent’s context (input − new is the reusable prefix, capped at what the parent actually had), output the tokens generated, gap the seconds from the parent’s completion to this arrival (ignored on the first step), kind free-form. Prefix identity is built from block hashes: the parent’s hashes over the shared prefix, fresh ones for the novel tail and the step’s own output, so a re-entry hits the parent’s generated tokens too. Whole blocks only: a partial block continued with new tokens is new content.

examples/sessions/tracelab_export.py exports TraceLab’s per_step_stats.parquet into this format.

Per-request CSV (--request-csv) carries, for session steps, session, step, worker (the memory-graph id of the worker that served it), gap, shared_toks (the most the prefix cache could serve), cached_toks (what it did), and two reuse distances: reuse_distance_bytes (KV bytes written into the caches between the parent’s completion and this arrival) and reuse_touched_bytes (the same plus the free blocks that hits pulled back into use in between). Fresh writes undercount the LRU stack distance and the touched count overcounts it, so the pair brackets it.

Closed-Loop Workload

Simulate a fixed number of concurrent users:

arrival_pattern = "closed_loop"
num_concurrent_users = 256
closed_loop_jitter_secs = 0.05  # stagger the initial arrivals
# ... length distributions ...

Common Configuration Patterns

High Throughput Setup

Maximize batch size and token throughput:

[scheduler]
max_num_batched_tokens = 16384
max_num_seqs = 512
enable_chunked_prefill = true

Low Latency Setup

Prioritize request completion speed:

[scheduler]
max_num_batched_tokens = 4096
max_num_seqs = 64
policy = "sof"  # Shortest Output First

Memory-Constrained Setup

Limit KV cache usage:

[scheduler]
kv_cache_capacity = 34359738368  # 32 GB explicit limit
max_num_seqs = 128

Next Steps

CLI Reference

Command-line interface reference for Inference Lab.

Usage

inference-lab [OPTIONS]

A binary built with --features serve (the Docker image) has subcommands instead: inference-lab sim [OPTIONS] takes the options below and inference-lab serve starts the OpenAI-compatible server. serve takes --config (a model config or a directory of them), --hardware (models without that entry are skipped) and an optional --workload, whose output_len_dist samples each response’s length; without one responses run to their max_tokens. --max-waiting <N> bounds the waiting queue so the server refuses arrivals past it with HTTP 529 (0, the default, queues without limit); it and max_num_seqs are retunable on a running server through POST /control/capacity — see Saturation and Capacity.

Options

Configuration

-c, --config <PATH>

Model config file (configs/<name>.toml).

  • Default: config.toml

--hardware <NAME>

Which [hardware.<name>] entry of the model config to run. Optional when the file has exactly one entry.

-w, --workload <PATH>

Workload file (workloads/<name>.toml). Required for sim.

inference-lab -c configs/gpt-oss-120b.toml --hardware gh200 -w workloads/quick.toml

Dataset Mode

-t, --tokenizer <PATH>

Path to tokenizer file (required for dataset mode).

  • Required when using dataset_path in configuration
  • Example: inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --tokenizer tokenizer.json

--chat-template <TEMPLATE>

Chat template for formatting messages in dataset mode.

  • Required when using datasets
  • Use "None" for simple message concatenation (no template)
  • Example: inference-lab -c configs/llama-3-70b.toml -w workloads/dataset-poisson.toml --tokenizer tokenizer.json --chat-template None
  • Example with template: inference-lab ... --tokenizer tokenizer.json --chat-template "{{system}}\n{{user}}\n{{assistant}}"

Output Options

-o, --output <PATH>

Path to output JSON file for results.

  • If not specified, results are only displayed to console
  • Example: inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -o results.json

-q, --quiet

Suppress progress output (only show final results).

  • Example: inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -q

-v, --verbose

Enable verbose output.

  • Example: inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -v

--debug

Enable debug logging.

  • Example: inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --debug

--no-color

Disable colored output.

  • Useful for logging to files or CI environments
  • Example: inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --no-color

Simulation Options

--seed <NUMBER>

Override the random seed from configuration.

  • Useful for reproducible runs with different seeds
  • Example: inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --seed 12345

Examples

Basic Simulation

inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml

Dataset Mode

inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml \
  --tokenizer tokenizer.json \
  --chat-template None

Save Results to File

inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -o results.json

Quiet Mode with Output

inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -q -o results.json

Multiple Runs with Different Seeds

for seed in 42 43 44; do
  inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --seed $seed -o results_$seed.json
done

Exit Codes

  • 0 - Simulation completed successfully
  • 1 - Error occurred (configuration error, file not found, etc.)

Configuration File Reference

Field-by-field reference for the two TOML files a simulation takes: a model config (configs/<name>.toml) and a workload (workloads/<name>.toml). Unknown fields anywhere in either file are rejected.

# configs/<name>.toml
model = "deepseek-v4-flash"    # catalog preset, or an inline [model] table

[scheduler]                    # engine args shared by every hardware entry
[speculative]                  # optional: speculative decoding, shared default
[router]                       # optional: how requests spread over replicas
[decode_router]                # optional: disagg decode pool (defaults to [router])
[memory]                       # optional: KV tiers beyond HBM, from the hardware's stores
[prefill]                      # optional: a prefill pool; the entry then decodes (disaggregated)
[fault]                        # optional: serve-mode static fault injection

[hardware.b200]                # one entry per hardware the model runs on
tp = 4
replicas = 4                   # identical workers behind the router
[hardware.gh200]
tp = 4
scheduler = { max_num_batched_tokens = 4096 }
# workloads/<name>.toml — the [workload] table at top level
arrival_pattern = "closed_loop"
num_concurrent_users = 256
num_requests = 2000
seed = 7
input_len_dist = { type = "lognormal", mean = 7.0, std_dev = 0.5 }
output_len_dist = { type = "lognormal", mean = 6.5, std_dev = 0.8 }

inference-lab --config <model config> --hardware <entry> --workload <workload> runs one; --hardware may be omitted when the file has one entry. In Rust, ModelConfig::from_file(..).deployment(Some("b200")) gives a Deployment (model on hardware) and .with_workload(WorkloadConfig::from_file(..)) a Config. The WASM API takes that resolved Config as JSON: hardware, parallel, model, scheduler, workload, optional replicas, router, decode_router, memory, prefill, speculative.

Catalog presets

Hardware and model presets ship inside the crate (catalog/hardware/*.toml, catalog/models/*.toml, embedded at build time). A config names one with model = "<name>" and [hardware.<name>]; inference_lab::catalog:: {hardware_names, model_names, hardware, model} list and load them from Rust. Presets carry only the physical hardware spec and the model architecture; everything about a deployment (TP, memory utilisation, batch limits) is in the config. To tweak a preset, copy its table inline.

[hardware.<name>]

One entry per hardware the model is deployed on. The entry name is the hardware preset unless spec says otherwise.

FieldTypeDefaultDescription
tpU321Replica world size: its GPUs pool FLOP rate, HBM bandwidth and memory; weights are sharded across them; each layer’s output is all-reduced twice (after attention, after the FFN)
epU321Experts sharded across ep of the ranks (divides tp). With TP attention every rank holds every token, so the MoE output is still combined by the FFN all-reduce (vLLM --enable-expert-parallel); under dp_attention the MoE layers dispatch + combine with all-to-alls over the ep group instead
dp_attentionBoolfalseAttention runs data-parallel over the tp ranks (sglang --enable-dp-attention): the attention projections are replicated (tp× resident and read per step), a sequence’s KV lives on one rank, no attention all-reduce; the TP-sharded FFN gathers the ranks’ tokens with an all-gather and returns them with a reduce-scatter (with ep > 1, DeepEP-style dispatch + combine all-to-alls). Each rank is its own worker: its own scheduler and KV cache over its GPU’s HBM (1/tp of the replica’s), its own port into the node’s [memory] stores; [scheduler] limits apply per rank. The [router] routes flat over ranks — the placement decision a DP-aware front end (Dynamo’s (worker, dp_rank) endpoints) makes. The ranks step in lockstep as one group: the step is priced on the replica-wide roofline over the union batch plus the attention skew, the slowest rank’s own-GPU attention time over the mean rank’s (the ranks meet at every layer’s FFN collective). Speculative decoding is not modelled with rank groups
replicasU321Identical workers of this deployment (each a tp-GPU replica with its own scheduler and KV cache; tp rank workers each under dp_attention) behind the router
specString or Tablethe entry nameAnother catalog preset, or an inline hardware table (fields below)
schedulerTable{}Keys merged over the shared [scheduler] for this entry
speculativeTableshared [speculative]Replaces the shared block for this entry
routerTableshared [router]Replaces the shared block for this entry
decode_routerTableshared [decode_router]Replaces the shared block for this entry
memoryTableshared [memory]Replaces the shared block for this entry
prefillTableshared [prefill]Replaces the shared block for this entry
time_correctionTableunsetStep-time calibration against a measured engine: { alpha = 1.0, beta = 0.0 } prices every roofline step at alpha × t + beta seconds (alpha = kernel-efficiency gap, beta = fixed per-iteration cost). Unset = pure roofline. Never applied on top of a measured step-cost table

Collectives are priced on the hardware’s [fabric] (below) and added serially to the step; an entry with tp > 1 or ep > 1 on hardware without one is rejected. Per layer: attention → one all-reduce over tp (none under dp_attention); dense FFN, or MoE without dp_attention (any ep) → one all-reduce; under dp_attention, dense FFN or MoE with ep = 1 → an all-gather + reduce-scatter, MoE with ep > 1 → dispatch and combine all-to-alls over ep, each rank moving its tokens / ep share of experts_per_tok hidden vectors. Expert reads and FLOPs are taken as balanced across ranks; there is no overlap of collectives with compute.

Hardware spec

Per-GPU physical spec: a catalog preset (catalog/hardware/<name>.toml) or an inline spec table.

FieldTypeDefaultDescription
nameString—Accelerator name
flops_fp4FloatunsetDense FLOP/s at FP4. Unset means the hardware has no FP4 rate; a model that declares an FP4 stream then fails at run time
flops_fp8FloatunsetDense FLOP/s at FP8
flops_bf16FloatunsetDense FLOP/s at BF16
flops_fp16FloatunsetDense FLOP/s at FP16 (FP32 is taken as half of it)
memory_bandwidthFloat—HBM bandwidth, bytes/s
memory_capacityU64—HBM capacity, bytes
fabricTableunsetCollective fabric, see below
memoryTableunsetKV memory beyond HBM this node class offers (stores and links), see below. A deployment picks tiers from it with its own [memory]

[fabric]

FieldTypeDefaultDescription
gpus_per_nodeU32—GPUs sharing one scale-up domain; parallel groups are packed node by node
scale_upTable—Inside a node (NVLink / NVSwitch): { bandwidth, latency, in_network_reduction } — per-GPU injection bytes/s per direction, seconds per collective call, and whether the switch reduces in-network (NVLink SHARP)
scale_outTableunsetAcross nodes, rail-optimised (GPU i drives NIC i): same fields. Required for a group wider than gpus_per_node

Cost per collective, added serially to the step (no overlap): a TP all-reduce of V bytes over g ≤ gpus_per_node ranks is latency + f·V / bandwidth with f = 2(g−1)/g (ring) or 1 (in-network reduction); over more ranks it is reduce-scatter and all-gather inside each node around an all-reduce of the V/k shard across the n nodes on the NIC. An EP all-to-all moves each rank’s (g−1)/g share at the scale-up rate, or, across nodes, its in-node and cross-node shares concurrently on their own links. dp_attention skips the per-layer all-reduce.

[memory] on the hardware

The stores a node offers for KV beyond HBM and the links that reach them, templated per GPU or per node: a graph. Presets ship one (host DRAM and a node NVMe pool behind each GPU’s PCIe port on the HGX parts; Grace LPDDR5X behind NVLink-C2C on gh200; NVLink ports into the node switch and one NIC per GPU to the network on all); an inline hardware table can declare its own.

FieldTypeDefaultDescription
gpus_per_nodeU32fabric’s, else 1GPUs sharing a node’s per = "node" stores
storesArray[]{ name, per = "gpu" | "node" | "cluster", kind = "store", capacity, bandwidth, stripe = 1, aggregate_bandwidth, latency = 0, pin = false } — bytes per instance; a gpu store is private to its GPU, a node store is one pool for the node’s GPUs, and a cluster store is one topology-wide pool behind the network. bandwidth is the store throughput per instance (per node for a cluster store). A cluster transfer can use stripe × bandwidth, while all its transfers share aggregate_bandwidth or, by default, nodes × bandwidth. latency is the fixed access cost in seconds on every fetch or write. The reserved { name = "peer_hbm", per = "node", kind = "peer_hbm" } entry is virtual: no capacity/bandwidth, read-only, and pin defaults to true
junctionsArray[]{ name, per } — a point with no capacity of its own, so that several links can share one (a GPU’s PCIe port feeding host DRAM and NVMe)
linksArray[]{ name, from, to, bandwidth, latency = 0 } — from is "gpu" (one port per GPU), "network", a store or a junction; to is a store, a junction, "switch" (the node’s scale-up fabric) or "network" (the scale-out core). One instance per instance of from, full duplex at bandwidth bytes/s each way

Instances: a tp-GPU worker pools its GPUs’ per-GPU stores, junctions and ports (capacity and bandwidth × tp); gpus_per_node / tp workers share a node’s per-node instances; every worker points at the same selected cluster store. A cluster access runs GPU → its NIC → network → store (and back), so the NIC remains the per-worker cap even when the store is striped. A transfer takes the shortest hop path between its ends and runs at its max-min fair share on every edge of it: the most contended edge fixes its transfers’ rate first, the residual is shared among the rest. Tier promotions run store → GPU; peer_hbm promotions run from the selected sibling GPU through the node switch to the destination GPU; hand-offs run GPU → network → GPU (see kv_link_bw).

Shipped presets, at datasheet figures: b200 (192 GB / 8 TB/s), b300 (288 GB / 8 TB/s), gh200 (96 GB / 4 TB/s), h100 (80 GB / 3.35 TB/s); each carries its node’s fabric (8-GPU NVSwitch + CX-7/CX-8 for the HGX boxes, 4-GPU NVLink + Slingshot for GH200).


[model]

A model is described by its per-token weight streams and its token-mixing layer classes; there are no named architectures. Any transformer the simulator serves is a composition of the pieces below.

FieldTypeDefaultDescription
nameString—
hidden_dimU32—Residual width (sizes collectives; prices MLA attention when its head shape is not given)
max_seq_lenU32—Architecture context limit (max_position_embeddings)
attention_precisionPrecision"bf16"Rate the attention score / AV matmuls run at; KV reads charge this stream
activation_bytesU322Bytes per activation element on the wire
weightsArray—One or more weight streams (below)
layersArray—One or more layer classes (below)

[[weights]] — a per-token GEMM stream at one precision

FieldTypeDefaultDescription
precision"fp4","fp8","bf16","fp16","fp32"—
active_paramsU64—Parameters touched per token (FLOPs = 2×)
resident_paramsU64—Parameters resident in HBM
routingTableunsetMoE routing: { routed_experts, experts_per_tok, moe_layers }. The per-step read then follows coupon-collector growth with the step’s tokens (per-expert and shared params are recovered from the active/resident split); EP all-to-alls (under dp_attention) = 2 × moe_layers

A dense fp8 model is one stream; DeepSeek-V4 is an fp4 expert stream with routing plus an fp8 non-expert stream; gpt-oss is fp4 experts + bf16 rest.

[[layers]] — a class of identical token-mixing layers

kind = "attention" — GQA / MHA with a growing KV cache:

FieldTypeDefaultDescription
countU32—Layers in the class
heads, head_dimU32—Query heads and head width (attention FLOPs = 4 × heads × head_dim per query-key pair)
kv_headsU32—KV heads (KV per token = 2 × kv_heads × head_dim × bytes)
kv_sharedBoolfalseK and V share one tensor (Gemma-4): half the KV
windowU320Sliding window: attend to and store only the last window tokens; 0 = full context
kv_precisionPrecision—

kind = "mla" — multi-head latent attention, optionally sparse:

FieldTypeDefaultDescription
countU32—
latent_dim, rope_dimU32— / 0KV per token = (latent + rope) × bytes
kv_precisionPrecision—
windowU320Recent tokens attended directly. Without a history path, 0 means the whole context; with one, 0 means no local window
historyTableunsetLong-range path { compress_ratio, index_topk, indexer }: the history at stride compress_ratio (1 = every position), all of it or the index_topk entries an indexer = { heads, head_dim, kv_precision, precision } selects (the indexer scores every entry and keeps its own KV; precision is the scoring GEMM’s rate, default attention_precision — DeepSeek-style indexers score in fp8)
heads, qk_head_dim, v_head_dimU32unsetHead shape for the score/AV FLOP count (2 × heads × (qk + v) per pair). All three or none; absent = 4 × hidden_dim per pair
q_latent_dim, o_latent_dimU32unsetLow-rank query / output projections (q_lora_rank, o_lora_rank); size the attention projections replicated under dp_attention. Absent = full-rank

Kimi-K2 is one mla class (full context); DeepSeek-V4 is three (window 128 only; window + top-k of the ÷4 history with an indexer; window + the whole ÷128 history); GLM-5 is top-2048 over the uncompressed history.

kind = "linear" — linear attention / SSM (GatedDeltaNet, Mamba, KDA):

FieldTypeDefaultDescription
countU32—
state_bytesU64—Fixed per-sequence state per layer (reserved for the sequence’s lifetime and read once per step); no context-scaling work

Shipped model presets: catalog/models/ (inference_lab::catalog::model_names()), each with the derivation of its numbers from the HF config in its header.


[scheduler]

FieldTypeDefaultDescription
max_num_batched_tokensU32—Token budget per iteration
max_num_seqsU32—Running-request cap. In serve mode this is retunable at runtime — see Saturation and capacity
max_waitingU320serve only: refuse arrivals with HTTP 529 once this many requests are waiting. 0 = unbounded (queue without limit, never refuse). See Saturation and capacity
policyString—fcfs, priority, sif, lif, sof, lof, stf, ltf (sjf = sof)
enable_chunked_prefillBool—Split long prefills across iterations
long_prefill_token_thresholdU320Prefill chunk cap; 0 = no cap. Defaults to 4% of max_seq_len when max_num_partial_prefills > 1
max_num_partial_prefillsU321vLLM’s knob; only its effect on the threshold default is modelled
block_sizeU32—KV block size, tokens
gpu_memory_utilizationFloat0.9Fraction of GPU memory the engine may use (vLLM’s --gpu-memory-utilization); the KV cache gets what is left after the weights
kv_cache_capacityU640Explicit KV cache bytes across the TP group; 0 derives it from gpu_memory_utilization
max_model_lenU32model’s max_seq_lenServing-time context limit (only the chunked-prefill threshold default depends on it)
enable_preemption_freeBoolfalseAdmit only what can grow to prompt + max_output without preemption
balance_setTableunsetBalance-set admission control (Denning’s medium-term scheduler): { high, low } as fractions of KV capacity. Admission stops when the running working set (resident context of running requests) reaches high and resumes only once it falls below low (hysteresis; low defaults to high). Holds the overflow in the queue instead of admitting-and-evicting, so recently-idle sessions’ cached prefixes survive in the reserved 1 − high headroom. Absent = overcommit (today’s behaviour)
enable_cascade_attentionBoolfalseLoad a batch’s shared prompt prefix once per iteration

Workload file

The [workload] table, at top level of workloads/<name>.toml.

FieldTypeDefaultDescription
arrival_patternString—poisson, uniform (= fixed_rate), burst, closed_loop, batched
arrival_rateFloat1.0Requests/s for the open-loop patterns
rate_scheduleTableunsetTime-varying rate: { type = "sine", min, max, period_secs }, { type = "square", low, high, period_secs, duty }, or { type = "trace", points = [[t, rate], ...] }
num_concurrent_usersU32unsetUsers for closed_loop
closed_loop_jitter_secsFloatunsetUniform stagger of the initial closed-loop arrivals
input_len_dist, output_len_distTable—{ type = "fixed", value }, { type = "uniform", min, max }, { type = "normal", mean, std_dev }, { type = "lognormal", mean, std_dev } (ignored in dataset mode for input)
num_requestsU32unsetMaximum total generated requests; in session mode every step counts (use num_sessions to bound starts instead)
duration_secsFloatunsetStop admitting arrivals after this many simulated seconds, then drain requests already in flight
dataset_pathStringunsetJSONL in OpenAI batch format; prompts are tokenised with --tokenizer and hashed per KV block so shared prefixes hit the prefix cache
sessions_pathStringunsetSession file (JSONL, one session per line, see Sessions). The arrival pattern then governs session starts (arrival_rate in sessions/s; closed_loop holds num_concurrent_users sessions in flight); each later step arrives at its parent’s completion plus the step’s gap. Mutually exclusive with dataset_path; length distributions are ignored
num_sessionsU32unsetSession mode: maximum session starts; does not limit the total request steps emitted by those sessions (the file is cycled, so it may exceed the file’s count)
stationary_start_sessionsU32unsetSession mode: start this many sessions at t=0 at a time-weighted step in their trace, with fresh inherited-context hashes, then resume the open-loop arrival clock
resample_sessionsBoolfalseSession mode: draw each session uniformly from the file with replacement instead of walking it in order; repeated instances receive fresh block hashes
seedU64—

[memory]

Optional; no tiering when absent. Which of the hardware’s stores hold KV evicted from HBM, closest first, and how much of each. A worker (a tp-GPU replica) reaches each tier over its own link — its GPUs’ ports pooled, so a tp = 2 worker on gh200 promotes at 2 × 450 GB/s from 2 × 120 GB of Grace memory. A per = "node" store is shared by the workers on that node (gpus_per_node / tp of them): what one demotes, its neighbours can promote. A worker wider than a node pools the node stores it spans.

[memory]
tiers = ["peer_hbm", "host_dram", "nvme"] # peer_hbm is declared by b200
preset = "reactive"                            # reactive | oracle (optional bundle)
source = { policy = "promote" }                # promote | min_time
hbm_eviction = { policy = "lru" }              # lru | outlook
write = { policy = "write_back" }              # write_back | write_through | selective | live
eviction = { policy = "fifo" }                 # fifo | lru | ttl | outlook
prefetch = { policy = "none" }                 # none | outlook
backup = "on_evict"                            # on_evict | on_land
hit_refresh = "first_tier"                     # first_tier | none
promote_fill = "through"                       # through | buffer | direct
storage_prefetch = { policy = "wait_complete" } # wait_complete | best_effort | timeout
load_overlap = "layerwise"                     # layerwise | none
hbm_evict_backed_first = false
[memory.capacity]
host_dram = 1.0e12          # bytes per instance given to KV
FieldTypeDefaultDescription
tiersArray[]Tier names from the hardware’s [memory], closest first. Each must be reachable from a GPU over the hardware’s links. peer_hbm consults same-node sibling HBM and has no capacity or write step
capacityTablefullPer-store cap on bytes per instance
presetStringunsetA named policy bundle; any field set explicitly overrides the preset’s choice. reactive: promote / lru / selective (min_hits 1) / lru / none / backed-first — decides only from what has already happened, as shipped stacks do. oracle: min_time / outlook / live / outlook / outlook / backed-first — reads every session’s announced re-entry. Both use staged through reads, wait for storage prefetch, and overlap the final load layer-wise. The preset covers KV movement only: on a pool with several workers or DP-attention ranks, pair it with a [router] that sends a re-entry to the worker holding (or prefetching) its prefix — kv_aware or prefix_affinity — or its evictions and prefetches serve arrivals that land elsewhere
sourceTablepromoteWhere a re-entry’s tier-held prefix comes from: promote — fetch it (a hit is a hit); min_time — fetch it only if the transfer, at the fetch path’s current fair share, beats recomputing those tokens at the worker’s roofline; otherwise recompute (the tier keeps its copy)
hbm_evictionTablelruWhich free HBM block is recycled first: lru — least recently freed; outlook — blocks with no announced re-entry first (LRU among them), then the farthest re-entry first, each sequence tail first
writeTablewrite_backWhen a block’s KV is written to the first tier: write_back — when its HBM block is recycled, if no tier holds it; write_through — as soon as it is produced; selective (min_hits, default 1) — on its min_hits-th HBM hit, and dropped on eviction otherwise (SGLang HiCache’s three positions); live — when recycled, only if its session has announced a re-entry (a finished trajectory is dropped)
evictionTablefifoHow every tier picks what to recycle: fifo (least recently inserted), lru (least recently inserted or promoted from), ttl (seconds: LRU, and any block untouched that long is dropped whether or not the store is full), outlook (no announced re-entry first, then farthest re-entry first). Stores hold ranges of a sequence, written and stamped together; a victim range is recycled from its tail under every policy, so what survives of a sequence in a store is a prefix
prefetchTablenoneWhether a demoted prefix is pulled back ahead of its re-entry: none; outlook (lead, default 0 s) — when a session step completes, plan a promotion of the prefix its next step re-enters with, starting so it lands lead seconds before that arrival at the fetch path’s fair share at planning time; whatever is still in HBM when the plan fires needs nothing, and a re-entry that arrives mid-transfer joins it
backupStringon_evictWhen a tier forwards a block to the tier below it: on_evict — only when it evicts the block (a store → store transfer of what would otherwise be dropped); on_land — as soon as the block’s write into it lands, so every tier below the first receives a copy within a transfer of production (SGLang HiCache backs a node up to its storage backend the moment its device → host DMA completes). Under on_land a private per-rank host tier’s fresh KV reaches a shared storage tier at once rather than when the host tier ages it out
hit_refreshStringfirst_tierWhether a prefix hit in HBM re-stamps tier copies as recently used: first_tier — the first tier below HBM ages with HBM (HiCache’s device and host tiers share one radix tree and one last_access_time; lower tiers see only the references that reach them); none — tier copies are re-stamped only when promoted from
promote_fillStringthroughHow a lower-tier read reaches HBM: through — transfer into each closer store in turn, publishing and retaining each copy before a separate final HBM load; buffer — use the same real staged transfers but release their intermediate copies after the final load; direct — one legacy source-to-HBM transfer with no staged cache fill
storage_prefetchTablewait_completeWhat a demand does while an external-store hit is staging: wait_complete — wait for the stage; best_effort — stage only while queued, then cancel an unfinished leg when an admission slot is available; timeout (seconds, finite and > 0) — wait no longer than the deadline. Cancellation keeps earlier completed stages, uses any prefix already in the closest tier, and recomputes the still-external suffix
load_overlapStringlayerwiseWhether the final closest-store-to-HBM load pipelines with the request’s first prefill pass: layerwise — aggregate approximation with elapsed time max(load, compute); none — the full load completes before compute begins
hbm_evict_backed_firstBoolfalseWhen HBM must recycle a block, take one whose KV a tier already holds (a free drop) over the policy’s first choice, looking 16 blocks up the free queue

The outlook, live and min_time/prefetch policies act on a session step’s outlook: on a session workload, when a step completes the simulator knows its successor’s arrival (completion plus the recorded gap) and how much of the context it re-enters with, and marks those blocks — in HBM and in the tiers — with that time. Other workloads announce nothing, so under live nothing is written and outlook eviction reduces to LRU. The gap between the reactive and oracle presets on the same replay is the value of knowing the re-entry.

Under promote_fill = "through", tiers are inclusive: a promoted block keeps every tier copy it staged through (KV is immutable), so its next eviction from HBM is a free drop and only blocks no tier holds ever cost a write. buffer keeps only the backing source; direct changes no store residency. Writes are transfers GPU → store on the same graph as promotions (full duplex: they share a port with fetches only in the reverse direction, but do share an NVMe pool’s drives); a block is resident once its write lands and a promotion of a block still arriving waits for it (the write-before-reuse race). A full store evicts its victim into the next tier as a store → store transfer, or drops it; a block dropped without ever having been promoted counts as dead bytes. A prompt whose blocks sit below the closest store is staged upward one store at a time, sharing every traversed edge with whatever else is in flight. These legs hold no HBM. The final closest-store-to-GPU leg reserves the request’s landing blocks; direct mode instead reserves them for its one source-to-GPU transfer.

A staged leg is admission-controlled against its destination store, the way HiCache bounds a storage prefetch by host-pool free space: it starts only if the store can take its blocks by evicting nothing that is pinned, and on start it pins its destination copy — in flight and after landing — until the request’s final HBM load consumes it (or the stage is cancelled, times out, or is abandoned). A request whose stage does not fit yet holds at the head of its worker’s queue until a stage completion or consumption frees pinned room; one whose external prefix exceeds the destination store outright abandons the prefetch and recomputes the external suffix. Under storage_prefetch = timeout, the deadline starts when the stage starts, not while it waits for room.

pin is set on each hardware store entry. Staged reads always protect their store source until that leg drains. On a direct read, pin = true prevents the source from being recycled; with pin = false, completion lands only the prefix still present, releases the unused HBM reservation, and recomputes the suffix. Peer HBM defaults to pinned; normal stores preserve the unpinned direct-read default.

The summary’s memory section reports, per store name, blocks held, bytes written / read / dead, evictions and expiries; per link name, bytes moved and utilisation; and totals of bytes written, bytes promoted, peer-HBM bytes promoted, pin stalls, partial landings, and promotions that waited on a write. The prefix_cache section counts lookups recomputed instead of fetched (min_time) and prefetches started (prefetch = outlook), with their tokens.

[prefill]

Optional. A disaggregated topology: this block is the prefill pool, and the hardware entry — its hardware, tp/ep/dp_attention, replicas and [memory] — becomes the decode pool. Arrivals enter the prefill pool through [router], prefill there against that pool’s HBM and [prefill] memory tiers, and hand their KV to a decoder chosen by [decode_router] over the network: the prefill worker’s nic link, the core, the decode worker’s nic. The first token rides with the hand-off. KV moves one way: a decoder’s HBM shortens later hand-offs of the same prefix, but nothing flows back to the prefill side, so a re-entry whose prefix exists only in decode HBM is recomputed by prefill. A per = "cluster" store named in both pools’ tiers is the one shared tier.

[prefill]
replicas = 2                       # prefill workers
parallel = { tp = 8, dp_attention = true }
memory = { tiers = ["host_dram", "nvme"], preset = "reactive", backup = "on_land" }
kv_link_bw = 4e11                  # optional core cap, bytes/s
FieldTypeDefaultDescription
hardwareString or Tablethe entry’s hardwareCatalog preset or inline hardware table for the prefill pool
parallelTablethe entry’s tp/ep/dp_attentionParallel layout of a prefill worker
replicasU321Prefill workers behind [router]
memoryTable{}KV tiers on the prefill side, from its hardware’s stores (same keys as [memory]). Absent: prefill keeps prefixes in HBM only
kv_link_bwF64unsetCapacity of the network core between the pools, bytes/s, shared by every hand-off in flight. Unset: the NICs alone bound them; an error when the hardware has no network links either

backup, hit_refresh and promote_fill are graph-wide and taken from the pools that have tiers. [scheduler] (including balance_set) applies to both pools.

[router] and [decode_router]

Optional; round_robin when absent. [router] picks the replica each arriving request enters (replicas on the hardware entry; the prefill pool on a disaggregated topology). [decode_router] picks the decode worker each hand-off goes to on a disaggregated topology, and defaults to [router]. The KV-reading policies look up each replica’s state for the prompt on every decision — an estimate from the replica’s block index, as a KV-aware front end sees it, not the scheduler’s admission-time lookup.

FieldTypeDefaultDescription
policyString"round_robin"round_robin, least_loaded, prefix_affinity, kv_aware, kv_aware_decode
max_load_ratioF64unsetprefix_affinity only: pass over the prefix holder for the least-loaded replica when its requests in system exceed max_load_ratio × the pool mean (bounded-load affinity)
load_weightF641.0 / 64.0kv_aware: weight on the replica’s queued prefill tokens. kv_aware_decode: tokens of transfer one running sequence is worth (default one 64-token block)
  • round_robin cycles through the replicas.
  • least_loaded picks the fewest requests in system (running + waiting), ties by queued prefill tokens, then index.
  • prefix_affinity picks the replica holding the longest cached prefix of the prompt (any tier); with none anywhere it falls back to least_loaded.
  • kv_aware minimises (prompt − cached prefix) + load_weight × queued prefill tokens: the prefill work the request adds plus the prefill work already ahead of it, in tokens. A prefill-side policy: on a decode pool the load term is always zero.
  • kv_aware_decode minimises (context − prompt prefix resident in the decoder's HBM) + load_weight × running sequences: the KV the hand-off must move plus the decode batch it joins, in tokens. Decoders whose free KV cannot hold the incoming context are passed over while any can.

The summary’s router section (and decode_router on a disaggregated topology) reports requests per replica and, for the KV-reading policies, how many decisions had a cached prefix on some replica, how many went to a holder, and how many went away from the longest holder. handoff reports transfers, bytes moved, and bytes skipped because the chosen decoder already held the prefix. For session workloads, reusable_kv reports the joint prefill/decode residency of reusable KV, derived hit, recompute and transfer fractions, a parent-prefill, parent-decode, and inherited-context split of prefiller misses, and per-decode-rank detail. Under same-rank session affinity, that split separates eviction-driven recomputation from decoder-output KV that was never written back. hbm reports capacity, resident prefix bytes, active/reserved bytes, and actual HBM eviction bytes per worker. sessions reports completed and deadline-censored sessions plus turns per started session; simulation records the arrival deadline snapshot and post-deadline drain; work separates logical prompt tokens from positions actually computed by prefill.

[speculative]

Optional. Decode steps then verify 1 + draft positions and advance by 1 + accepted.

FieldTypeDefaultDescription
gammaU32—Draft length (fixed) or maximum candidate depth (budget policies)
acceptanceTable—{ kind = "constant", alpha }, { kind = "per_position", a = [...] }, or { kind = "trace_rounds", path } (CSV bank of real rounds: commits,category,a0..aD-1)
policyString"fixed"fixed, goodput_budget, gated_budget, gated_aggregate
measured_costTableunset{ path, ref_seq_len }: measured (batch_size, num_draft_tokens, step_seconds) grid that prices decode steps and the policy’s cost curve instead of the roofline
switchTableunconstrained{ cooldown_rounds, max_step, cost_ms } for gated_aggregate
drafterTablefree drafter{ kind = "fraction", frac }, { kind = "autoregressive", dense_params, expert_params, num_experts, experts_per_tok, shared_experts }, or { kind = "block_parallel", params, block }

Fault Injection

The serve mode can kill a streaming chat completion mid-generation in every way a real downstream does — on demand, deterministically, per request. It exists as the test double for gateway resilience work (mid-stream continuation/resume middleware): if the sim can produce every death signature, that middleware can be e2e-tested without waiting for real incidents.

Every faulting stream first emits real partial output — the initial role frame plus after_chunks content-bearing delta frames of deterministic placeholder text — so there is always something to resume. The fault path bypasses the simulation engine (like echo-directives) and contains no randomness: the same trigger produces the same frames and the same death at the same byte, every time.

Scope: streaming POST /v1/chat/completions and streaming POST /v1/completions. A fault header on a non-streaming request is rejected with a 400 (invalid_fault_directive) — never silently ignored. With no trigger present, behavior is completely unchanged.

Both endpoints matter because a mid-stream-continuation resume leg is a streaming /v1/completions request, so chain-resume tests need to kill one. The frames follow the endpoint’s own wire shape: /v1/completions emits text_completion chunks carrying text and has no role frame, so after_chunks=N puts exactly N frames on the wire before the death (one fewer than the chat flavor, which leads with its role frame). mid_reasoning and mid_tool_call have no delta object to use there, so they stream the same partial payloads as raw text — an unterminated <think> block and a tool call the model never finished writing — which is how both actually appear on a real base-model completions stream.

Trigger: the x-inference-lab-fault header

The header is only honored when the server runs with --enable-directives — it is client-controlled and can stall or abort connections at will, so it sits behind the same “untrusted clients must not reach this server” gate as echo-directives. Without the flag the header is a 400 (invalid_fault_directive), never silently ignored. The staging deployment already sets the flag.

x-inference-lab-fault: <mode>[;after_chunks=<u32>][;delay_ms=<u64>][;utf8=<bool>]
ParameterDefaultMeaning
<mode>requiredone of the eleven mode names below
after_chunks3content-bearing delta frames emitted before the fault fires (the initial role-only frame is always sent and not counted)
delay_ms10fixed pacing between frames, milliseconds
utf8falsecut_mid_frame only: cut inside a multi-byte UTF-8 character

Parts are ;-separated; whitespace around parts and = values is ignored. Unknown modes, unknown keys, or malformed values are a 400 listing the valid modes.

A header was chosen over a body extension because the platform’s proxy layer runs strict request sanitization that strips unknown body fields; headers pass through. Verify this empirically through the full stack (client → dwctl → onwards → sim) when the branch reaches a preview environment — full-platform verification is out of scope here and happens after deploy.

Fallback trigger: static per-model config

For clients that cannot set a header, a model’s TOML config can apply one fault to every streaming chat completion on that model (non-streaming requests are served normally; an explicit header on the request still wins). Unlike the header, this is operator input validated at server boot, so it does not require --enable-directives:

[fault]
mode = "cut_mid_frame"   # same names as the header
after_chunks = 5         # optional, default 3
delay_ms = 10            # optional, default 10
utf8 = true              # optional, cut_mid_frame only

An invalid [fault] block fails server startup, not individual requests.

Precedence per request: header > model [fault] config > (echo-directives >) normal path.

Modes

ModeAfter the N content frames…Client observes (curl)
cut_between_framesconnection closes on a frame boundary, without the chunked-encoding terminator (FIN)exit 18, transfer closed with outstanding read data
cut_mid_framehalf of the next frame’s bytes, then close — torn JSON. utf8=true cuts one byte into a 2-byte UTF-8 character (é) in the delta textexit 18, partial data: line
resetabortive close: SO_LINGER=0 then drop → TCP RST, not FINexit 56, connection reset by peer
stallnothing, forever; connection stays open until the client gives upexit 28 (client timeout)
error_envelope_200OpenRouter-style error envelope as an SSE data frame, then [DONE]; HTTP status stays 200exit 0, {"error":{"message":…,"code":502,"metadata":{…}}}
error_400_in_ssevLLM-style 400 object in-stream (the nemotron-incident signature), then clean close, no [DONE]exit 0, {"object":"error",…,"code":400}
no_donefinish_reason frame, then the separate choices: [] usage frame if requested, then clean close — [DONE] never comesexit 0, stream just ends
no_usagefinish_reason frame then [DONE], but the usage frame never arrives (even when stream_options.include_usage was set)exit 0, usage missing
cancelled_499the exact dynamo frontend-cancellation body, then clean closeexit 0, {"error":{"code":499,"message":"CancelledError: ","type":"request_cancelled"}}
mid_reasoningframes carry reasoning_content deltas instead of content; dies cut_between_frames-styleexit 18, last deltas are reasoning
mid_tool_callframe 1 announces a tool call (id + name), later frames stream arguments fragments that never terminate; dies cut_between_frames-styleexit 18, partial tool call

Notes:

  • The error body shapes for error_envelope_200 and error_400_in_sse are representative; exact shapes sync with the death-taxonomy workstream as it lands. The cancelled_499 body is exact.
  • delay_ms=0 is fine for the graceful modes, but the abrupt modes always wait a short flush grace (~25 ms) before killing the connection so the partial output reliably reaches the wire first.
  • reset needs the raw socket, which the server threads through per-connection; on non-unix platforms (or when handlers are driven outside the real server, as in unit tests) it degrades to a FIN with a warning log.

Examples

All against a local sim (inference-lab serve --config configs/ --hardware b200 --port 8080); $BODY is any streaming chat request:

BODY='{"model":"DeepSeek-V4-Flash","stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hello"}],"max_tokens":16}'
URL=http://localhost:8080/v1/chat/completions
# 1. cut_between_frames — 5 frames then FIN (curl exit 18)
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cut_between_frames;after_chunks=5' -d "$BODY"

# 2a. cut_mid_frame — torn JSON frame
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cut_mid_frame' -d "$BODY"

# 2b. cut_mid_frame, cut inside a multi-byte UTF-8 character (pipe through xxd to see it)
curl -sN $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cut_mid_frame;utf8=true' -d "$BODY" | xxd | tail

# 3. reset — TCP RST (curl exit 56, "connection reset by peer")
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: reset' -d "$BODY"

# 4. stall — 3 frames then silence; bound the wait client-side
curl -N --max-time 10 $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: stall' -d "$BODY"

# 5. error_envelope_200 — OpenRouter-style envelope then [DONE]
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: error_envelope_200' -d "$BODY"

# 6. error_400_in_sse — 400-shaped object inside the 200 stream
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: error_400_in_sse' -d "$BODY"

# 7. no_done — finish_reason + usage frame, then the stream ends without [DONE]
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: no_done' -d "$BODY"

# 8. no_usage — [DONE] arrives but the requested usage never does
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: no_usage' -d "$BODY"

# 9. cancelled_499 — exact dynamo cancellation body
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cancelled_499' -d "$BODY"

# 10. mid_reasoning — dies while streaming reasoning_content deltas
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: mid_reasoning' -d "$BODY"

# 11. mid_tool_call — dies with a tool call's arguments unterminated
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: mid_tool_call;after_chunks=4' -d "$BODY"

Killing a resume leg

A resume leg is a streaming /v1/completions request whose prompt is token ids. Every mode above works against it; only the frame shape differs (no role frame, text instead of a delta):

RESUME='{"model":"DeepSeek-V4-Flash","prompt":[1,2,3,4],"stream":true,"priority":0,"stream_options":{"include_usage":true},"max_tokens":64}'

# exactly 2 text frames, then FIN (curl exit 18)
curl -N http://localhost:8080/v1/completions -H 'content-type: application/json' \
  -H 'x-inference-lab-fault: cut_between_frames;after_chunks=2' -d "$RESUME"

Saturation and Capacity

serve mode queues arrivals without bound by default: there is no request volume at which it refuses work. That is fine for latency experiments, but it makes the server useless for testing a client that is supposed to react to overload — a controller that discovers a model’s sustainable concurrency from downstream rejections has nothing to react to, because no rejection ever arrives.

Two knobs change that, and they do different things:

KnobWhat it changesTypical use
max_waitingHow long the queue may get before the server refusesSwitch 529s on and off
max_num_seqsHow fast the engine drains workModel a scale up/down

max_waiting does not change capacity. Lowering it makes rejections start sooner, but the engine serves at exactly the same rate, so the concurrency a client settles at is set by queue policy rather than by anything physical. max_num_seqs is the dial that changes the drain rate: halve it and the queue backs up at roughly half the offered load.

The waiting bound

Set it in a model’s [scheduler] table:

[scheduler]
max_num_seqs = 256
max_waiting = 64   # 0 (the default) = unbounded, never refuse

or override every model’s value at the command line:

inference-lab serve --config configs/ --hardware b200 --max-waiting 64

Once max_waiting requests are queued, further arrivals get:

HTTP/1.1 529
content-type: application/json

{"error": {
  "message": "Server is at capacity: 64 requests waiting (max_waiting = 64). Retry with reduced concurrency.",
  "type": "overloaded_error",
  "code": "queue_saturated"
}}

Why 529, and not 503 or 429

529 is the convention for “the engine has nowhere to put this request”. Clients that adapt their concurrency generally key on 529 alone, because the other two mean something a client should not answer by shedding load: 503 is “this service is unavailable” and 429 is “you exceeded a quota or a proxy’s own limit”. serve already returns 503 when the engine channel is closed, which is a liveness failure rather than saturation, and stays a distinct code.

Two properties the rejection is built to have

It is returned before the response starts. A real engine that admits a request, sends 200 plus streaming headers, and only then discovers it cannot schedule it has spent its status code: the failure can then only be an error object inside the stream, or a stream that stops with no content. Neither is classifiable as overload, so the client learns nothing. The check therefore runs in submit_engine_request, the single funnel into the engine, before any handler has built a response — a refused request returns a bare status and envelope, never a text/event-stream.

It is immediate, never a stall. A bounded queue that parks requests until some later timeout produces a client that waits and eventually gives up, which consumes a client slot for the whole timeout and still carries no overload signal. The bound is a synchronous check against a published queue depth.

What “waiting” counts

The depth compared against max_waiting is every worker’s num_waiting() (queued requests plus those parked on a KV transfer or a staged read), plus arrivals the HTTP layer has admitted that the engine has not stepped into a scheduler yet. The second term matters: Engine::submit only queues an Arrival event, so without it a burst would read as depth 0 and be admitted wholesale.

Runtime capacity control

GET /control/capacity reports every model’s knobs and live depth:

curl localhost:8080/control/capacity
[{"model":"gpt-oss-20b","max_waiting":64,"max_num_seqs":256,"waiting":12,"running":256}]

POST /control/capacity retunes them without a restart — which matters because a restart drops every in-flight request, destroying the before-and-after that a capacity-change experiment depends on. Both fields are optional, and model defaults to every loaded model:

# Scale down: the engine now drains at a sixteenth of the rate.
curl -X POST localhost:8080/control/capacity \
  -H 'content-type: application/json' -d '{"max_num_seqs": 16}'

# Turn 529s on (or off again with 0) against a running server.
curl -X POST localhost:8080/control/capacity \
  -H 'content-type: application/json' -d '{"max_waiting": 64}'

Both changes act on live state:

  • Lowering max_num_seqs drains, it does not evict. The cap gates admission only, so requests already running above the new cap run to completion and the batch shrinks to the new size. Nothing in flight is lost.
  • Lowering max_waiting only affects requests that have not arrived yet. Anything already queued keeps its place; its status code is long since spent.

max_num_seqs must be at least 1 — a cap of 0 would accept requests and then never schedule them, which is exactly the stall the bound exists to replace. max_waiting: 0 is meaningful (unbounded) and allowed.

Cold start: the first burst after idle is unpaced

serve paces simulated time to wall-clock from an epoch fixed when the engine starts, and simulated time only advances while there is work to do. An idle server therefore accumulates a deficit, and the first requests after a quiet period are served as fast as the CPU allows rather than at the rate the model predicts — a single 200-token stream into a freshly booted server returns in single-digit milliseconds.

The deficit burns off once there is continuous work: under sustained load the sim clock catches up within about a second, after which pacing is correct and steady. Measured on gpt-oss-20b / b200 at 24 concurrent requests, a 200-token response settles at a flat ~293 ms; drop max_num_seqs to 8 and the same load settles at ~646 ms with throughput down by roughly the same factor.

The practical consequence is only for short experiments: a benchmark that fires one burst at a just-started server measures the CPU, not the modelled hardware. Give it a second of warm-up load first.