Introduction
Inference Lab is a simulation framework designed to evaluate and analyze LLM workloads.
It uses discrete-event simulation to model the behavior of a multi-GPU node serving LLM inference requests with the vLLM library. It contains a facsimile of the vLLM queueing, scheduling, and execution logic, with only the actual model inference replaced by a performance model based on the supplied GPU specs and model architecture.
Within each simulation step, the simulator:
- Processes any newly arrived requests, adding them to the scheduling queue.
- Schedules requests to serve based on the selected scheduling policy.
- Calculates the compute and memory bandwidth usage for the workload that the scheduled requests represent, and the theoretical time required to execute the workload on the specified hardware.
- Increments the simulation time by the calculated execution time, updating the state of all requests accordingly.
Caveats:
- Step times are a datasheet roofline: peak FLOP rate and HBM bandwidth
per precision stream, collectives on the fabric preset added serially,
and no kernel-efficiency or fixed per-step overhead term. Every
latency and throughput is an upper bound; the optional
time_correction = { alpha, beta }on a hardware entry calibrates the step (alpha × roofline + beta) against a measured engine.
Features
- Roofline Performance Modeling: compute (FLOPS) and memory bandwidth constraints per precision stream
- Multiple Scheduling Policies: FCFS, Priority, SJF, and more
- Chunked Prefill: Simulates realistic request interleaving
- KV Cache Management: Models GPU memory and KV cache utilization
- Workload Generation: Supports Poisson, Gamma, and closed-loop patterns
- WebAssembly Support: Run simulations in the browser via WASM
Quick Start
See the Getting Started guide to begin using Inference Lab.
Getting Started
This guide will help you get started with Inference Lab.
Installation
Install from crates.io:
cargo install --locked inference-lab
Or build from source:
cargo build --release
Running Your First Simulation
From a checkout of the repository:
inference-lab --config configs/llama-3-70b.toml --workload workloads/quick.toml
configs/ holds one file per model (each with its hardware entries) and
workloads/ the arrival patterns and request shapes; pass --hardware <name> when a model config has more than one entry.
Next Steps
- Learn about configuration options
- Explore running simulations
Running Simulations
This guide covers how to run simulations and interpret results.
Basic Usage
Run a model config on one of its hardware entries against a workload:
inference-lab -c configs/gpt-oss-120b.toml --hardware b200 -w workloads/chat-closed-256.toml
For dataset mode, add tokenizer and chat template:
inference-lab -c configs/llama-3-70b.toml -w workloads/dataset-poisson.toml \
--tokenizer tokenizer.json \
--chat-template None
See Configuration for details on configuring workloads, policies, and hardware.
Output Modes
Console Output (Default)
By default, the simulator displays:
- Real-time progress bar
- Current simulation time
- Queue status (running/waiting requests)
- KV cache utilization
Final output includes:
- Latency metrics (TTFT, E2E, per-token)
- Throughput metrics (tokens/sec, requests/sec)
- Utilization statistics (KV cache, FLOPS, bandwidth)
- Preemption statistics
JSON Output
Save results to a file:
inference-lab -c configs/gpt-oss-120b.toml --hardware b200 -w workloads/chat-closed-256.toml -o results.json
Combine with -q for batch processing:
inference-lab -c configs/gpt-oss-120b.toml --hardware b200 -w workloads/chat-closed-256.toml -q -o results.json
Running Multiple Experiments
Comparing Hardware
for hw in b200 b300 gh200; do
inference-lab -c configs/gpt-oss-120b.toml --hardware $hw \
-w workloads/chat-closed-256.toml -q -o results_$hw.json
done
Sweeping Engine Args
for batch_size in 4096 8192 16384; do
sed "s/max_num_batched_tokens = .*/max_num_batched_tokens = $batch_size/" \
configs/gpt-oss-120b.toml > /tmp/gpt-oss-120b_$batch_size.toml
inference-lab -c /tmp/gpt-oss-120b_$batch_size.toml --hardware b200 \
-w workloads/chat-closed-256.toml -o results_$batch_size.json
done
Multiple Seeds
Override the workload’s seed:
for seed in {1..10}; do
inference-lab -c configs/gpt-oss-120b.toml --hardware b200 \
-w workloads/chat-closed-256.toml --seed $seed -q -o results_$seed.json
done
Understanding Results
Latency Metrics
Time to First Token (TTFT)
- Time from request arrival to first token generation
- Lower is better for interactive applications
- Affected by: queue wait time, prefill computation
End-to-End (E2E) Latency
- Total time from request arrival to completion
- Includes prefill and all decode steps
- Key metric for overall user experience
Per-Token Latency
- Average time between consecutive output tokens
- Lower is better for streaming applications
- Primarily affected by batch size and model size
Throughput Metrics
Input Tokens/sec
- Rate of processing prompt tokens
- Indicates prefill throughput
Output Tokens/sec
- Rate of generating output tokens
- Indicates decode throughput
Requests/sec
- Overall request completion rate
- Key metric for capacity planning
Utilization Metrics
KV Cache
- Percentage of KV cache memory in use
- High utilization may lead to preemptions
FLOPS
- Percentage of compute capacity utilized
- Low FLOPS may indicate memory bottleneck
Bandwidth
- Percentage of memory bandwidth utilized
- High bandwidth utilization indicates memory-bound workload
Preemption Statistics
Preemptions occur when new requests need memory but the KV cache is full:
- Total number of preemptions
- Average preemptions per request
- Can significantly impact TTFT for preempted requests
Troubleshooting
Simulation running slowly?
- Reduce
num_requestsor use-qflag - Build with
--release
Too many preemptions?
- Raise
gpu_memory_utilizationor setkv_cache_capacityin[scheduler] - Reduce
max_num_seqsormax_num_batched_tokensin scheduler config
Dataset loading errors?
- Verify
--tokenizerand--chat-templateflags are provided - Check JSONL format matches OpenAI batch API format
For more details, see CLI Reference and Configuration.
Configuration
A simulation is a model config × one of its hardware entries × a
workload. Model configs live in configs/, one file per model
deployment; workloads in workloads/. Unknown fields in either are rejected.
inference-lab --config configs/qwen3.6-35b-a3b-fp8.toml --hardware b200 \
--workload workloads/chat-closed-256.toml
Model config
model = "qwen3.6-35b-a3b-fp8" # catalog preset, or an inline [model] table
[scheduler] # engine args shared by every hardware entry
max_num_batched_tokens = 16384
max_num_seqs = 4096
policy = "priority"
enable_chunked_prefill = true
block_size = 64
[hardware.b200] # one entry per hardware this model runs on
tp = 1
[hardware.b300]
tp = 1
[hardware.gh200]
tp = 1
scheduler = { max_num_batched_tokens = 8192 } # per-entry override
- model — a catalog preset name or an inline
[model]table: weight streams and token-mixing layer classes. - [scheduler] — engine arguments: batching, KV blocks, memory utilisation, scheduling policy.
- [hardware.<name>] — one per hardware the model is deployed on. The
name is a hardware preset (
b200,b300,gh200,h100) unless the entry setsspec. Each entry gives the parallel layout —tp(replica world size, weights sharded, per-layer all-reduces),ep(experts sharded overepranks),dp_attention(data-parallel attention: replicated attention weights, all-gather/reduce-scatter around the FFN, or dispatch/combine all-to-alls whenep > 1) — and may overrideschedulerkeys or carry its ownspeculativeblock. The shipped configs carry the layouts production runs (tp8 + dp_attentionfor DeepSeek-V4-Pro / Kimi / GLM-5,tp = epfor the Qwen3.5/VL MoEs and Nemotron Ultra, plaintpelsewhere). - [speculative] — speculative decoding, optional; a shared default that
an entry’s
speculativereplaces (acceptance traces and measured step costs are per hardware). - [memory] — KV tiers beyond HBM, picked from the stores the hardware
offers (host DRAM over PCIe, Grace memory over NVLink-C2C, NVMe): evicted
blocks fall through them and are promoted back instead of recomputed. A
per = "node"store is shared by the workers on a node. How KV moves — fetch or recompute, what HBM and each tier evict, when blocks are written, whether a re-entry’s prefix is prefetched — is a set of policies with two presets,reactive(decides from the past, like shipped stacks) andoracle(knows every session’s next re-entry); see the reference. - replicas / [router] / [decode_router] — an entry’s
replicas(default 1) runs that many identical workers, each with its own scheduler and KV cache; the shared[router](or an entry’srouter) picks which one each request enters:round_robin,least_loaded,prefix_affinity, orkv_aware. On a disaggregated topology[decode_router](default:[router]) picks the decode worker each hand-off goes to;kv_aware_decodeprices the transfer and the decode batch (see the reference).
--hardware picks the entry; it can be omitted when a file has one.
inference-lab serve --config configs/ --hardware b200 serves every model
with a b200 entry.
Hardware
Name a shipped preset as the entry:
[hardware.b200]
tp = 2
or point an entry at another preset, or at an inline per-GPU spec (a FLOP
rate for every precision the model uses, bandwidth, capacity, optional
[fabric] and [memory]):
[hardware.isambard]
spec = "gh200"
tp = 4
[hardware.custom]
tp = 1
[hardware.custom.spec]
name = "H100"
flops_fp8 = 1.979e15 # dense FLOP/s at fp8
flops_bf16 = 9.895e14 # dense FLOP/s at bf16
memory_bandwidth = 3.35e12 # bytes/sec
memory_capacity = 85899345920 # 80 GB
How much of that memory the engine may use, and how much goes to KV, are
deployment settings and live in [scheduler] (gpu_memory_utilization,
kv_cache_capacity).
A preset also carries its node’s collective fabric — gpus_per_node,
scale_up (NVLink: bandwidth, latency, in-network reduction) and
scale_out (per-GPU NIC across nodes) — which prices the TP all-reduces and
EP all-to-alls of any entry with tp > 1 or ep > 1. An inline spec that
omits [fabric] can only be used with tp = 1, ep = 1.
Model
Name a shipped preset:
model = "gemma-4-31b-it"
or describe the architecture inline as weight streams plus layer classes:
[model]
name = "Llama-3-70B"
hidden_dim = 8192
max_seq_len = 8192
attention_precision = "fp8"
[[model.weights]] # one per-token GEMM stream per precision
precision = "fp8"
active_params = 70000000000
resident_params = 70000000000
[[model.layers]] # token-mixing layer classes
kind = "attention" # or "mla", "linear"
count = 80
heads = 64
head_dim = 128
kv_heads = 8 # GQA
kv_precision = "fp8"
MoE adds routing = { routed_experts, experts_per_tok, moe_layers } to the
expert stream; sliding-window layers are an attention class with
window; MLA / DeepSeek sparse attention is the mla kind; GatedDeltaNet or
Mamba layers are linear with their per-sequence state_bytes. See the
Configuration Reference for every field.
Scheduler
Control request scheduling and batching:
[scheduler]
max_num_batched_tokens = 8192
max_num_seqs = 256
policy = "fcfs"
enable_chunked_prefill = true
block_size = 16
Scheduling Policies
Available policies:
fcfs- First-Come-First-Served (default)sof- Shortest Output Firstsif- Shortest Input Firststf- Shortest Total Firstlif- Longest Input Firstlof- Longest Output Firstltf- Longest Total First
Chunked Prefill
Enable chunked prefill to allow interleaving prompt processing with generation:
enable_chunked_prefill = true
long_prefill_token_threshold = 512 # Optional: chunk size limit
max_num_partial_prefills = 1 # Max concurrent partial prefills
Preemption-Free Mode
Enable conservative admission control to guarantee zero preemptions:
enable_preemption_free = true
Workload
A workload file is the workload table at top level: how requests arrive and their shapes.
Synthetic Workload
# workloads/chat-poisson-5rps.toml
arrival_pattern = "poisson"
arrival_rate = 5.0
num_requests = 100
seed = 42
[input_len_dist]
type = "lognormal"
mean = 6.9
std_dev = 0.7
[output_len_dist]
type = "lognormal"
mean = 5.3
std_dev = 0.8
Arrival Patterns
poisson- Poisson process with exponential inter-arrival timesuniform- Uniform random inter-arrival timesburst- Bursty trafficfixed_rate- Fixed interval between requestsclosed_loop- Fixed number of concurrent usersbatched- Requests arrive in batches
Length Distributions
Four distribution types are supported:
Fixed:
[input_len_dist]
type = "fixed"
value = 1000
Uniform:
[input_len_dist]
type = "uniform"
min = 100
max = 2000
Normal:
[input_len_dist]
type = "normal"
mean = 1000.0
std_dev = 200.0
LogNormal:
[input_len_dist]
type = "lognormal"
mean = 6.9 # ln(1000)
std_dev = 0.7
Dataset Mode
Use real request traces instead of synthetic workloads:
dataset_path = "path/to/dataset.jsonl"
arrival_pattern = "poisson"
arrival_rate = 1.0
# These are used for sampling actual generation length
input_len_dist = { type = "fixed", value = 100 } # Ignored
output_len_dist = { type = "fixed", value = 50 } # Samples EOS
Dataset Format: JSONL file in OpenAI batch API format. Each line may target either /v1/chat/completions with a messages array or /v1/completions with a string prompt.
Example:
{"custom_id": "req-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-3.5-turbo", "messages": [{"role": "user", "content": "Hello!"}], "max_tokens": 100}}
{"custom_id": "req-2", "method": "POST", "url": "/v1/completions", "body": {"model": "gpt-3.5-turbo-instruct", "prompt": "Write a haiku about Rust.", "max_tokens": 80}}
Tokenizer: Dataset mode requires a tokenizer file to convert text to tokens. You’ll need to provide this via the --tokenizer flag:
inference-lab -c configs/llama-3-70b.toml -w workloads/dataset-poisson.toml --tokenizer tokenizer.json
The tokenizer should be a HuggingFace tokenizers JSON file (typically tokenizer.json from the model repository).
Chat Template: You’ll also need to specify how to format chat-style requests via --chat-template:
- Use
"None"for simple concatenation of messages - Use a Jinja2 template string for custom formatting (e.g.,
"{{user}}\n{{assistant}}") - Most models have their own chat template format
- Plain
/v1/completionsprompts are tokenized directly and do not use the chat template
Example with no template:
inference-lab -c configs/llama-3-70b.toml -w workloads/dataset-poisson.toml \
--tokenizer tokenizer.json \
--chat-template None
Sessions
Agentic traffic: chains of requests where each step re-enters with its parent’s whole context as prefix (prompt and the parent’s output) plus some novel tokens, after the gap the harness spent between the parent’s completion and this arrival (a tool call running, a user typing).
sessions_path = "data/sessions/tracelab.jsonl"
arrival_pattern = "poisson" # governs session starts
arrival_rate = 0.05 # sessions/s
num_sessions = 200 # stop starting sessions after this many
# Optional: seed the first live population part-way through its traces.
stationary_start_sessions = 128
# Optional: draw sessions uniformly with replacement instead of in file order.
resample_sessions = true
seed = 42
input_len_dist = { type = "fixed", value = 1 } # ignored in session mode
output_len_dist = { type = "fixed", value = 1 } # ignored in session mode
The arrival pattern decides when sessions start: poisson / uniform /
burst at arrival_rate sessions per second, closed_loop keeps
num_concurrent_users sessions in flight (a slot starts a fresh session when
its session’s last step completes), batched starts every session at t=0.
Every later step of a session arrives at its parent’s completion plus the
step’s gap, so the simulated latency feeds back into the arrival process
and long gaps are preserved. By default sessions are taken from the file in
order and the file is cycled. num_sessions bounds session starts;
num_requests separately bounds total emitted request steps across all
sessions.
stationary_start_sessions avoids waiting a session-lifetime tail for an
open-loop live population to reach stationarity. All N seeded sessions enter
at t=0; the arrival clock is held until they have started, then resumes
normally. Each starts at a trace step sampled in proportion to that step’s
gap (the time the session waits before issuing it). Its inherited context
receives fresh block hashes, so the first emitted request reports its shared
prefix but must prefill that context once; later sessions start at step 0
normally.
resample_sessions = true draws every new session uniformly from the file
with replacement. It preserves each sampled session’s within-trace structure
while avoiding file-order effects. Repeated instances receive fresh block
hashes and therefore do not share cache state.
Session file: JSONL, one session per line:
{"id": "claude:000adcd5", "steps": [
{"input": 15524, "new": 15524, "output": 111, "gap": 0.0, "kind": "user"},
{"input": 17079, "new": 1444, "output": 96, "gap": 0.104, "kind": "tool"}
]}
input is the step’s prompt length, new the tokens of it that are not the
parent’s context (input − new is the reusable prefix, capped at what the
parent actually had), output the tokens generated, gap the seconds from
the parent’s completion to this arrival (ignored on the first step), kind
free-form. Prefix identity is built from block hashes: the parent’s hashes
over the shared prefix, fresh ones for the novel tail and the step’s own
output, so a re-entry hits the parent’s generated tokens too. Whole blocks
only: a partial block continued with new tokens is new content.
examples/sessions/tracelab_export.py exports TraceLab’s
per_step_stats.parquet into this format.
Per-request CSV (--request-csv) carries, for session steps, session,
step, worker (the memory-graph id of the worker that served it), gap,
shared_toks (the most the prefix cache could serve),
cached_toks (what it did), and two reuse distances: reuse_distance_bytes
(KV bytes written into the caches between the parent’s completion and this
arrival) and reuse_touched_bytes (the same plus the free blocks that hits
pulled back into use in between). Fresh writes undercount the LRU stack
distance and the touched count overcounts it, so the pair brackets it.
Closed-Loop Workload
Simulate a fixed number of concurrent users:
arrival_pattern = "closed_loop"
num_concurrent_users = 256
closed_loop_jitter_secs = 0.05 # stagger the initial arrivals
# ... length distributions ...
Common Configuration Patterns
High Throughput Setup
Maximize batch size and token throughput:
[scheduler]
max_num_batched_tokens = 16384
max_num_seqs = 512
enable_chunked_prefill = true
Low Latency Setup
Prioritize request completion speed:
[scheduler]
max_num_batched_tokens = 4096
max_num_seqs = 64
policy = "sof" # Shortest Output First
Memory-Constrained Setup
Limit KV cache usage:
[scheduler]
kv_cache_capacity = 34359738368 # 32 GB explicit limit
max_num_seqs = 128
Next Steps
- See the Configuration Reference for exhaustive field documentation
- Learn about Running Simulations
CLI Reference
Command-line interface reference for Inference Lab.
Usage
inference-lab [OPTIONS]
A binary built with --features serve (the Docker image) has subcommands
instead: inference-lab sim [OPTIONS] takes the options below and
inference-lab serve starts the OpenAI-compatible server. serve takes
--config (a model config or a directory of them), --hardware (models
without that entry are skipped) and an optional --workload, whose
output_len_dist samples each response’s length; without one responses run
to their max_tokens. --max-waiting <N> bounds the waiting queue so the
server refuses arrivals past it with HTTP 529 (0, the default, queues without
limit); it and max_num_seqs are retunable on a running server through
POST /control/capacity — see Saturation and Capacity.
Options
Configuration
-c, --config <PATH>
Model config file (configs/<name>.toml).
- Default:
config.toml
--hardware <NAME>
Which [hardware.<name>] entry of the model config to run. Optional when
the file has exactly one entry.
-w, --workload <PATH>
Workload file (workloads/<name>.toml). Required for sim.
inference-lab -c configs/gpt-oss-120b.toml --hardware gh200 -w workloads/quick.toml
Dataset Mode
-t, --tokenizer <PATH>
Path to tokenizer file (required for dataset mode).
- Required when using
dataset_pathin configuration - Example:
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --tokenizer tokenizer.json
--chat-template <TEMPLATE>
Chat template for formatting messages in dataset mode.
- Required when using datasets
- Use
"None"for simple message concatenation (no template) - Example:
inference-lab -c configs/llama-3-70b.toml -w workloads/dataset-poisson.toml --tokenizer tokenizer.json --chat-template None - Example with template:
inference-lab ... --tokenizer tokenizer.json --chat-template "{{system}}\n{{user}}\n{{assistant}}"
Output Options
-o, --output <PATH>
Path to output JSON file for results.
- If not specified, results are only displayed to console
- Example:
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -o results.json
-q, --quiet
Suppress progress output (only show final results).
- Example:
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -q
-v, --verbose
Enable verbose output.
- Example:
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -v
--debug
Enable debug logging.
- Example:
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --debug
--no-color
Disable colored output.
- Useful for logging to files or CI environments
- Example:
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --no-color
Simulation Options
--seed <NUMBER>
Override the random seed from configuration.
- Useful for reproducible runs with different seeds
- Example:
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --seed 12345
Examples
Basic Simulation
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml
Dataset Mode
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml \
--tokenizer tokenizer.json \
--chat-template None
Save Results to File
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -o results.json
Quiet Mode with Output
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml -q -o results.json
Multiple Runs with Different Seeds
for seed in 42 43 44; do
inference-lab -c configs/llama-3-70b.toml -w workloads/quick.toml --seed $seed -o results_$seed.json
done
Exit Codes
0- Simulation completed successfully1- Error occurred (configuration error, file not found, etc.)
Configuration File Reference
Field-by-field reference for the two TOML files a simulation takes: a
model config (configs/<name>.toml) and a workload
(workloads/<name>.toml). Unknown fields anywhere in either file are
rejected.
# configs/<name>.toml
model = "deepseek-v4-flash" # catalog preset, or an inline [model] table
[scheduler] # engine args shared by every hardware entry
[speculative] # optional: speculative decoding, shared default
[router] # optional: how requests spread over replicas
[decode_router] # optional: disagg decode pool (defaults to [router])
[memory] # optional: KV tiers beyond HBM, from the hardware's stores
[prefill] # optional: a prefill pool; the entry then decodes (disaggregated)
[fault] # optional: serve-mode static fault injection
[hardware.b200] # one entry per hardware the model runs on
tp = 4
replicas = 4 # identical workers behind the router
[hardware.gh200]
tp = 4
scheduler = { max_num_batched_tokens = 4096 }
# workloads/<name>.toml — the [workload] table at top level
arrival_pattern = "closed_loop"
num_concurrent_users = 256
num_requests = 2000
seed = 7
input_len_dist = { type = "lognormal", mean = 7.0, std_dev = 0.5 }
output_len_dist = { type = "lognormal", mean = 6.5, std_dev = 0.8 }
inference-lab --config <model config> --hardware <entry> --workload <workload>
runs one; --hardware may be omitted when the file has one entry. In Rust,
ModelConfig::from_file(..).deployment(Some("b200")) gives a Deployment
(model on hardware) and .with_workload(WorkloadConfig::from_file(..)) a
Config. The WASM API takes that resolved Config as JSON: hardware,
parallel, model, scheduler, workload, optional replicas, router,
decode_router, memory, prefill, speculative.
Catalog presets
Hardware and model presets ship inside the crate (catalog/hardware/*.toml,
catalog/models/*.toml, embedded at build time). A config names one with
model = "<name>" and [hardware.<name>]; inference_lab::catalog:: {hardware_names, model_names, hardware, model} list and load them from
Rust. Presets carry only the physical hardware spec and the model
architecture; everything about a deployment (TP, memory utilisation, batch
limits) is in the config. To tweak a preset, copy its table inline.
[hardware.<name>]
One entry per hardware the model is deployed on. The entry name is the
hardware preset unless spec says otherwise.
| Field | Type | Default | Description |
|---|---|---|---|
tp | U32 | 1 | Replica world size: its GPUs pool FLOP rate, HBM bandwidth and memory; weights are sharded across them; each layer’s output is all-reduced twice (after attention, after the FFN) |
ep | U32 | 1 | Experts sharded across ep of the ranks (divides tp). With TP attention every rank holds every token, so the MoE output is still combined by the FFN all-reduce (vLLM --enable-expert-parallel); under dp_attention the MoE layers dispatch + combine with all-to-alls over the ep group instead |
dp_attention | Bool | false | Attention runs data-parallel over the tp ranks (sglang --enable-dp-attention): the attention projections are replicated (tp× resident and read per step), a sequence’s KV lives on one rank, no attention all-reduce; the TP-sharded FFN gathers the ranks’ tokens with an all-gather and returns them with a reduce-scatter (with ep > 1, DeepEP-style dispatch + combine all-to-alls). Each rank is its own worker: its own scheduler and KV cache over its GPU’s HBM (1/tp of the replica’s), its own port into the node’s [memory] stores; [scheduler] limits apply per rank. The [router] routes flat over ranks — the placement decision a DP-aware front end (Dynamo’s (worker, dp_rank) endpoints) makes. The ranks step in lockstep as one group: the step is priced on the replica-wide roofline over the union batch plus the attention skew, the slowest rank’s own-GPU attention time over the mean rank’s (the ranks meet at every layer’s FFN collective). Speculative decoding is not modelled with rank groups |
replicas | U32 | 1 | Identical workers of this deployment (each a tp-GPU replica with its own scheduler and KV cache; tp rank workers each under dp_attention) behind the router |
spec | String or Table | the entry name | Another catalog preset, or an inline hardware table (fields below) |
scheduler | Table | {} | Keys merged over the shared [scheduler] for this entry |
speculative | Table | shared [speculative] | Replaces the shared block for this entry |
router | Table | shared [router] | Replaces the shared block for this entry |
decode_router | Table | shared [decode_router] | Replaces the shared block for this entry |
memory | Table | shared [memory] | Replaces the shared block for this entry |
prefill | Table | shared [prefill] | Replaces the shared block for this entry |
time_correction | Table | unset | Step-time calibration against a measured engine: { alpha = 1.0, beta = 0.0 } prices every roofline step at alpha × t + beta seconds (alpha = kernel-efficiency gap, beta = fixed per-iteration cost). Unset = pure roofline. Never applied on top of a measured step-cost table |
Collectives are priced on the hardware’s [fabric] (below) and added
serially to the step; an entry with tp > 1 or ep > 1 on hardware without
one is rejected. Per layer: attention → one all-reduce over tp (none under
dp_attention); dense FFN, or MoE without dp_attention (any ep) → one
all-reduce; under dp_attention, dense FFN or MoE with ep = 1 → an
all-gather + reduce-scatter, MoE with ep > 1 → dispatch and combine
all-to-alls over ep, each rank moving its tokens / ep share of
experts_per_tok hidden vectors. Expert reads and FLOPs are taken as
balanced across ranks; there is no overlap of collectives with compute.
Hardware spec
Per-GPU physical spec: a catalog preset (catalog/hardware/<name>.toml) or
an inline spec table.
| Field | Type | Default | Description |
|---|---|---|---|
name | String | — | Accelerator name |
flops_fp4 | Float | unset | Dense FLOP/s at FP4. Unset means the hardware has no FP4 rate; a model that declares an FP4 stream then fails at run time |
flops_fp8 | Float | unset | Dense FLOP/s at FP8 |
flops_bf16 | Float | unset | Dense FLOP/s at BF16 |
flops_fp16 | Float | unset | Dense FLOP/s at FP16 (FP32 is taken as half of it) |
memory_bandwidth | Float | — | HBM bandwidth, bytes/s |
memory_capacity | U64 | — | HBM capacity, bytes |
fabric | Table | unset | Collective fabric, see below |
memory | Table | unset | KV memory beyond HBM this node class offers (stores and links), see below. A deployment picks tiers from it with its own [memory] |
[fabric]
| Field | Type | Default | Description |
|---|---|---|---|
gpus_per_node | U32 | — | GPUs sharing one scale-up domain; parallel groups are packed node by node |
scale_up | Table | — | Inside a node (NVLink / NVSwitch): { bandwidth, latency, in_network_reduction } — per-GPU injection bytes/s per direction, seconds per collective call, and whether the switch reduces in-network (NVLink SHARP) |
scale_out | Table | unset | Across nodes, rail-optimised (GPU i drives NIC i): same fields. Required for a group wider than gpus_per_node |
Cost per collective, added serially to the step (no overlap): a TP
all-reduce of V bytes over g ≤ gpus_per_node ranks is latency + f·V / bandwidth with f = 2(g−1)/g (ring) or 1 (in-network reduction); over
more ranks it is reduce-scatter and all-gather inside each node around an
all-reduce of the V/k shard across the n nodes on the NIC. An EP
all-to-all moves each rank’s (g−1)/g share at the scale-up rate, or, across
nodes, its in-node and cross-node shares concurrently on their own links.
dp_attention skips the per-layer all-reduce.
[memory] on the hardware
The stores a node offers for KV beyond HBM and the links that reach them,
templated per GPU or per node: a graph. Presets ship one (host DRAM and a
node NVMe pool behind each GPU’s PCIe port on the HGX parts; Grace LPDDR5X
behind NVLink-C2C on gh200; NVLink ports into the node switch and one NIC
per GPU to the network on all); an inline hardware table can declare its
own.
| Field | Type | Default | Description |
|---|---|---|---|
gpus_per_node | U32 | fabric’s, else 1 | GPUs sharing a node’s per = "node" stores |
stores | Array | [] | { name, per = "gpu" | "node" | "cluster", kind = "store", capacity, bandwidth, stripe = 1, aggregate_bandwidth, latency = 0, pin = false } — bytes per instance; a gpu store is private to its GPU, a node store is one pool for the node’s GPUs, and a cluster store is one topology-wide pool behind the network. bandwidth is the store throughput per instance (per node for a cluster store). A cluster transfer can use stripe × bandwidth, while all its transfers share aggregate_bandwidth or, by default, nodes × bandwidth. latency is the fixed access cost in seconds on every fetch or write. The reserved { name = "peer_hbm", per = "node", kind = "peer_hbm" } entry is virtual: no capacity/bandwidth, read-only, and pin defaults to true |
junctions | Array | [] | { name, per } — a point with no capacity of its own, so that several links can share one (a GPU’s PCIe port feeding host DRAM and NVMe) |
links | Array | [] | { name, from, to, bandwidth, latency = 0 } — from is "gpu" (one port per GPU), "network", a store or a junction; to is a store, a junction, "switch" (the node’s scale-up fabric) or "network" (the scale-out core). One instance per instance of from, full duplex at bandwidth bytes/s each way |
Instances: a tp-GPU worker pools its GPUs’ per-GPU stores, junctions and
ports (capacity and bandwidth × tp); gpus_per_node / tp workers share a
node’s per-node instances; every worker points at the same selected cluster
store. A cluster access runs GPU → its NIC → network → store (and back), so
the NIC remains the per-worker cap even when the store is striped. A transfer
takes the shortest hop path between its ends and runs at its max-min fair
share on every edge of it: the most contended edge fixes its transfers’ rate
first, the residual is shared among the rest. Tier promotions run store →
GPU; peer_hbm promotions run
from the selected sibling GPU through the node switch to the destination
GPU; hand-offs run GPU → network → GPU (see kv_link_bw).
Shipped presets, at datasheet figures: b200 (192 GB / 8 TB/s), b300
(288 GB / 8 TB/s), gh200 (96 GB / 4 TB/s), h100 (80 GB / 3.35 TB/s);
each carries its node’s fabric (8-GPU NVSwitch + CX-7/CX-8 for the HGX
boxes, 4-GPU NVLink + Slingshot for GH200).
[model]
A model is described by its per-token weight streams and its token-mixing layer classes; there are no named architectures. Any transformer the simulator serves is a composition of the pieces below.
| Field | Type | Default | Description |
|---|---|---|---|
name | String | — | |
hidden_dim | U32 | — | Residual width (sizes collectives; prices MLA attention when its head shape is not given) |
max_seq_len | U32 | — | Architecture context limit (max_position_embeddings) |
attention_precision | Precision | "bf16" | Rate the attention score / AV matmuls run at; KV reads charge this stream |
activation_bytes | U32 | 2 | Bytes per activation element on the wire |
weights | Array | — | One or more weight streams (below) |
layers | Array | — | One or more layer classes (below) |
[[weights]] — a per-token GEMM stream at one precision
| Field | Type | Default | Description |
|---|---|---|---|
precision | "fp4","fp8","bf16","fp16","fp32" | — | |
active_params | U64 | — | Parameters touched per token (FLOPs = 2×) |
resident_params | U64 | — | Parameters resident in HBM |
routing | Table | unset | MoE routing: { routed_experts, experts_per_tok, moe_layers }. The per-step read then follows coupon-collector growth with the step’s tokens (per-expert and shared params are recovered from the active/resident split); EP all-to-alls (under dp_attention) = 2 × moe_layers |
A dense fp8 model is one stream; DeepSeek-V4 is an fp4 expert stream with routing plus an fp8 non-expert stream; gpt-oss is fp4 experts + bf16 rest.
[[layers]] — a class of identical token-mixing layers
kind = "attention" — GQA / MHA with a growing KV cache:
| Field | Type | Default | Description |
|---|---|---|---|
count | U32 | — | Layers in the class |
heads, head_dim | U32 | — | Query heads and head width (attention FLOPs = 4 × heads × head_dim per query-key pair) |
kv_heads | U32 | — | KV heads (KV per token = 2 × kv_heads × head_dim × bytes) |
kv_shared | Bool | false | K and V share one tensor (Gemma-4): half the KV |
window | U32 | 0 | Sliding window: attend to and store only the last window tokens; 0 = full context |
kv_precision | Precision | — |
kind = "mla" — multi-head latent attention, optionally sparse:
| Field | Type | Default | Description |
|---|---|---|---|
count | U32 | — | |
latent_dim, rope_dim | U32 | — / 0 | KV per token = (latent + rope) × bytes |
kv_precision | Precision | — | |
window | U32 | 0 | Recent tokens attended directly. Without a history path, 0 means the whole context; with one, 0 means no local window |
history | Table | unset | Long-range path { compress_ratio, index_topk, indexer }: the history at stride compress_ratio (1 = every position), all of it or the index_topk entries an indexer = { heads, head_dim, kv_precision, precision } selects (the indexer scores every entry and keeps its own KV; precision is the scoring GEMM’s rate, default attention_precision — DeepSeek-style indexers score in fp8) |
heads, qk_head_dim, v_head_dim | U32 | unset | Head shape for the score/AV FLOP count (2 × heads × (qk + v) per pair). All three or none; absent = 4 × hidden_dim per pair |
q_latent_dim, o_latent_dim | U32 | unset | Low-rank query / output projections (q_lora_rank, o_lora_rank); size the attention projections replicated under dp_attention. Absent = full-rank |
Kimi-K2 is one mla class (full context); DeepSeek-V4 is three (window 128
only; window + top-k of the ÷4 history with an indexer; window + the whole
÷128 history); GLM-5 is top-2048 over the uncompressed history.
kind = "linear" — linear attention / SSM (GatedDeltaNet, Mamba, KDA):
| Field | Type | Default | Description |
|---|---|---|---|
count | U32 | — | |
state_bytes | U64 | — | Fixed per-sequence state per layer (reserved for the sequence’s lifetime and read once per step); no context-scaling work |
Shipped model presets: catalog/models/ (inference_lab::catalog::model_names()),
each with the derivation of its numbers from the HF config in its header.
[scheduler]
| Field | Type | Default | Description |
|---|---|---|---|
max_num_batched_tokens | U32 | — | Token budget per iteration |
max_num_seqs | U32 | — | Running-request cap. In serve mode this is retunable at runtime — see Saturation and capacity |
max_waiting | U32 | 0 | serve only: refuse arrivals with HTTP 529 once this many requests are waiting. 0 = unbounded (queue without limit, never refuse). See Saturation and capacity |
policy | String | — | fcfs, priority, sif, lif, sof, lof, stf, ltf (sjf = sof) |
enable_chunked_prefill | Bool | — | Split long prefills across iterations |
long_prefill_token_threshold | U32 | 0 | Prefill chunk cap; 0 = no cap. Defaults to 4% of max_seq_len when max_num_partial_prefills > 1 |
max_num_partial_prefills | U32 | 1 | vLLM’s knob; only its effect on the threshold default is modelled |
block_size | U32 | — | KV block size, tokens |
gpu_memory_utilization | Float | 0.9 | Fraction of GPU memory the engine may use (vLLM’s --gpu-memory-utilization); the KV cache gets what is left after the weights |
kv_cache_capacity | U64 | 0 | Explicit KV cache bytes across the TP group; 0 derives it from gpu_memory_utilization |
max_model_len | U32 | model’s max_seq_len | Serving-time context limit (only the chunked-prefill threshold default depends on it) |
enable_preemption_free | Bool | false | Admit only what can grow to prompt + max_output without preemption |
balance_set | Table | unset | Balance-set admission control (Denning’s medium-term scheduler): { high, low } as fractions of KV capacity. Admission stops when the running working set (resident context of running requests) reaches high and resumes only once it falls below low (hysteresis; low defaults to high). Holds the overflow in the queue instead of admitting-and-evicting, so recently-idle sessions’ cached prefixes survive in the reserved 1 − high headroom. Absent = overcommit (today’s behaviour) |
enable_cascade_attention | Bool | false | Load a batch’s shared prompt prefix once per iteration |
Workload file
The [workload] table, at top level of workloads/<name>.toml.
| Field | Type | Default | Description |
|---|---|---|---|
arrival_pattern | String | — | poisson, uniform (= fixed_rate), burst, closed_loop, batched |
arrival_rate | Float | 1.0 | Requests/s for the open-loop patterns |
rate_schedule | Table | unset | Time-varying rate: { type = "sine", min, max, period_secs }, { type = "square", low, high, period_secs, duty }, or { type = "trace", points = [[t, rate], ...] } |
num_concurrent_users | U32 | unset | Users for closed_loop |
closed_loop_jitter_secs | Float | unset | Uniform stagger of the initial closed-loop arrivals |
input_len_dist, output_len_dist | Table | — | { type = "fixed", value }, { type = "uniform", min, max }, { type = "normal", mean, std_dev }, { type = "lognormal", mean, std_dev } (ignored in dataset mode for input) |
num_requests | U32 | unset | Maximum total generated requests; in session mode every step counts (use num_sessions to bound starts instead) |
duration_secs | Float | unset | Stop admitting arrivals after this many simulated seconds, then drain requests already in flight |
dataset_path | String | unset | JSONL in OpenAI batch format; prompts are tokenised with --tokenizer and hashed per KV block so shared prefixes hit the prefix cache |
sessions_path | String | unset | Session file (JSONL, one session per line, see Sessions). The arrival pattern then governs session starts (arrival_rate in sessions/s; closed_loop holds num_concurrent_users sessions in flight); each later step arrives at its parent’s completion plus the step’s gap. Mutually exclusive with dataset_path; length distributions are ignored |
num_sessions | U32 | unset | Session mode: maximum session starts; does not limit the total request steps emitted by those sessions (the file is cycled, so it may exceed the file’s count) |
stationary_start_sessions | U32 | unset | Session mode: start this many sessions at t=0 at a time-weighted step in their trace, with fresh inherited-context hashes, then resume the open-loop arrival clock |
resample_sessions | Bool | false | Session mode: draw each session uniformly from the file with replacement instead of walking it in order; repeated instances receive fresh block hashes |
seed | U64 | — |
[memory]
Optional; no tiering when absent. Which of the hardware’s stores hold KV
evicted from HBM, closest first, and how much of each. A worker (a
tp-GPU replica) reaches each tier over its own link — its GPUs’ ports
pooled, so a tp = 2 worker on gh200 promotes at 2 × 450 GB/s from 2 ×
120 GB of Grace memory. A per = "node" store is shared by the workers on
that node (gpus_per_node / tp of them): what one demotes, its neighbours
can promote. A worker wider than a node pools the node stores it spans.
[memory]
tiers = ["peer_hbm", "host_dram", "nvme"] # peer_hbm is declared by b200
preset = "reactive" # reactive | oracle (optional bundle)
source = { policy = "promote" } # promote | min_time
hbm_eviction = { policy = "lru" } # lru | outlook
write = { policy = "write_back" } # write_back | write_through | selective | live
eviction = { policy = "fifo" } # fifo | lru | ttl | outlook
prefetch = { policy = "none" } # none | outlook
backup = "on_evict" # on_evict | on_land
hit_refresh = "first_tier" # first_tier | none
promote_fill = "through" # through | buffer | direct
storage_prefetch = { policy = "wait_complete" } # wait_complete | best_effort | timeout
load_overlap = "layerwise" # layerwise | none
hbm_evict_backed_first = false
[memory.capacity]
host_dram = 1.0e12 # bytes per instance given to KV
| Field | Type | Default | Description |
|---|---|---|---|
tiers | Array | [] | Tier names from the hardware’s [memory], closest first. Each must be reachable from a GPU over the hardware’s links. peer_hbm consults same-node sibling HBM and has no capacity or write step |
capacity | Table | full | Per-store cap on bytes per instance |
preset | String | unset | A named policy bundle; any field set explicitly overrides the preset’s choice. reactive: promote / lru / selective (min_hits 1) / lru / none / backed-first — decides only from what has already happened, as shipped stacks do. oracle: min_time / outlook / live / outlook / outlook / backed-first — reads every session’s announced re-entry. Both use staged through reads, wait for storage prefetch, and overlap the final load layer-wise. The preset covers KV movement only: on a pool with several workers or DP-attention ranks, pair it with a [router] that sends a re-entry to the worker holding (or prefetching) its prefix — kv_aware or prefix_affinity — or its evictions and prefetches serve arrivals that land elsewhere |
source | Table | promote | Where a re-entry’s tier-held prefix comes from: promote — fetch it (a hit is a hit); min_time — fetch it only if the transfer, at the fetch path’s current fair share, beats recomputing those tokens at the worker’s roofline; otherwise recompute (the tier keeps its copy) |
hbm_eviction | Table | lru | Which free HBM block is recycled first: lru — least recently freed; outlook — blocks with no announced re-entry first (LRU among them), then the farthest re-entry first, each sequence tail first |
write | Table | write_back | When a block’s KV is written to the first tier: write_back — when its HBM block is recycled, if no tier holds it; write_through — as soon as it is produced; selective (min_hits, default 1) — on its min_hits-th HBM hit, and dropped on eviction otherwise (SGLang HiCache’s three positions); live — when recycled, only if its session has announced a re-entry (a finished trajectory is dropped) |
eviction | Table | fifo | How every tier picks what to recycle: fifo (least recently inserted), lru (least recently inserted or promoted from), ttl (seconds: LRU, and any block untouched that long is dropped whether or not the store is full), outlook (no announced re-entry first, then farthest re-entry first). Stores hold ranges of a sequence, written and stamped together; a victim range is recycled from its tail under every policy, so what survives of a sequence in a store is a prefix |
prefetch | Table | none | Whether a demoted prefix is pulled back ahead of its re-entry: none; outlook (lead, default 0 s) — when a session step completes, plan a promotion of the prefix its next step re-enters with, starting so it lands lead seconds before that arrival at the fetch path’s fair share at planning time; whatever is still in HBM when the plan fires needs nothing, and a re-entry that arrives mid-transfer joins it |
backup | String | on_evict | When a tier forwards a block to the tier below it: on_evict — only when it evicts the block (a store → store transfer of what would otherwise be dropped); on_land — as soon as the block’s write into it lands, so every tier below the first receives a copy within a transfer of production (SGLang HiCache backs a node up to its storage backend the moment its device → host DMA completes). Under on_land a private per-rank host tier’s fresh KV reaches a shared storage tier at once rather than when the host tier ages it out |
hit_refresh | String | first_tier | Whether a prefix hit in HBM re-stamps tier copies as recently used: first_tier — the first tier below HBM ages with HBM (HiCache’s device and host tiers share one radix tree and one last_access_time; lower tiers see only the references that reach them); none — tier copies are re-stamped only when promoted from |
promote_fill | String | through | How a lower-tier read reaches HBM: through — transfer into each closer store in turn, publishing and retaining each copy before a separate final HBM load; buffer — use the same real staged transfers but release their intermediate copies after the final load; direct — one legacy source-to-HBM transfer with no staged cache fill |
storage_prefetch | Table | wait_complete | What a demand does while an external-store hit is staging: wait_complete — wait for the stage; best_effort — stage only while queued, then cancel an unfinished leg when an admission slot is available; timeout (seconds, finite and > 0) — wait no longer than the deadline. Cancellation keeps earlier completed stages, uses any prefix already in the closest tier, and recomputes the still-external suffix |
load_overlap | String | layerwise | Whether the final closest-store-to-HBM load pipelines with the request’s first prefill pass: layerwise — aggregate approximation with elapsed time max(load, compute); none — the full load completes before compute begins |
hbm_evict_backed_first | Bool | false | When HBM must recycle a block, take one whose KV a tier already holds (a free drop) over the policy’s first choice, looking 16 blocks up the free queue |
The outlook, live and min_time/prefetch policies act on a session
step’s outlook: on a session workload, when a step completes the
simulator knows its successor’s arrival (completion plus the recorded gap)
and how much of the context it re-enters with, and marks those blocks —
in HBM and in the tiers — with that time. Other workloads announce
nothing, so under live nothing is written and outlook eviction reduces
to LRU. The gap between the reactive and oracle presets on the same
replay is the value of knowing the re-entry.
Under promote_fill = "through", tiers are inclusive: a promoted block keeps
every tier copy it staged through (KV is immutable), so its next eviction from
HBM is a free drop and only blocks no tier holds ever cost a write. buffer
keeps only the backing source; direct changes no store residency. Writes are
transfers GPU → store on the same
graph as promotions (full duplex: they share a port with fetches only in
the reverse direction, but do share an NVMe pool’s drives); a block is
resident once its write lands and a promotion of a block still arriving
waits for it (the write-before-reuse race). A full store evicts its victim
into the next tier as a store → store transfer, or drops it; a block
dropped without ever having been promoted counts as dead bytes. A prompt
whose blocks sit below the closest store is staged upward one store at a time,
sharing every traversed edge with whatever else is in flight. These legs hold
no HBM. The final closest-store-to-GPU leg reserves the request’s landing
blocks; direct mode instead reserves them for its one source-to-GPU transfer.
A staged leg is admission-controlled against its destination store, the way
HiCache bounds a storage prefetch by host-pool free space: it starts only if
the store can take its blocks by evicting nothing that is pinned, and on
start it pins its destination copy — in flight and after landing — until the
request’s final HBM load consumes it (or the stage is cancelled, times out,
or is abandoned). A request whose stage does not fit yet holds at the head
of its worker’s queue until a stage completion or consumption frees pinned
room; one whose external prefix exceeds the destination store outright
abandons the prefetch and recomputes the external suffix. Under
storage_prefetch = timeout, the deadline starts when the stage starts, not
while it waits for room.
pin is set on each hardware store entry. Staged reads always protect their
store source until that leg drains. On a direct read, pin = true prevents
the source from being recycled; with pin = false, completion lands only the
prefix still present, releases the unused HBM reservation, and recomputes the
suffix. Peer HBM defaults to pinned; normal stores preserve the unpinned
direct-read default.
The summary’s memory section reports, per store name, blocks held, bytes
written / read / dead, evictions and expiries; per link name, bytes moved
and utilisation; and totals of bytes written, bytes promoted, peer-HBM bytes
promoted, pin stalls, partial landings, and promotions that waited on a
write. The prefix_cache section counts
lookups recomputed instead of fetched (min_time) and prefetches started
(prefetch = outlook), with their tokens.
[prefill]
Optional. A disaggregated topology: this block is the prefill pool, and
the hardware entry — its hardware, tp/ep/dp_attention, replicas
and [memory] — becomes the decode pool. Arrivals enter the prefill pool
through [router], prefill there against that pool’s HBM and [prefill] memory tiers, and hand their KV to a decoder chosen by [decode_router]
over the network: the prefill worker’s nic link, the core, the decode
worker’s nic. The first token rides with the hand-off. KV moves one way:
a decoder’s HBM shortens later hand-offs of the same prefix, but nothing
flows back to the prefill side, so a re-entry whose prefix exists only in
decode HBM is recomputed by prefill. A per = "cluster" store named in
both pools’ tiers is the one shared tier.
[prefill]
replicas = 2 # prefill workers
parallel = { tp = 8, dp_attention = true }
memory = { tiers = ["host_dram", "nvme"], preset = "reactive", backup = "on_land" }
kv_link_bw = 4e11 # optional core cap, bytes/s
| Field | Type | Default | Description |
|---|---|---|---|
hardware | String or Table | the entry’s hardware | Catalog preset or inline hardware table for the prefill pool |
parallel | Table | the entry’s tp/ep/dp_attention | Parallel layout of a prefill worker |
replicas | U32 | 1 | Prefill workers behind [router] |
memory | Table | {} | KV tiers on the prefill side, from its hardware’s stores (same keys as [memory]). Absent: prefill keeps prefixes in HBM only |
kv_link_bw | F64 | unset | Capacity of the network core between the pools, bytes/s, shared by every hand-off in flight. Unset: the NICs alone bound them; an error when the hardware has no network links either |
backup, hit_refresh and promote_fill are graph-wide and taken from
the pools that have tiers. [scheduler] (including balance_set) applies
to both pools.
[router] and [decode_router]
Optional; round_robin when absent. [router] picks the replica each
arriving request enters (replicas on the hardware entry; the prefill
pool on a disaggregated topology). [decode_router] picks the decode
worker each hand-off goes to on a disaggregated topology, and defaults to
[router]. The KV-reading policies look up each replica’s state for the
prompt on every decision — an estimate from the replica’s block index, as
a KV-aware front end sees it, not the scheduler’s admission-time lookup.
| Field | Type | Default | Description |
|---|---|---|---|
policy | String | "round_robin" | round_robin, least_loaded, prefix_affinity, kv_aware, kv_aware_decode |
max_load_ratio | F64 | unset | prefix_affinity only: pass over the prefix holder for the least-loaded replica when its requests in system exceed max_load_ratio × the pool mean (bounded-load affinity) |
load_weight | F64 | 1.0 / 64.0 | kv_aware: weight on the replica’s queued prefill tokens. kv_aware_decode: tokens of transfer one running sequence is worth (default one 64-token block) |
round_robincycles through the replicas.least_loadedpicks the fewest requests in system (running + waiting), ties by queued prefill tokens, then index.prefix_affinitypicks the replica holding the longest cached prefix of the prompt (any tier); with none anywhere it falls back toleast_loaded.kv_awareminimises(prompt − cached prefix) + load_weight × queued prefill tokens: the prefill work the request adds plus the prefill work already ahead of it, in tokens. A prefill-side policy: on a decode pool the load term is always zero.kv_aware_decodeminimises(context − prompt prefix resident in the decoder's HBM) + load_weight × running sequences: the KV the hand-off must move plus the decode batch it joins, in tokens. Decoders whose free KV cannot hold the incoming context are passed over while any can.
The summary’s router section (and decode_router on a disaggregated
topology) reports requests per replica and, for the KV-reading policies,
how many decisions had a cached prefix on some replica, how many went to
a holder, and how many went away from the longest holder. handoff
reports transfers, bytes moved, and bytes skipped because the chosen
decoder already held the prefix. For session workloads, reusable_kv
reports the joint prefill/decode residency of reusable KV, derived hit,
recompute and transfer fractions, a parent-prefill, parent-decode, and
inherited-context split of prefiller misses, and per-decode-rank detail. Under
same-rank session affinity, that split separates eviction-driven recomputation
from decoder-output KV that was never written back. hbm reports capacity,
resident prefix bytes, active/reserved bytes, and actual HBM eviction bytes per
worker. sessions reports completed and deadline-censored sessions plus turns
per started session; simulation records the arrival deadline snapshot and
post-deadline drain; work separates logical prompt tokens from positions
actually computed by prefill.
[speculative]
Optional. Decode steps then verify 1 + draft positions and advance by
1 + accepted.
| Field | Type | Default | Description |
|---|---|---|---|
gamma | U32 | — | Draft length (fixed) or maximum candidate depth (budget policies) |
acceptance | Table | — | { kind = "constant", alpha }, { kind = "per_position", a = [...] }, or { kind = "trace_rounds", path } (CSV bank of real rounds: commits,category,a0..aD-1) |
policy | String | "fixed" | fixed, goodput_budget, gated_budget, gated_aggregate |
measured_cost | Table | unset | { path, ref_seq_len }: measured (batch_size, num_draft_tokens, step_seconds) grid that prices decode steps and the policy’s cost curve instead of the roofline |
switch | Table | unconstrained | { cooldown_rounds, max_step, cost_ms } for gated_aggregate |
drafter | Table | free drafter | { kind = "fraction", frac }, { kind = "autoregressive", dense_params, expert_params, num_experts, experts_per_tok, shared_experts }, or { kind = "block_parallel", params, block } |
Fault Injection
The serve mode can kill a streaming chat completion mid-generation in every way a real downstream does — on demand, deterministically, per request. It exists as the test double for gateway resilience work (mid-stream continuation/resume middleware): if the sim can produce every death signature, that middleware can be e2e-tested without waiting for real incidents.
Every faulting stream first emits real partial output — the initial role frame plus
after_chunks content-bearing delta frames of deterministic placeholder text — so there
is always something to resume. The fault path bypasses the simulation engine (like
echo-directives) and contains no randomness: the same trigger produces the same frames
and the same death at the same byte, every time.
Scope: streaming POST /v1/chat/completions and streaming POST /v1/completions.
A fault header on a non-streaming request is rejected with a 400
(invalid_fault_directive) — never silently ignored. With no trigger present, behavior is
completely unchanged.
Both endpoints matter because a mid-stream-continuation resume leg is a streaming
/v1/completions request, so chain-resume tests need to kill one. The frames follow the
endpoint’s own wire shape: /v1/completions emits text_completion chunks carrying
text and has no role frame, so after_chunks=N puts exactly N frames on the wire
before the death (one fewer than the chat flavor, which leads with its role frame).
mid_reasoning and mid_tool_call have no delta object to use there, so they stream the
same partial payloads as raw text — an unterminated <think> block and a tool call the
model never finished writing — which is how both actually appear on a real base-model
completions stream.
Trigger: the x-inference-lab-fault header
The header is only honored when the server runs with --enable-directives — it is
client-controlled and can stall or abort connections at will, so it sits behind the same
“untrusted clients must not reach this server” gate as echo-directives. Without the flag
the header is a 400 (invalid_fault_directive), never silently ignored. The staging
deployment already sets the flag.
x-inference-lab-fault: <mode>[;after_chunks=<u32>][;delay_ms=<u64>][;utf8=<bool>]
| Parameter | Default | Meaning |
|---|---|---|
<mode> | required | one of the eleven mode names below |
after_chunks | 3 | content-bearing delta frames emitted before the fault fires (the initial role-only frame is always sent and not counted) |
delay_ms | 10 | fixed pacing between frames, milliseconds |
utf8 | false | cut_mid_frame only: cut inside a multi-byte UTF-8 character |
Parts are ;-separated; whitespace around parts and = values is ignored. Unknown modes,
unknown keys, or malformed values are a 400 listing the valid modes.
A header was chosen over a body extension because the platform’s proxy layer runs strict request sanitization that strips unknown body fields; headers pass through. Verify this empirically through the full stack (client → dwctl → onwards → sim) when the branch reaches a preview environment — full-platform verification is out of scope here and happens after deploy.
Fallback trigger: static per-model config
For clients that cannot set a header, a model’s TOML config can apply one fault to every
streaming chat completion on that model (non-streaming requests are served normally;
an explicit header on the request still wins). Unlike the header, this is operator input
validated at server boot, so it does not require --enable-directives:
[fault]
mode = "cut_mid_frame" # same names as the header
after_chunks = 5 # optional, default 3
delay_ms = 10 # optional, default 10
utf8 = true # optional, cut_mid_frame only
An invalid [fault] block fails server startup, not individual requests.
Precedence per request: header > model [fault] config > (echo-directives >) normal path.
Modes
| Mode | After the N content frames… | Client observes (curl) |
|---|---|---|
cut_between_frames | connection closes on a frame boundary, without the chunked-encoding terminator (FIN) | exit 18, transfer closed with outstanding read data |
cut_mid_frame | half of the next frame’s bytes, then close — torn JSON. utf8=true cuts one byte into a 2-byte UTF-8 character (é) in the delta text | exit 18, partial data: line |
reset | abortive close: SO_LINGER=0 then drop → TCP RST, not FIN | exit 56, connection reset by peer |
stall | nothing, forever; connection stays open until the client gives up | exit 28 (client timeout) |
error_envelope_200 | OpenRouter-style error envelope as an SSE data frame, then [DONE]; HTTP status stays 200 | exit 0, {"error":{"message":…,"code":502,"metadata":{…}}} |
error_400_in_sse | vLLM-style 400 object in-stream (the nemotron-incident signature), then clean close, no [DONE] | exit 0, {"object":"error",…,"code":400} |
no_done | finish_reason frame, then the separate choices: [] usage frame if requested, then clean close — [DONE] never comes | exit 0, stream just ends |
no_usage | finish_reason frame then [DONE], but the usage frame never arrives (even when stream_options.include_usage was set) | exit 0, usage missing |
cancelled_499 | the exact dynamo frontend-cancellation body, then clean close | exit 0, {"error":{"code":499,"message":"CancelledError: ","type":"request_cancelled"}} |
mid_reasoning | frames carry reasoning_content deltas instead of content; dies cut_between_frames-style | exit 18, last deltas are reasoning |
mid_tool_call | frame 1 announces a tool call (id + name), later frames stream arguments fragments that never terminate; dies cut_between_frames-style | exit 18, partial tool call |
Notes:
- The error body shapes for
error_envelope_200anderror_400_in_sseare representative; exact shapes sync with the death-taxonomy workstream as it lands. Thecancelled_499body is exact. delay_ms=0is fine for the graceful modes, but the abrupt modes always wait a short flush grace (~25 ms) before killing the connection so the partial output reliably reaches the wire first.resetneeds the raw socket, which the server threads through per-connection; on non-unix platforms (or when handlers are driven outside the real server, as in unit tests) it degrades to a FIN with a warning log.
Examples
All against a local sim (inference-lab serve --config configs/ --hardware b200 --port 8080); $BODY is
any streaming chat request:
BODY='{"model":"DeepSeek-V4-Flash","stream":true,"stream_options":{"include_usage":true},"messages":[{"role":"user","content":"hello"}],"max_tokens":16}'
URL=http://localhost:8080/v1/chat/completions
# 1. cut_between_frames — 5 frames then FIN (curl exit 18)
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cut_between_frames;after_chunks=5' -d "$BODY"
# 2a. cut_mid_frame — torn JSON frame
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cut_mid_frame' -d "$BODY"
# 2b. cut_mid_frame, cut inside a multi-byte UTF-8 character (pipe through xxd to see it)
curl -sN $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cut_mid_frame;utf8=true' -d "$BODY" | xxd | tail
# 3. reset — TCP RST (curl exit 56, "connection reset by peer")
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: reset' -d "$BODY"
# 4. stall — 3 frames then silence; bound the wait client-side
curl -N --max-time 10 $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: stall' -d "$BODY"
# 5. error_envelope_200 — OpenRouter-style envelope then [DONE]
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: error_envelope_200' -d "$BODY"
# 6. error_400_in_sse — 400-shaped object inside the 200 stream
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: error_400_in_sse' -d "$BODY"
# 7. no_done — finish_reason + usage frame, then the stream ends without [DONE]
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: no_done' -d "$BODY"
# 8. no_usage — [DONE] arrives but the requested usage never does
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: no_usage' -d "$BODY"
# 9. cancelled_499 — exact dynamo cancellation body
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: cancelled_499' -d "$BODY"
# 10. mid_reasoning — dies while streaming reasoning_content deltas
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: mid_reasoning' -d "$BODY"
# 11. mid_tool_call — dies with a tool call's arguments unterminated
curl -N $URL -H 'content-type: application/json' -H 'x-inference-lab-fault: mid_tool_call;after_chunks=4' -d "$BODY"
Killing a resume leg
A resume leg is a streaming /v1/completions request whose prompt is token ids. Every
mode above works against it; only the frame shape differs (no role frame, text instead
of a delta):
RESUME='{"model":"DeepSeek-V4-Flash","prompt":[1,2,3,4],"stream":true,"priority":0,"stream_options":{"include_usage":true},"max_tokens":64}'
# exactly 2 text frames, then FIN (curl exit 18)
curl -N http://localhost:8080/v1/completions -H 'content-type: application/json' \
-H 'x-inference-lab-fault: cut_between_frames;after_chunks=2' -d "$RESUME"
Saturation and Capacity
serve mode queues arrivals without bound by default: there is no request
volume at which it refuses work. That is fine for latency experiments, but it
makes the server useless for testing a client that is supposed to react to
overload — a controller that discovers a model’s sustainable concurrency from
downstream rejections has nothing to react to, because no rejection ever
arrives.
Two knobs change that, and they do different things:
| Knob | What it changes | Typical use |
|---|---|---|
max_waiting | How long the queue may get before the server refuses | Switch 529s on and off |
max_num_seqs | How fast the engine drains work | Model a scale up/down |
max_waiting does not change capacity. Lowering it makes rejections start
sooner, but the engine serves at exactly the same rate, so the concurrency a
client settles at is set by queue policy rather than by anything physical.
max_num_seqs is the dial that changes the drain rate: halve it and the queue
backs up at roughly half the offered load.
The waiting bound
Set it in a model’s [scheduler] table:
[scheduler]
max_num_seqs = 256
max_waiting = 64 # 0 (the default) = unbounded, never refuse
or override every model’s value at the command line:
inference-lab serve --config configs/ --hardware b200 --max-waiting 64
Once max_waiting requests are queued, further arrivals get:
HTTP/1.1 529
content-type: application/json
{"error": {
"message": "Server is at capacity: 64 requests waiting (max_waiting = 64). Retry with reduced concurrency.",
"type": "overloaded_error",
"code": "queue_saturated"
}}
Why 529, and not 503 or 429
529 is the convention for “the engine has nowhere to put this request”.
Clients that adapt their concurrency generally key on 529 alone, because the
other two mean something a client should not answer by shedding load: 503 is
“this service is unavailable” and 429 is “you exceeded a quota or a proxy’s own
limit”. serve already returns 503 when the engine channel is closed, which is
a liveness failure rather than saturation, and stays a distinct code.
Two properties the rejection is built to have
It is returned before the response starts. A real engine that admits a
request, sends 200 plus streaming headers, and only then discovers it cannot
schedule it has spent its status code: the failure can then only be an error
object inside the stream, or a stream that stops with no content. Neither is
classifiable as overload, so the client learns nothing. The check therefore
runs in submit_engine_request, the single funnel into the engine, before any
handler has built a response — a refused request returns a bare status and
envelope, never a text/event-stream.
It is immediate, never a stall. A bounded queue that parks requests until some later timeout produces a client that waits and eventually gives up, which consumes a client slot for the whole timeout and still carries no overload signal. The bound is a synchronous check against a published queue depth.
What “waiting” counts
The depth compared against max_waiting is every worker’s num_waiting()
(queued requests plus those parked on a KV transfer or a staged read), plus
arrivals the HTTP layer has admitted that the engine has not stepped into a
scheduler yet. The second term matters: Engine::submit only queues an
Arrival event, so without it a burst would read as depth 0 and be admitted
wholesale.
Runtime capacity control
GET /control/capacity reports every model’s knobs and live depth:
curl localhost:8080/control/capacity
[{"model":"gpt-oss-20b","max_waiting":64,"max_num_seqs":256,"waiting":12,"running":256}]
POST /control/capacity retunes them without a restart — which matters
because a restart drops every in-flight request, destroying the before-and-after
that a capacity-change experiment depends on. Both fields are optional, and
model defaults to every loaded model:
# Scale down: the engine now drains at a sixteenth of the rate.
curl -X POST localhost:8080/control/capacity \
-H 'content-type: application/json' -d '{"max_num_seqs": 16}'
# Turn 529s on (or off again with 0) against a running server.
curl -X POST localhost:8080/control/capacity \
-H 'content-type: application/json' -d '{"max_waiting": 64}'
Both changes act on live state:
- Lowering
max_num_seqsdrains, it does not evict. The cap gates admission only, so requests already running above the new cap run to completion and the batch shrinks to the new size. Nothing in flight is lost. - Lowering
max_waitingonly affects requests that have not arrived yet. Anything already queued keeps its place; its status code is long since spent.
max_num_seqs must be at least 1 — a cap of 0 would accept requests and then
never schedule them, which is exactly the stall the bound exists to replace.
max_waiting: 0 is meaningful (unbounded) and allowed.
Cold start: the first burst after idle is unpaced
serve paces simulated time to wall-clock from an epoch fixed when the engine
starts, and simulated time only advances while there is work to do. An idle
server therefore accumulates a deficit, and the first requests after a quiet
period are served as fast as the CPU allows rather than at the rate the model
predicts — a single 200-token stream into a freshly booted server returns in
single-digit milliseconds.
The deficit burns off once there is continuous work: under sustained load the
sim clock catches up within about a second, after which pacing is correct and
steady. Measured on gpt-oss-20b / b200 at 24 concurrent requests, a
200-token response settles at a flat ~293 ms; drop max_num_seqs to 8 and the
same load settles at ~646 ms with throughput down by roughly the same factor.
The practical consequence is only for short experiments: a benchmark that fires one burst at a just-started server measures the CPU, not the modelled hardware. Give it a second of warm-up load first.