Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Onwards

Crates.io Documentation GitHub

A Rust-based AI Gateway that provides a unified interface for routing requests to OpenAI-compatible targets. The goal is to be as “transparent” as possible.

Features

  • Unified routing to any OpenAI-compatible provider
  • Hot-reloading configuration with automatic file watching
  • Authentication with global and per-target API keys
  • Rate limiting per-target and per-API-key with token bucket algorithm
  • Concurrency limiting per-target and per-API-key
  • Load balancing with weighted random and priority strategies
  • Automatic failover across multiple providers
  • Strict mode for request validation, response sanitization, and error standardization
  • Response sanitization for OpenAI schema compliance
  • Prometheus metrics for monitoring
  • Custom response headers for pricing and metadata
  • Upstream auth customization for non-standard providers

Quickstart

Create a configuration file

Create a config.json file with your target configurations:

{
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-your-openai-key",
      "onwards_model": "gpt-4"
    },
    "claude-3": {
      "url": "https://api.anthropic.com",
      "onwards_key": "sk-ant-your-anthropic-key"
    },
    "local-model": {
      "url": "http://localhost:8080"
    }
  }
}

Start the gateway

cargo run -- -f config.json

Modifying the file will automatically and atomically reload the configuration. To disable this, set the --watch flag to false.

Send a request

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Configuration

Onwards is configured through a JSON file. Each key in the targets object defines a model alias that clients can request.

Global options

OptionTypeRequiredDescription
strict_modeboolNoEnable strict mode globally for all targets (see Strict Mode). Default: false
authobjectNoGlobal authentication configuration (see Authentication)
targetsobjectYesMap of model aliases to target configurations

Target options

OptionTypeRequiredDescription
urlstringYesBase URL of the AI provider
onwards_keystringNoAPI key to include in requests to the target
onwards_modelstringNoModel name to use when forwarding requests
keysstring[]NoAPI keys required for authentication to this target
rate_limitobjectNoPer-target rate limiting (see Rate Limiting)
concurrency_limitobjectNoPer-target concurrency limiting (see Concurrency Limiting)
upstream_auth_header_namestringNoCustom header name for upstream auth (default: Authorization)
upstream_auth_header_prefixstringNoCustom prefix for upstream auth header value (default: Bearer )
response_headersobjectNoKey-value pairs to add or override in the response headers
sanitize_responseboolNoEnforce strict OpenAI schema compliance for responses only (see Sanitization)
propagate_trace_contextoptional boolNoInject W3C traceparent / tracestate headers on outbound requests; omit to inherit from the resolved trusted value. Provider-scoped: valid on a single-provider target and on each entry of a pool’s providers array — not as a top-level key on a pool that uses providers. See Trace context propagation.
reasoning_translationobjectNoTranslate canonical OpenAI reasoning efforts into this provider’s request shape. Provider-scoped in load-balanced pools.
strategystringNoLoad balancing strategy: weighted_random or priority
fallbackobjectNoRetry configuration (see Load Balancing)
providersarrayNoArray of provider configurations for load balancing

Reasoning translation

Clients use reasoning_effort on Chat Completions and reasoning.effort on Responses. The complete OpenAI-compatible effort set is none, minimal, low, medium, high, xhigh, and max.

Each configured surface has a required writes array and a required unsupported_efforts array. Every write must map the same supported efforts. The mapped efforts and unsupported_efforts must be disjoint and together explicitly account for all seven values. This makes accepting, collapsing, or disabling an OpenAI effort a conscious provider configuration choice.

For a model with native effort levels, preserve reasoning_effort and reject levels the model does not support:

{
  "targets": {
    "gpt-oss": {
      "url": "https://inference.example.com/v1",
      "reasoning_translation": {
        "chat_completions": {
          "unsupported_efforts": ["none", "minimal", "xhigh", "max"],
          "writes": [{
            "target_path": "/reasoning_effort",
            "values": {
              "low": "low",
              "medium": "medium",
              "high": "high"
            }
          }]
        }
      }
    }
  }
}

For vLLM or SGLang models controlled by an absolute reasoning budget, retain reasoning_effort to activate thinking and add thinking_token_budget:

{
  "chat_completions": {
    "unsupported_efforts": [],
    "writes": [
      {
        "target_path": "/reasoning_effort",
        "values": {
          "none": "none",
          "minimal": "minimal",
          "low": "low",
          "medium": "medium",
          "high": "high",
          "xhigh": "xhigh",
          "max": "max"
        }
      },
      {
        "target_path": "/thinking_token_budget",
        "values": {
          "none": 0,
          "minimal": 512,
          "low": 1024,
          "medium": 4096,
          "high": 8192,
          "xhigh": 12288,
          "max": 16384
        }
      }
    ]
  }
}

Budget values are model-specific and should be established through evaluation; the values above are illustrative, not defaults. Chat Completions requests using a budget mapping must set non-null max_completion_tokens (or legacy max_tokens) above the selected budget. Responses requests must set max_output_tokens above it. A missing limit returns 422; an equal or smaller limit returns 400. Budgets are never silently clipped.

Binary providers can deliberately collapse all enabled efforts to one boolean while still naming every mapping:

{
  "chat_completions": {
    "unsupported_efforts": [],
    "writes": [{
      "target_path": "/thinking",
      "values": {
        "none": false,
        "minimal": true,
        "low": true,
        "medium": true,
        "high": true,
        "xhigh": true,
        "max": true
      }
    }]
  }
}

An omitted client effort injects nothing. If a requested effort is unsupported by any provider in a fallback pool, or its absolute budget is incompatible with the request limit, the request is rejected before an upstream attempt.

Provider-native reasoning controls in client requests, including thinking_token_budget, are rejected. Legacy Completions does not support reasoning controls. Per-model capability discovery through /v1/models is intentionally left to the control layer.

Rate limit object

FieldTypeDescription
requests_per_secondfloatNumber of requests allowed per second
burst_sizeintegerMaximum burst size of requests

Concurrency limit object

FieldTypeDescription
max_concurrent_requestsintegerMaximum number of concurrent requests

Auth configuration

The top-level auth object configures global authentication:

FieldTypeDescription
global_keysstring[]Keys that grant access to all authenticated targets
key_definitionsobjectNamed key definitions with per-key rate/concurrency limits

See Authentication for details.

Minimal example

{
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-your-openai-key"
    }
  }
}

Full example

{
  "strict_mode": true,
  "auth": {
    "global_keys": ["global-api-key-1"],
    "key_definitions": {
      "premium_user": {
        "key": "sk-premium-67890",
        "rate_limit": {
          "requests_per_second": 100,
          "burst_size": 200
        },
        "concurrency_limit": {
          "max_concurrent_requests": 10
        }
      }
    }
  },
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-your-openai-key",
      "onwards_model": "gpt-4-turbo",
      "keys": ["premium_user"],
      "rate_limit": {
        "requests_per_second": 50,
        "burst_size": 100
      },
      "concurrency_limit": {
        "max_concurrent_requests": 20
      },
      "response_headers": {
        "Input-Price-Per-Token": "0.0001",
        "Output-Price-Per-Token": "0.0002"
      },
      "sanitize_response": true
    }
  }
}

Reasoning controls through internal gateways

Non-strict gateways forward provider-native reasoning fields such as chat_template_kwargs.enable_thinking unchanged. This permits an internal gateway to receive requests already translated by a public gateway. Strict mode rejects these fields in favor of canonical reasoning controls. Configured reasoning_effort translations still apply in either mode.

Streaming response buffering

Onwards forwards complete SSE events before checking the unfinished remainder. A single network chunk can contain many complete events and is not subject to a 64 KiB aggregate limit. The default limit on an unfinished event is 65,536 bytes; it is not a total response or token limit. Increase it for providers that emit large fragmented tool-call events. Zero permits no unfinished bytes.

  • Standalone: --sse-buffer-limit 1048576 or ONWARDS_SSE_BUFFER_LIMIT=1048576.
  • Embedded library: AppState::with_sse_buffer_limit(1048576).
  • Control layer: onwards.sse_buffer_limit: 1048576 in config.yaml, or DWCTL_ONWARDS__SSE_BUFFER_LIMIT=1048576.

The limit applies wherever Onwards buffers SSE, including strict response sanitization, its leading-event inspection, and non-strict sanitization. Exceeding it produces a response-body error and drops the upstream stream, rather than reporting a clean end of stream. Once HTTP headers have been sent, this cannot change the HTTP status or transparently retry already forwarded output.

Command Line Options

FlagDescriptionDefault
--targets <file> / -f <file>Path to configuration fileRequired
--port <port>Port to listen on3000
--watchEnable configuration file watching for hot-reloadingtrue
--metricsEnable Prometheus metrics endpointtrue
--metrics-port <port>Port for metrics and /healthz, /readyz9090
--metrics-prefix <prefix>Prefix for metric namesonwards
--shutdown-delay-secs <seconds>Continue serving after readiness fails, before closing admission5
--shutdown-timeout-secs <seconds>Deadline for draining accepted requests after closing admission; must be positive300

The shutdown settings also accept ONWARDS_SHUTDOWN_DELAY_SECS and ONWARDS_SHUTDOWN_TIMEOUT_SECS. See Graceful shutdown for the signal sequence, health probes and Kubernetes termination budget.

Examples

Start with defaults:

cargo run -- -f config.json

Custom port, metrics disabled:

cargo run -- -f config.json --port 8080 --metrics false

Custom metrics configuration:

cargo run -- -f config.json --metrics-port 9100 --metrics-prefix gateway

Graceful shutdown

The standalone onwards binary handles SIGTERM and SIGINT using Axum’s with_graceful_shutdown. Applications embedding the Onwards library still own their HTTP server and its shutdown policy.

On receipt of a signal:

  1. /readyz on the metrics port changes from HTTP 200 to HTTP 503.
  2. For --shutdown-delay-secs (default 5), requests continue to be served while load balancers observe the readiness change.
  3. Axum stops accepting new proxy connections, closes idle keep-alive connections, and waits for accepted requests and response bodies to finish. An SSE response keeps the process alive until its body ends, even though its headers were sent earlier.
  4. Metrics and /healthz remain available while the proxy drains. When the proxy finishes, the metrics server closes and tracing is flushed.

If draining exceeds --shutdown-timeout-secs (default 300), the process logs the timeout and exits unsuccessfully. Remaining streams are cut off; this is a forced termination, not a successful drain. The metrics server has a separate five-second shutdown ceiling after the proxy stops.

Use HTTP probes on the metrics port, not TCP probes on the proxy listener:

terminationGracePeriodSeconds: 360
containers:
  - name: onwards
    # Pin the release containing this feature, or its immutable image digest.
    args:
      - --targets
      - /etc/onwards/targets.json
      - --shutdown-delay-secs
      - "5"
      - --shutdown-timeout-secs
      - "300"
    readinessProbe:
      httpGet:
        path: /readyz
        port: 9090
      periodSeconds: 2
      failureThreshold: 1
    livenessProbe:
      httpGet:
        path: /healthz
        port: 9090

The pod’s termination budget must exceed the propagation delay, proxy drain, metrics shutdown and tracing flush. Set it using the request durations your service supports. The example allows 55 seconds beyond the default 305-second propagation-plus-drain budget. The propagation delay replaces a sleep-only preStop hook; adding such a hook would consume extra time before SIGTERM. The health routes reveal only process readiness/liveness, are unauthenticated, and should remain internal alongside metrics. Readiness does not test every configured upstream model.

Do not apply a normal rolling update to old gateway versions that lack this behavior while they have active streams. Create a separate deployment and Service, move callers to its distinct hostname without restarting them, then verify the old gateways have no new requests and zero in-flight work before retiring them. Merely changing a Service selector does not move established connections. Validate client configuration propagation and retry behavior in a rehearsal before using this procedure in production.

After all replicas support draining, use required pod anti-affinity on kubernetes.io/hostname, a suitable PodDisruptionBudget, and enough spare capacity for rolling updates. These controls and signal handling do not preserve TCP streams through an abrupt node failure.

The Unix integration tests start the actual binary and a gated HTTP backend. They cover SSE completion across SIGTERM, a unary request still waiting for headers, deadline expiry with a stuck stream, and idle HTTP/1.1 keep-alive closure on SIGINT. They use synthetic traffic and do not terminate production pods. Run them with:

cargo test --package onwards --test graceful_shutdown

See Axum’s server API and Kubernetes pod termination.

API Usage

List available models

Get a list of all configured targets in the OpenAI models format:

curl http://localhost:3000/v1/models

Sending requests

Send requests to the gateway using the standard OpenAI API format:

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

The model field determines which target receives the request.

Model override header

Override the target using the model-override header. This routes the request to a different target regardless of the model field in the body:

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "model-override: claude-3" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

This is also used for routing requests without bodies – for example, to get the embeddings usage for your organization:

curl -X GET http://localhost:3000/v1/organization/usage/embeddings \
  -H "model-override: claude-3"

Metrics

When the --metrics flag is enabled (the default), Prometheus metrics are exposed on a separate port:

curl http://localhost:9090/metrics

See Command Line Options for metrics configuration flags.

Rejected requests

Every client error that onwards decides on itself, such as an unknown model, a failed reasoning check, a strict-mode schema error or a rate limit, increments onwards_rejections_total{model, status, code, traffic}:

  • code is the error code returned to the client. Strict-mode body errors carry no code, so they are counted as:
    • invalid_json: malformed JSON;
    • schema_mismatch: valid JSON that doesn’t match the schema;
    • invalid_content_type: a /v1/responses body not sent as JSON;
    • invalid_body: a /v1/responses body that couldn’t be read.
  • model is the configured model the request names, without any serving-class suffix such as :interactive. It is empty when the request names no configured model.
  • traffic is dispatched for requests carrying the first-token-timeout exempt header and realtime otherwise.

Each rejection is also logged at info with its status, code, parameter, model, account and API key ID. Values and request bodies are never logged. A client error that reports an upstream’s response isn’t counted.

Authentication

Onwards supports bearer token authentication to control access to your AI targets. You can configure authentication keys both globally and per-target.

Global authentication keys

Global keys apply to all targets that have authentication enabled:

{
  "auth": {
    "global_keys": ["global-api-key-1", "global-api-key-2"]
  },
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-your-openai-key",
      "keys": ["target-specific-key"]
    }
  }
}

Per-target authentication

You can specify authentication keys for individual targets:

{
  "targets": {
    "secure-gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-your-openai-key",
      "keys": ["secure-key-1", "secure-key-2"]
    },
    "open-local": {
      "url": "http://localhost:8080"
    }
  }
}

In this example:

  • secure-gpt-4 requires a valid bearer token from the keys array
  • open-local has no authentication requirements

If both global and local keys are supplied, either global or local keys will be valid for accessing models with local keys.

How authentication works

When a target has keys configured, requests must include a valid Authorization: Bearer <token> header where <token> matches one of the configured keys. If global keys are configured, they are automatically added to each target’s key set.

Successful authenticated request:

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Authorization: Bearer secure-key-1" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "secure-gpt-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Failed authentication (invalid key):

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Authorization: Bearer wrong-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "secure-gpt-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
# Returns: 401 Unauthorized

Failed authentication (missing header):

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "secure-gpt-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
# Returns: 401 Unauthorized

No authentication required:

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "open-local",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
# Success - no authentication required for this target

Upstream Authentication

By default, Onwards sends upstream API keys using the standard Authorization: Bearer <key> header format. Some AI providers use different authentication header formats. You can customize both the header name and prefix per target.

Custom header name

Some providers use custom header names for authentication:

{
  "targets": {
    "custom-api": {
      "url": "https://api.custom-provider.com",
      "onwards_key": "your-api-key-123",
      "upstream_auth_header_name": "X-API-Key"
    }
  }
}

This sends: X-API-Key: Bearer your-api-key-123

Custom header prefix

Some providers use different prefixes or no prefix at all:

{
  "targets": {
    "api-with-prefix": {
      "url": "https://api.provider1.com",
      "onwards_key": "token-xyz",
      "upstream_auth_header_prefix": "ApiKey "
    },
    "api-without-prefix": {
      "url": "https://api.provider2.com",
      "onwards_key": "plain-key-456",
      "upstream_auth_header_prefix": ""
    }
  }
}

This sends:

  • To provider1: Authorization: ApiKey token-xyz
  • To provider2: Authorization: plain-key-456

Combining custom name and prefix

You can customize both the header name and prefix:

{
  "targets": {
    "fully-custom": {
      "url": "https://api.custom.com",
      "onwards_key": "secret-key",
      "upstream_auth_header_name": "X-Custom-Auth",
      "upstream_auth_header_prefix": "Token "
    }
  }
}

This sends: X-Custom-Auth: Token secret-key

Default behavior

If these options are not specified, Onwards uses the standard OpenAI-compatible format:

{
  "targets": {
    "standard-api": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-openai-key"
    }
  }
}

This sends: Authorization: Bearer sk-openai-key

Rate Limiting

Onwards supports rate limiting using a token bucket algorithm. You can configure limits per-target and per-API-key.

Per-target rate limiting

Add rate limiting to any target in your config.json:

{
  "targets": {
    "rate-limited-model": {
      "url": "https://api.provider.com",
      "onwards_key": "your-api-key",
      "rate_limit": {
        "requests_per_second": 5.0,
        "burst_size": 10
      }
    }
  }
}

How it works

Each target gets its own token bucket. Tokens are refilled at a rate determined by requests_per_second. The maximum number of tokens in the bucket is determined by burst_size. When the bucket is empty, requests to that target are rejected with a 429 Too Many Requests response.

Examples

// Allow 1 request per second with burst of 5
"rate_limit": {
  "requests_per_second": 1.0,
  "burst_size": 5
}

// Allow 100 requests per second with burst of 200
"rate_limit": {
  "requests_per_second": 100.0,
  "burst_size": 200
}

Rate limiting is optional – targets without rate_limit configuration have no rate limiting applied.

Per-API-key rate limiting

In addition to per-target rate limiting, Onwards supports individual rate limits for different API keys. This allows you to provide different service tiers – for example, basic users might have lower limits while premium users get higher limits.

Configuration

Per-key rate limiting uses a key_definitions section in the auth configuration:

{
  "auth": {
    "global_keys": ["fallback-key"],
    "key_definitions": {
      "basic_user": {
        "key": "sk-user-12345",
        "rate_limit": {
          "requests_per_second": 10,
          "burst_size": 20
        }
      },
      "premium_user": {
        "key": "sk-premium-67890",
        "rate_limit": {
          "requests_per_second": 100,
          "burst_size": 200
        }
      },
      "enterprise_user": {
        "key": "sk-enterprise-abcdef",
        "rate_limit": {
          "requests_per_second": 500,
          "burst_size": 1000
        }
      }
    }
  },
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-your-openai-key",
      "keys": ["basic_user", "premium_user", "enterprise_user", "fallback-key"]
    }
  }
}

Priority order

Rate limits are checked in this order:

  1. Per-key rate limits (if the API key has limits configured)
  2. Per-target rate limits (if the target has limits configured)

If either limit is exceeded, the request returns 429 Too Many Requests.

Usage examples

Basic user request (10/sec limit):

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Authorization: Bearer sk-user-12345" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4", "messages": [{"role": "user", "content": "Hello!"}]}'

Premium user request (100/sec limit):

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Authorization: Bearer sk-premium-67890" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4", "messages": [{"role": "user", "content": "Hello!"}]}'

Legacy key (no per-key limits, only target limits apply):

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Authorization: Bearer fallback-key" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4", "messages": [{"role": "user", "content": "Hello!"}]}'

Concurrency Limiting

In addition to rate limiting (which controls how fast requests are made), concurrency limiting controls how many requests are processed simultaneously. This is useful for managing resource usage and preventing overload.

Per-target concurrency limiting

Limit the number of concurrent requests to a specific target:

{
  "targets": {
    "resource-limited-model": {
      "url": "https://api.provider.com",
      "onwards_key": "your-api-key",
      "concurrency_limit": {
        "max_concurrent_requests": 5
      }
    }
  }
}

With this configuration, only 5 requests will be processed concurrently for this target. Additional requests will receive a 429 Too Many Requests response until an in-flight request completes.

Per-API-key concurrency limiting

You can set different concurrency limits for different API keys:

{
  "auth": {
    "key_definitions": {
      "basic_user": {
        "key": "sk-user-12345",
        "concurrency_limit": {
          "max_concurrent_requests": 2
        }
      },
      "premium_user": {
        "key": "sk-premium-67890",
        "concurrency_limit": {
          "max_concurrent_requests": 10
        },
        "rate_limit": {
          "requests_per_second": 100,
          "burst_size": 200
        }
      }
    }
  },
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-your-openai-key"
    }
  }
}

Per-account in-flight limits

A pool can cap how many realtime requests each account has in flight on its alias. A key’s account is its account label, so every key of an account shares one count. inflight_limit is the default and account_inflight_limits overrides it for named accounts:

{
  "auth": {
    "global_keys": [],
    "key_definitions": {
      "acme_backend": { "key": "sk-acme-1", "labels": { "account": "acme" } },
      "acme_batch": { "key": "sk-acme-2", "labels": { "account": "acme" } }
    }
  },
  "targets": {
    "gpt-4": {
      "inflight_limit": 20,
      "account_inflight_limits": { "acme": 200 },
      "providers": [{ "url": "https://api.openai.com", "onwards_key": "sk-your-openai-key" }]
    }
  }
}

The slot is taken after routing rules and counts against the alias the request named, even when a rule redirects it to another alias, so the named alias’s limits apply. It is held across failover attempts and released when the response body finishes or the client disconnects. A request over the limit receives 429 with code inflight_limit_exceeded and a Retry-After: 1 header. The limits belong to the alias as a whole: they are read from its default pool, whichever request-class pool serves the request. Requests carrying the header set with AppState::with_first_token_timeout_exempt_header (dispatched batch and async work) and requests from keys with no account label are not counted, and the plugged-in limiter can admit others without counting them. Counts are per process by default; AppState::with_inflight_limiter plugs in a shared counter so several instances enforce one limit.

Batch in-flight cap

Realtime and batch are deliberately asymmetric. Realtime has no alias-wide total in-flight cap: only the per-account limits described above, which divide an alias’s capacity between tenants rather than capping the alias as a whole. A full downstream answers 529, and realtime traffic is allowed to grow into whatever capacity exists, with the per-account limit stopping one tenant from taking it all. Batch instead gets a per-model cap, because batch must never crowd realtime out of a model and can always be processed later — a refused batch request is rescheduled, not lost.

A pool can also cap how many batch (dispatched) requests are in flight on its alias. Unlike the per-account limit above, this is a single count for the alias rather than a per-tenant one. It is shared across every dispatcher and proxy instance only when a shared limiter is plugged in; with the default local limiter the count is per process (see below). It is configured with batch_inflight_limit:

{
  "targets": {
    "gpt-4": {
      "inflight_limit": 20,
      "batch_inflight_limit": 8,
      "providers": [{ "url": "https://api.openai.com", "onwards_key": "sk-your-openai-key" }]
    }
  }
}

A batch request is one onwards classifies as dispatched: it carries the header set with AppState::with_first_token_timeout_exempt_header. Realtime requests are never counted against this cap, and batch requests are never counted against the per-account inflight_limit. The two counts are independent, so an alias can serve realtime traffic while its batch slots are full. None (or an absent field) means batch is uncapped. The control layer resolves this field for each virtual model before writing the onwards config: a positive batch_capacity is used, otherwise it falls back to limits.batch_inflight.default_capacity (default 200), so an enforced deployment caps virtual models that never set a value unless that default is turned off.

The cap is read from the alias’s default pool and counts against the alias the request named, even when a routing rule redirects it, so a redirected request still consumes a slot on the model it asked for. The slot is held across failover attempts for the life of the response body and released when the body finishes or the client disconnects.

A request over the cap receives 529 (the shared overload status) with type overloaded_error, code batch_capacity_exceeded, and a Retry-After: 1 header. It is deliberately not a 429: the dispatcher treats 529 as a downstream overload and reduces its adaptive concurrency, whereas a 429 is a per-key rate limit and would not. A batch_capacity_exceeded 529 is admission control rather than a failed attempt, so the dispatcher reschedules it with normal backoff without spending a retry attempt; an ordinary 529 still spends one. The refusal is counted in onwards_batch_inflight_refusals_total{model}.

Counts are per process by default. AppState::with_batch_inflight_limiter plugs in a shared counter (the control layer passes the same Redis-backed limiter used for realtime inflight_limit, under a reserved __batch__ scope so the two key spaces never collide). Enforcement is switched with AppState::with_batch_inflight_enforce; when it is off, batch requests are never refused and the limiter is not consulted at all.

batch_inflight_limit is consumed only by an embedding process that both plugs in a limiter and turns enforcement on — the control layer does this when it resolves virtual models. The standalone onwards binary does not: it loads batch_inflight_limit from its config but never calls AppState::with_batch_inflight_limiter or AppState::with_batch_inflight_enforce, so enforcement stays off and the field has no effect there. To use the cap in a standalone process, call both builders on the constructed AppState.

Combining rate limiting and concurrency limiting

You can use both rate limiting and concurrency limiting together:

  • Rate limiting controls how fast requests are made over time
  • Concurrency limiting controls how many requests are active at once
{
  "targets": {
    "balanced-model": {
      "url": "https://api.provider.com",
      "onwards_key": "your-api-key",
      "rate_limit": {
        "requests_per_second": 10,
        "burst_size": 20
      },
      "concurrency_limit": {
        "max_concurrent_requests": 5
      }
    }
  }
}

How it works

Concurrency limits use a semaphore-based approach:

  1. When a request arrives, it tries to acquire a permit
  2. If a permit is available, the request proceeds (holding the permit)
  3. If no permits are available, the request is rejected with 429 Too Many Requests
  4. When the request completes, the permit is automatically released

The error response distinguishes between rate limiting and concurrency limiting:

  • Rate limit: "code": "rate_limit"
  • Concurrency limit: "code": "concurrency_limit_exceeded"

Both use HTTP 429 status code for consistency.

Response Headers

Onwards can include custom headers in the response for each target. These can override existing headers or add new ones.

Configuration

{
  "targets": {
    "model-with-headers": {
      "url": "https://api.provider.com",
      "onwards_key": "your-api-key",
      "response_headers": {
        "X-Custom-Header": "custom-value",
        "X-Provider": "my-gateway"
      }
    }
  }
}

Pricing headers

One use of this feature is to set pricing information. If you have a dynamic token price, when a user’s request is accepted the price is agreed and can be recorded in the HTTP headers:

{
  "targets": {
    "priced-model": {
      "url": "https://api.provider.com",
      "onwards_key": "your-api-key",
      "response_headers": {
        "Input-Price-Per-Token": "0.0001",
        "Output-Price-Per-Token": "0.0002"
      }
    }
  }
}

When using load balancing, response headers can be configured at both the pool level and provider level. Provider-level headers take precedence.

Strict Mode

Strict mode provides enhanced security and API compliance by using typed request/response handlers instead of the default wildcard passthrough router. This feature:

  • Validates all requests against OpenAI API schemas before forwarding
  • Sanitizes all responses by removing third-party provider metadata
  • Standardizes error messages to prevent information leakage
  • Ensures model field consistency between requests and responses
  • Supports streaming and non-streaming for all endpoints

This is useful when you need guaranteed API compatibility, security hardening, or protection against third-party response variations.

Enabling strict mode

Strict mode is a global configuration that applies to all targets in your gateway. Add strict_mode: true at the top level of your configuration (not inside individual targets).

{
  "strict_mode": true,
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-openai-key"
    },
    "claude": {
      "url": "https://api.anthropic.com",
      "onwards_key": "sk-ant-key"
    }
  }
}

When enabled, all requests to all targets will use strict mode validation and sanitization.

How it works

When strict_mode: true is enabled:

  1. Request validation: Incoming requests are deserialized through OpenAI schemas. Invalid requests receive immediate 400 Bad Request errors with clear messages.
  2. Response sanitization: Third-party responses are deserialized (automatically dropping unknown fields). If deserialization fails, a standard error is returned - malformed responses are never passed through. The model field is rewritten to match the original request, then re-serialized as clean OpenAI responses with correct Content-Length headers.
  3. Error standardization: Third-party errors are logged internally but never forwarded to clients. Clients receive standardized OpenAI-format errors based only on HTTP status codes.
  4. Streaming support: SSE streams are parsed line-by-line to handle multi-line data events and strip comment lines, each chunk is sanitized, and re-emitted as clean events.

Security benefits

Prevents information leakage:

  • Third-party stack traces, database errors, and debug information are never exposed
  • Error responses contain only standard HTTP status codes and generic messages
  • No provider-specific metadata (trace IDs, internal IDs, costs) reaches clients
  • Malformed provider responses fail closed with standard errors (never leaked)
  • SSE comment lines stripped to prevent metadata leakage in streaming responses

Ensures consistency:

  • Responses always match OpenAI’s API format exactly
  • The model field always reflects what the client requested, not what the provider returned
  • Extra fields like provider, cost, trace_id are automatically dropped

Fast failure:

  • Invalid requests fail immediately with clear, actionable error messages
  • No wasted upstream requests for malformed input
  • Reduces debugging time for integration issues

Error standardization

When strict mode is enabled, all error responses follow OpenAI’s error format exactly:

{
  "error": {
    "message": "Invalid request",
    "type": "invalid_request_error",
    "param": null,
    "code": null
  }
}

Status code mapping:

HTTP StatusError TypeMessage
400invalid_request_errorInvalid request
401authentication_errorAuthentication failed
403permission_errorPermission denied
404not_found_errorNot found
429rate_limit_errorRate limit exceeded
500api_errorInternal server error
502api_errorBad gateway
503api_errorService unavailable

For untrusted upstreams, account-class status codes are remapped before this table is applied — see Account-class status code masking.

Third-party error details are always logged server-side but never sent to clients.

Account-class status code masking

Untrusted providers can return status codes that describe the operator’s relationship with the provider — not the caller’s relationship with onwards. Surfacing those would let callers probe the operator’s account state. Onwards rewrites:

Upstream statusSurfaced asRationale
401 Unauthorized502 Bad GatewayThe caller’s auth is fine; ours isn’t — don’t expose that
402 Payment Required502 Bad GatewayProvider billing state is internal
403 Forbidden502 Bad GatewayPermission failures are operator-scoped
451 Unavailable For Legal Reasons502 Bad GatewayJurisdictional restrictions are operator-scoped
408 Request Timeout504 Gateway TimeoutUpstream timeouts are gateway timeouts from the caller’s view

User-facing codes (400, 404, 413, 422, 429) pass through unchanged — they’re real signal about the caller’s request and downstream retry logic depends on seeing them.

For untrusted providers, the embedded error’s code is read whether the provider encodes it as a number (429), a numeric string ("429"), or a named string — the rate-limit family (rate_limit, rate_limit_error, rate_limit_exceeded) maps to 429 so retry semantics survive, and unrecognized names fall back to 500. The resulting code is then masked per the table above.

Masking is applied to both non-streaming error responses (where the upstream status is the outer HTTP status) and to embedded error.code values in SSE streams (so a downstream reassembler that reclassifies on the embedded code surfaces the masked code as the HTTP status). Trusted targets bypass masking entirely.

Errors embedded in 2xx SSE streams

Some providers return HTTP 200 OK and start an SSE stream even when the upstream of the upstream has failed. The failure surfaces as a chunk with shape:

data: {"id":"...","object":"chat.completion.chunk","choices":[],"error":{"code":429,"message":"..."}}

The naive strict deserializer would parse this as a valid (empty) chunk and silently drop the error field, leaving a downstream stream reassembler with no completion content and no signal that anything went wrong.

Onwards detects the embedded error envelope before strict deserialization and forwards it as a stand-alone event:

data: {"error":{"message":"Rate limit exceeded","type":"rate_limit_error","param":null,"code":429}}

The chunk wrapper is stripped so the emitted SSE data line begins with {"error" — that prefix is what the fusillade reassembler matches on to reclassify the response from HTTP 200 to the embedded code. End-to-end, an upstream-of-upstream 429 becomes a real HTTP 429 to the client, with retry semantics intact.

This applies to both bare error chunks and chunks that carry completion fields alongside the error. Untrusted upstreams additionally have the embedded error.code masked (per the table above) and the prose replaced with a generic message; trusted targets forward the envelope verbatim.

Supported endpoints

Strict mode currently supports:

  • /v1/chat/completions (streaming and non-streaming) - Full sanitization
  • /v1/embeddings - Full sanitization
  • /v1/responses (Open Responses API, non-streaming) - Full sanitization
  • /v1/models - Model listing (no sanitization needed)

All supported endpoints include:

  • Request validation - Invalid requests fail immediately with clear error messages
  • Response sanitization - Third-party metadata automatically removed
  • Model field rewriting - Ensures consistency with client request
  • Error standardization - Third-party error details never exposed

Requests to unsupported endpoints will return 404 Not Found when strict mode is enabled.

Parameters not every backend supports

Some chat parameters are engine extensions or options that not every backend behind a model can honour, for example n greater than 1, logprobs, guided_json or min_tokens. The full list, with the values that count as absent (such as n: 1, logprobs: false or top_logprobs: 0), is the catalog in onwards::unsupported_params. chat_template_kwargs isn’t in it: strict mode refuses that field outright and points to reasoning_effort instead.

Strict mode checks every Chat Completions request against that list after schema validation and before reasoning validation:

  • Each matching parameter increments onwards_unsupported_params_total{param, model, traffic, action}. action is logged or rejected. A logged request can still be refused by a later check, such as reasoning validation.
  • The request is logged at info with the parameter names, model, account and API key ID. Values and request bodies are never logged.
  • If any parameter is in the gateway’s reject list, the whole request gets a 400 with code unsupported_parameter naming the rejected parameters, for example Unsupported parameter(s): `logprobs` , and is not forwarded. A request whose flagged parameters are all outside the reject list is forwarded unchanged.

The check covers Chat Completions requests, including Responses and Messages requests that an edge such as dwctl has translated into Chat Completions before they reach onwards. Native /v1/responses requests are forwarded to an upstream that speaks the Responses API itself and aren’t checked.

The reject list is empty by default, so the check only observes. Set it with AppState::with_rejected_params; in dwctl, use onwards.rejected_params. A name outside the catalog is a configuration error.

Comparison with response sanitization

FeatureResponse SanitizationStrict Mode
Request validation✗ No✓ Yes
Response sanitization✓ Yes✓ Yes
Error standardization✗ No✓ Yes
Endpoint coverage/v1/chat/completions onlyChat, Embeddings, Responses, Models
Router typeWildcard passthroughTyped handlers
Use caseSimple response cleaningProduction security & compliance

Important: When strict mode is enabled globally, the per-target sanitize_response flag is automatically ignored. Strict mode handlers perform complete sanitization themselves, so enabling sanitize_response: true on individual targets has no effect and won’t cause double sanitization.

When to use strict mode:

  • Production deployments requiring security hardening
  • Compliance requirements around error message content
  • Multi-provider setups needing guaranteed response consistency
  • Applications that need request validation before forwarding

When to use response sanitization:

  • Simple use cases where you only need response cleaning
  • Non-security-critical deployments
  • Maximum flexibility with endpoint coverage

Trusted Providers

In strict mode, you can mark providers as trusted to bypass error sanitization while keeping success response sanitization. This is useful when you have providers you fully control (e.g., your own OpenAI account) and want their detailed error messages to help with debugging, while still ensuring response consistency.

Trust can be set at two levels:

  • Pool level (trusted on the target) — default for all providers in the pool
  • Provider level (trusted inside a provider entry) — overrides the pool default for that specific provider

This is the only exception to strict mode’s error standardization guarantees: when a provider is effectively trusted, its errors may be forwarded with full third-party details instead of being standardized.

Configuration

Single-provider (pool-level trusted):

{
  "strict_mode": true,
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-...",
      "trusted": true
    },
    "third-party": {
      "url": "https://some-provider.com",
      "onwards_key": "sk-..."
    }
  }
}

Uniform pool-level trusted:

{
  "strict_mode": true,
  "targets": {
    "gpt-4-pool": {
      "trusted": true,
      "providers": [
        { "url": "https://api.openai.com", "onwards_key": "sk-primary-..." },
        { "url": "https://api.openai.com", "onwards_key": "sk-backup-..." }
      ]
    }
  }
}

Mixed trust within a pool (per-provider override):

{
  "strict_mode": true,
  "targets": {
    "gpt-4": {
      "trusted": false,
      "providers": [
        { "url": "https://internal.example.com", "trusted": true },
        { "url": "https://external.example.com" }
      ]
    }
  }
}

Here, the internal provider’s error responses pass through unchanged. The external provider omits trusted, so it inherits the pool default (false) and has its error responses sanitized. This lets you mix trusted internal infrastructure with untrusted external providers inside a single pool.

Behavior

When a pool is marked as trusted: true:

Success responses (200 OK) are STILL sanitized:

  • Model field IS rewritten to match the client’s request
  • Provider-specific metadata IS removed (costs, trace IDs, custom fields)
  • Response IS validated against OpenAI schemas
  • Content-Length headers ARE updated correctly
  • Streaming responses ARE parsed and sanitized line-by-line

Error responses (4xx, 5xx) bypass sanitization:

  • Original error messages and metadata forwarded to clients
  • Provider-specific error details preserved (stack traces, debug info)
  • Custom error fields passed through unchanged

This allows you to get detailed debugging information from errors while maintaining response consistency for successful requests.

Security Warning

⚠️ Use trusted providers carefully. Marking a provider as trusted bypasses error sanitization for that provider:

What is exposed for trusted providers:

  • Error details and stack traces from the provider
  • Provider-specific error metadata (trace IDs, internal error codes)
  • Debug information in error responses

What is NOT exposed (still sanitized):

  • Success responses are fully sanitized (model rewritten, metadata removed)
  • Provider metadata in successful requests (costs, trace IDs) is still stripped
  • Responses still match OpenAI schema exactly for successful requests

Only mark providers as trusted when you fully control or trust them. This typically means:

  • Your own OpenAI/Anthropic accounts (providers using your API keys)
  • Self-hosted models you operate
  • Internal services you maintain

Do not mark third-party providers as trusted unless you want their detailed error messages exposed to your clients. Trusted providers are designed for debugging your own infrastructure, not for production use with external providers.

Interaction with Model Override Header

Onwards supports the model-override header to route requests to different pools than specified in the request body. Trust is resolved from the provider that actually handles the request (after routing and provider selection), so it correctly reflects the resolved model rather than what the client specified in the body.

This means if a client sends:

  • Request body with "model": "trusted-pool"
  • Header with model-override: untrusted-pool

The request will route to untrusted-pool and sanitization will be applied based on that pool’s provider trust settings, preventing metadata leakage. Clients cannot bypass sanitization by exploiting mismatches between body and header model resolution.

Implementation details

For developers working on the Onwards codebase:

Router architecture:

  • Strict mode uses typed Axum handlers defined in src/strict/handlers.rs
  • Each endpoint has dedicated request/response schema types in src/strict/schemas/
  • Requests are deserialized using serde, which automatically validates structure
  • response_transform_fn is skipped when strict mode is enabled to prevent double sanitization

Response sanitization:

  • Responses are deserialized through strict schemas (extra fields automatically dropped by serde)
  • Malformed responses fail closed with standard errors - never passed through
  • Model field is rewritten to match the original request model
  • Re-serialized to ensure only defined fields are present
  • Content-Length headers updated to match sanitized response size
  • Applies to both non-streaming responses and SSE chunks
  • SSE streams processed line-by-line to handle multi-line events and strip comments

Error handling:

  • Third-party errors are intercepted in sanitize_error_response()
  • Original error logged with error!() macro for server-side debugging
  • Standard error generated based only on HTTP status code
  • OpenAI-compatible format guaranteed via error_response() helper
  • Deserialization failures return standard errors, never leak malformed responses

Trust resolution:

  • target_message_handler resolves effective trust as provider.trusted.unwrap_or(pool.trusted) after provider selection
  • The resolved trust is attached to the response via a ResolvedTrust extension
  • Strict mode handlers read it via ForwardResult.trusted — no separate pool lookup needed
  • Ensures trust reflects the actual provider that handled the request, including after fallback retries

Testing:

  • Request/response schema tests in each schema module
  • Integration tests in src/strict/handlers.rs verify sanitization behavior
  • Tests verify fail-closed behavior on malformed responses (no passthrough)
  • Tests verify SSE multi-line events and comment stripping
  • Tests verify Content-Length header correctness after sanitization
  • Tests verify per-provider trusted overrides pool-level setting in both directions

Response Sanitization

Onwards can enforce strict OpenAI API schema compliance for /v1/chat/completions responses. This feature:

  • Removes provider-specific fields from responses
  • Rewrites the model field to match what the client originally requested
  • Supports both streaming and non-streaming responses
  • Validates responses against OpenAI’s official API schema
  • Sanitizes error responses to prevent upstream provider details from leaking to clients

This is useful when proxying to non-OpenAI providers that add custom fields, or when using onwards_model to rewrite model names upstream.

Note: For production deployments requiring additional security (request validation, error standardization), consider using Strict Mode instead, which includes response sanitization plus comprehensive security features.

Enabling response sanitization

Add sanitize_response: true to any target or provider in your configuration.

Single provider:

{
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-your-key",
      "onwards_model": "gpt-4-turbo-2024-04-09",
      "sanitize_response": true
    }
  }
}

Pool with multiple providers:

{
  "targets": {
    "gpt-4": {
      "sanitize_response": true,
      "providers": [
        {
          "url": "https://api1.example.com",
          "onwards_key": "sk-key-1"
        },
        {
          "url": "https://api2.example.com",
          "onwards_key": "sk-key-2"
        }
      ]
    }
  }
}

How it works

When sanitize_response: true and a client requests model: gpt-4:

  1. Request sent upstream with model: gpt-4
  2. Upstream responds with custom fields and model: gpt-4-turbo-2024-04-09
  3. Onwards sanitizes:
    • Parses response using OpenAI schema (removes unknown fields)
    • Rewrites model field to gpt-4 (matches original request)
    • Reserializes clean response
  4. Client receives standard OpenAI response with model: gpt-4

Common use cases

Third-party providers (e.g., Together AI) often add extra fields like provider, native_finish_reason, cost, etc. Sanitization strips these.

Provider comparison – normalize responses from different providers for consistent handling.

Debugging – reduce noise by filtering to only standard OpenAI fields.

Error sanitization

When sanitize_response: true, error responses from upstream providers are also sanitized. This prevents information leakage – upstream error bodies can contain provider names, internal URLs, and model identifiers that you may not want exposed to clients.

How it works

Onwards replaces the upstream error body with a generic OpenAI-compatible error, while preserving the original HTTP status code:

  • 4xx errors are replaced with:
{
  "error": {
    "message": "The upstream provider rejected the request.",
    "type": "invalid_request_error",
    "param": null,
    "code": "upstream_error"
  }
}
  • 5xx errors (and any other non-2xx status) are replaced with:
{
  "error": {
    "message": "An internal error occurred. Please try again later.",
    "type": "internal_error",
    "param": null,
    "code": "internal_error"
  }
}

The original error body is logged at ERROR level (up to 64 KB) for debugging, so operators can still investigate upstream failures without exposing details to clients.

Errors embedded in 2xx SSE streams

Some providers return HTTP 200 OK and start an SSE stream even when the upstream of the upstream has failed, embedding the failure in a chunk alongside (or instead of) normal completion fields:

data: {"id":"...","object":"chat.completion.chunk","choices":[],"error":{"code":429,"message":"..."}}

Without handling, the lenient deserializer absorbs the error field into its unknown-fields map and drops it on re-serialize, leaving the caller with a content-less stream and no signal that anything went wrong. With sanitize_response: true, Onwards detects the embedded error envelope and forwards it as a stand-alone event with the chunk wrapper stripped:

data: {"error":{"code":429,"message":"..."}}

The emitted data line begins with {"error", the prefix downstream reassemblers match on to reclassify the response from HTTP 200 to the embedded code.

Non-strict mode forwards the error verbatim. The provider’s message and the original status code pass through unchanged — non-strict mode does not mask account-class codes (401/402/403/451) or replace the provider’s prose. If you need that protection so callers can’t probe the operator’s auth/billing/jurisdictional state, use Strict Mode, which masks account-class codes and replaces untrusted error messages with generic ones.

⚠️ Security warning: Verbatim forwarding passes the entire error object through to your clients — every field, including the provider’s message, billing/auth state hints, and any nested metadata/raw/unknown fields. Do not enable sanitize_response: true on targets proxying untrusted third parties if you need to hide upstream internals. Use Strict Mode for production deployments requiring information-leakage prevention.

Error format

All Onwards error responses (both sanitized upstream errors and errors generated by Onwards itself) use the OpenAI-compatible {"error": {...}} envelope:

{
  "error": {
    "message": "...",
    "type": "...",
    "param": null,
    "code": "..."
  }
}
FieldDescription
messageHuman-readable error description
typeError category (invalid_request_error, rate_limit_error, internal_error)
paramThe request parameter that caused the error, if applicable
codeMachine-readable error code

Supported endpoints

Currently supports:

  • /v1/chat/completions (streaming and non-streaming)

Load Balancing

Onwards supports load balancing across multiple providers for a single alias, with automatic failover, weighted distribution, and configurable retry behavior.

Configuration

{
  "targets": {
    "gpt-4": {
      "strategy": "weighted_random",
      "fallback": {
        "enabled": true,
        "on_status": [429, 5],
        "on_rate_limit": true
      },
      "providers": [
        { "url": "https://api.openai.com", "onwards_key": "sk-key-1", "weight": 3 },
        { "url": "https://api.openai.com", "onwards_key": "sk-key-2", "weight": 1 }
      ]
    }
  }
}

Strategy

  • weighted_random (default): Distributes traffic randomly based on weights. A provider with weight: 3 receives ~3x the traffic of weight: 1.
  • priority: Always routes to the first provider. Falls through to subsequent providers only when fallback is triggered.

Fallback

Controls automatic retry on other providers when requests fail:

OptionTypeDefaultDescription
enabledboolfalseMaster switch for fallback
on_statusint[]–Status codes that trigger fallback (supports wildcards)
realtime_on_statusint[][]Extra statuses that trigger fallback for realtime requests only
on_rate_limitboolfalseFallback when hitting local rate limits
first_token_timeout_msint–Failover deadline for the first token of a streamed response; 0 disables (see below)
affinityobject–Priority pools only: choose preferred-first once per conversation instead of per request (see below)

Status code wildcards:

  • 5 matches all 5xx (500-599)
  • 50 matches 500-509
  • 502 matches exact 502

When fallback triggers, the next provider is selected based on strategy (weighted random resamples from remaining pool; priority uses definition order).

First-token failover

A provider can accept a streamed request and then stall: the headers arrive, but no token ever does. first_token_timeout_ms bounds that wait, so the request fails over instead of hanging. Set a proxy-wide default with AppState::with_first_token_timeout; a pool’s own value overrides it.

The deadline is only armed when all of these hold, so it can reroute a stalled request but never fail one that would otherwise have succeeded:

  • The request is "stream": true. A non-streaming response only sends headers once the whole completion is done, so a deadline would cut off long answers.
  • The pool has more than one provider, and the attempt is not the last one the attempt budget allows.
  • The request doesn’t carry the header set with AppState::with_first_token_timeout_exempt_header.

It bounds the wait for response headers and, in strict mode, the wait for the first real SSE frame. Keep-alive comments don’t count. Nothing has reached the client at that point, so the next provider starts cleanly.

First-token observations

Two metrics expose the existing first-token checks without changing provider selection or failover:

MetricTypeMeaning
onwards_first_token_secondsHistogramTime from the start of an attempt to its first observed non-[DONE] SSE data frame
onwards_first_token_breaches_totalCounterFirst-token failover deadlines that expired, while waiting for headers or an SSE frame

Both have model, pool, and role labels. model is the requested, validated alias, as for onwards_model_inflight; pool is the resolved pool name (default or a named pool such as completions, including after a routing redirect). role is preferred for the first provider in definition order and alternate for every other provider, regardless of attempt order or selection strategy. Alternate providers are aggregated; provider URLs are never labels.

Samples come only from the existing strict-mode 2xx SSE lead-frame check. They include the wait for headers and skip keep-alive comments, embedded errors, and [DONE]. A [DONE]-only stream remains valid and does not become retryable. The measurement is time to a data frame, which can contain metadata such as a role delta; it does not require generated text.

This histogram has partial coverage: non-strict streams and non-SSE responses produce no samples. Without an armed first-token deadline, the existing peek has time and event limits; a first frame arriving after those limits is also unobserved. Observing those streams would require additional instrumentation. Do not interpret the histogram as the full latency distribution or combine its count with the breach counter to estimate an overall breach rate.

A deadline expiry is a censored observation, so it increments only the breach counter, never the histogram. The counter measures the configured failover deadline, not a separate latency budget. Network errors, HTTP or embedded upstream errors, and the provider’s independent request timeout do not increment it. Non-strict requests can still increment it while waiting for headers.

Pool-level options

Settings that apply to the entire alias:

OptionDescription
keysAccess control keys for this alias
rate_limitRate limit for all requests to this alias
concurrency_limitMax concurrent requests to this alias
response_headersHeaders added to all responses
strategyweighted_random or priority
fallbackRetry configuration (see above)
providersArray of provider configurations

Provider-level options

Settings specific to each provider:

OptionDescription
urlProvider endpoint URL
onwards_keyAPI key for this provider
onwards_modelModel name override
weightTraffic weight (default: 1)
rate_limitProvider-specific rate limit
concurrency_limitProvider-specific concurrency limit
response_headersProvider-specific headers
trustedOverride pool-level trust for strict mode error sanitization (true/false; omit to inherit from pool)
propagate_trace_contextWhether to inject W3C traceparent / tracestate headers on outbound requests to this provider (true/false; omit to inherit from the resolved trusted value). Useful for preventing trace IDs from leaking to third-party providers whose downstream HTTP fetches would re-emit them. See Trace context propagation below.

Trace context propagation

onwards forwards W3C trace context (traceparent and tracestate headers) on outbound requests to upstream providers, so a downstream service that participates in your distributed tracing fabric can stitch its spans into the calling trace.

Whether the headers are sent is controlled by propagate_trace_context:

  • propagate_trace_context: true — always propagate
  • propagate_trace_context: false — never propagate
  • omitted (default) — inherit from the resolved trusted value:
    • per-provider trusted: true|false overrides
    • falling back to the pool-level trusted (default false)

In effect: trusted upstreams receive trace context by default; untrusted upstreams do not. This prevents trace IDs from leaking to third-party services that may re-emit them on their own outbound calls (e.g., a provider’s image fetcher echoing your traceparent back to whatever URL the caller supplied).

Migration note. Prior to onwards v0.28, traceparent was propagated to every upstream unconditionally. After this change, non-trusted upstreams no longer propagate by default (and any inbound trace context is stripped before forwarding to them). If you rely on trace continuity across onwards → upstream and the upstream isn’t marked trusted: true, set propagate_trace_context: true on that provider. The field is provider-scoped: set it on each relevant entry of a pool’s providers array, or on a legacy single-provider target. There is no pool-level propagate_trace_context key — for a whole pool, mark the pool trusted: true (which both bypasses error sanitization and enables propagation) or set the field on each provider entry.

Examples

Primary/backup failover

{
  "targets": {
    "gpt-4": {
      "strategy": "priority",
      "fallback": { "enabled": true, "on_status": [5], "on_rate_limit": true },
      "providers": [
        { "url": "https://primary.example.com", "onwards_key": "sk-primary" },
        { "url": "https://backup.example.com", "onwards_key": "sk-backup" }
      ]
    }
  }
}

Multiple API keys with pool-level rate limit

{
  "targets": {
    "gpt-4": {
      "rate_limit": { "requests_per_second": 100, "burst_size": 200 },
      "providers": [
        { "url": "https://api.openai.com", "onwards_key": "sk-key-1" },
        { "url": "https://api.openai.com", "onwards_key": "sk-key-2" }
      ]
    }
  }
}

Backwards compatibility

Single-provider configs still work unchanged:

{
  "targets": {
    "gpt-4": {
      "url": "https://api.openai.com",
      "onwards_key": "sk-key"
    }
  }
}

Load-aware priority share

Priority pools with enabled fallback and multiple providers automatically adjust the share of requests that try the preferred provider first. The share decreases when the preferred provider’s breach rate — first frames later than the budget, plus overload statuses such as 429, 503 and 529 — rises above a target, holds inside a hysteresis band, and recovers in steps once the rate stays low or the pool has too few samples to judge. A pool that temporarily drops to one provider keeps its controller and resumes it when the preferred provider returns. Set fallback.aimd.enabled: false to opt out, or override the default parameters. The controller applies only to eligible strict-mode streams, and it preserves the preferred provider in subsequent failover attempts. See load-aware failover for configuration, eligibility, sampling limits, reload behavior and rollout.

Conversation affinity

An agentic conversation sends many turns, each repeating the conversation so far. The provider that served the previous turn usually still holds that prefix in its cache; another provider has to process it again. When the preferred provider cannot take every request, choosing per request moves conversations back and forth and loses the cache on both sides.

fallback.affinity makes the choice per conversation. Each realtime request is keyed by an explicit identifier when the client sends one (x-session-id header, or session_id / prompt_cache_key in the body), otherwise by the conversation’s opening: its first system or developer message and its first other message. The key maps to a point in [0, 1), and the preferred provider is tried first when the point is below the pool’s affinity share. The share admits about target_conversations of the conversations active in the last active_window_ms, and moves only when the admitted count leaves target_conversations ± margin, so admitted conversations stay put while the population is steady.

"fallback": {
  "enabled": true,
  "affinity": { "target_conversations": 50 }
}
OptionDefaultDescription
target_conversationsrequiredConversations to keep on the preferred provider
margina tenth of the targetDrift allowed before the share moves; at most the target
active_window_ms600000How long a conversation stays active after its last request
update_interval_ms60000The share is recomputed at multiples of this wall-clock interval
max_tracked100000Conversations tracked per process; when full, the tracker keeps the ones that decide the share
enabledtrueSet false to keep per-request selection

Replicas share no state. Each records the conversations it sees and recomputes the share at the same wall-clock instants; because a conversation’s requests are spread across replicas, they see the same conversations and reach the same share. The load-aware share still caps the preferred side for eligible streams, so an overloaded preferred provider sheds conversations. Continuation pools and requests without a key keep ordinary priority selection.

target_conversations does not adapt: set it to what the preferred provider can serve concurrently, and change it when that capacity changes.

Realtime-only failover statuses

fallback.realtime_on_status adds statuses that fail a request over to the next provider only when the request is realtime — it lacks the header set with AppState::with_first_token_timeout_exempt_header — on top of fallback.on_status. It accepts the same wildcards. Use it for a provider’s over-capacity status: realtime callers are rerouted, while dispatched traffic that runs its own retries receives the upstream response.

Contributing

Testing

Run the test suite:

cargo test

Release process

This project uses automated releases through release-plz.

How releases work

  1. Make changes using conventional commits:

    • feat: for new features (minor version bump)
    • fix: for bug fixes (patch version bump)
    • feat!: or fix!: for breaking changes (major version bump)
  2. Create a pull request with your changes

  3. Merge the PR – this triggers the release-plz workflow

  4. Release PR appears – release-plz automatically creates a PR with:

    • Updated version in Cargo.toml
    • Generated changelog
    • All changes since last release
  5. Review and merge the release PR

  6. Automated publishing – when the release PR is merged:

    • release-plz publishes the crate to crates.io
    • Creates a GitHub release with changelog

Conventional commit examples

feat: add new proxy authentication method
fix: resolve connection timeout issues
docs: update API documentation
chore: update dependencies
feat!: change configuration file format (BREAKING CHANGE)

The release workflow automatically handles version bumping and publishing based on your commit messages.

Load-Aware Failover

Status: AIMD is enabled by default for priority pools with fallback enabled and at least two providers. Nonstrict, nonstreaming and exempt requests retain ordinary routing. Set aimd.enabled: false to opt a pool/model out.

Operating the controller

Existing eligible pools immediately use these defaults without a database backfill:

{
  "enabled": true,
  "latency_budget_ms": 10000,
  "breach_rate_target": 0.10,
  "recovery_breach_rate": 0.03,
  "window_samples": 100,
  "min_samples": 20,
  "share_step": 0.05,
  "share_decay": 0.8,
  "share_floor": 0.05,
  "dwell_ms": 30000,
  "idle_recovery_ms": 300000,
  "overload_statuses": [429, 503, 529]
}

The default 10-second latency budget and dwctl’s default 20-second first-token failover deadline are separate thresholds. An attempt whose first frame arrives between them keeps streaming from the preferred provider and is recorded as a breach; only an attempt with no first frame by the deadline is cut off and fails over. Breaches lower the share once the window’s breach rate exceeds the target, so sustained slowness shifts later requests to the alternates without cutting slow attempts off. Controllers and sample minima are per process, not aggregated across replicas. With the defaults:

  • Decrease the share by 20% when more than 10% of the window’s completed samples breach.
  • Hold while the breach rate is above 3% and at most 10%.
  • Increase by five percentage points once the rate has stayed at or below 3% for the 30-second dwell, and again each dwell while it stays there.
  • Idle recovery adds a step for each five minutes the pool goes without the 20 samples needed to judge, so a pool that sees little traffic after an incident climbs back instead of staying degraded until a restart.
  • The share never drops below 0.05, so real traffic keeps measuring the preferred provider. There are no synthetic probes.

Override fallback.aimd on a native onwards priority pool, or the top-level aimd field when creating/updating a dwctl priority composite. An override object replaces the prior object; omitted object members use the defaults above. For an enabled override, set an explicit first_token_timeout_ms: either 0 to disable that deadline, or at least the latency budget. An absent AIMD override inherits the controller defaults and permits an inherited deadline. If that inherited deadline is shorter than the budget, its censored result is unknown, not a budget breach. The controller never silently lengthens a configured deadline.

Dwctl PATCH semantics: omitted fields are unchanged; aimd: null restores the default controller; aimd: {"enabled": false} disables it. A null first_token_timeout_ms restores deadline inheritance. Overrides appear under fallback in model responses; null means inherited defaults, not disabled. The dashboard has no AIMD editor; use the model API for overrides and opt-out.

Only strict-mode stream: true requests without the configured timeout-exempt header use the share or contribute observations. Single-provider and weighted-selection pools, and pools without enabled fallback, are inert. Eligible preferred attempts at any position in the retry cascade, including the final attempt, can contribute; alternate attempts never do. Nonstrict/nonstreaming/exempt traffic uses ordinary selection even when the pool has a demoted share. Each preferred attempt ends as one of:

  • healthy: its first non-sentinel data frame arrived within the budget;
  • breach: that frame arrived after the budget, the first-token deadline expired at or after the budget, or the provider answered with a status listed in overload_statuses, either on the response line or embedded in a 2xx stream. A provider shedding load is the clearest overload signal there is;
  • unknown: anything else — a shorter inherited deadline, other upstream errors, network errors and provider request timeouts, non-SSE responses, empty or DONE-only streams, and cancellation.

Unknown outcomes are excluded from the breach rate. They count as neither healthy nor a breach, and they never block a decision.

Controller observations inspect the already parsed strict SSE stream, including content arriving after the lead peek’s event/time caps. They do not add buffering, change bytes, or move the failover deadline. The exported first-token histogram still has its documented lead-peek coverage; it is not the controller’s denominator. No token-content parsing is added: the first non-sentinel data frame can still be metadata. A stream that never produces data and has no armed failover deadline remains unknown until it ends or is cancelled. Latency is measured when the gateway polls the frame, so gateway scheduling or downstream backpressure can contribute; it is not a measurement of backend execution time alone.

The window holds up to window_samples completed healthy/breach outcomes, and decisions need at least min_samples. Attempts still in flight never hold a decision back, and a full window keeps admitting samples, so the controller keeps deciding under sustained concurrency. A decrease clears the window and starts a new generation: attempts that began at the old share are ignored when they complete. An increase keeps the window, so a window that is still healthy supports the next step one dwell later. Updates are constant-time under a shared pool mutex, and each process controls its own share independently.

Reloads keep a controller when its AIMD parameters, explicit deadline and preferred provider identity (URL, key, upstream model) are unchanged. When a pool is still configured for AIMD but drops to a single provider — for example because an autoscaler disabled the preferred provider — its controller is parked rather than discarded. It resumes with its learned share when a later reload restores the same preferred provider, and idle recovery credits the time it spent parked. Changing the settings, replacing or reordering the preferred provider, or opting out retires the controller, and its successor starts at 1.0. Cloned request pools share state; process restarts reset it.

Validation bounds: budget 1–3,600,000 ms; dwell 1–86,400,000 ms; idle recovery 0–86,400,000 ms (0 disables it); 2 <= min_samples <= window_samples <= 100000; target in [0,1) and 0 <= recovery_breach_rate <= breach_rate_target; decay in (0,1); step and floor in (0,1]; at most 32 overload_statuses, each 400–599. Choose windows and dwell for sample volume across individual gateway replicas, including traffic remaining at the floor.

Realtime-only failover statuses

Separately from the controller, a pool’s fallback.realtime_on_status lists upstream statuses that fail a realtime request over to the next provider, on top of on_status. Requests carrying the exempt header keep the upstream response and retry on their own terms. Dwctl stores this per model as fallback_realtime_on_status (the dashboard’s “Overloaded (529, realtime only)” switch); new composite models default to [529]. Whether or not a status fails over, a listed overload status from the preferred provider still counts as a controller breach.

Monitor client latency, errors, preferred-first share, adjustment rate and alternate spend after deployment. A low share is an indicator of capacity shortfall, not proof. Disable with aimd: {"enabled": false} to restore ordinary priority selection for new requests after the routing configuration reloads.

Problem

first_token_timeout_ms bounds how long a streamed attempt may go without producing its first frame, then fails over. It is a good hang detector and a poor latency control, because every firing is a tax: the client waits the full deadline on the first provider and then waits again for the second.

That cost is acceptable when the first provider is broken. It is not acceptable when the first provider is merely slow, which is the common case for a self-hosted upstream whose time-to-first-token is load-dependent:

  • Under load, first-token latency rises. A deadline chosen to catch stalls starts catching the ordinary upper tail instead.
  • Each failover then adds the whole deadline to a request that would have completed shortly after it.
  • The result is a cluster of client latencies just above the deadline — caused by the mitigation rather than the upstream.

Raising the deadline removes the manufactured cluster but abandons the slow tail. Lowering it converts more ordinary requests into double-waits. There is no good value, because a per-request deadline can only ever react after paying its own cost.

Why a binary circuit breaker is not enough

The obvious fix is a circuit breaker: watch first-token latency, and when it degrades, send traffic to the next provider instead of paying the deadline per request. That removes the per-request tax — the cost collapses from “every affected request” to “one probe per interval”.

But a naive breaker oscillates when the upstream’s latency is load-dependent:

  1. Latency degrades under full load; the breaker opens.
  2. All traffic moves away, so the upstream goes idle.
  3. A probe arrives at an idle upstream and is fast, so the breaker closes.
  4. Full load returns, latency degrades, and the breaker opens again.

The measurement taken at trickle load does not predict behaviour at full load, so the breaker hunts forever. This is not a tuning problem: the stable answer is usually a split — the upstream serves the share it can serve within budget, and the remainder goes elsewhere — and a two-position control cannot express a split.

Design: a controlled share

Replace the binary open/closed state with a share f ∈ [0, 1]: the fraction of eligible requests for which the preferred provider is tried first.

The control signal is a rate, not an event

A single slow request must not move the share. Any realistic first-token distribution has a tail, so at every sustainable share some requests exceed any fixed budget. If each breach triggered a decrease, the share would decay to its floor regardless of actual capacity, and it would decay faster at higher request volumes — making the control a function of traffic rather than of service quality.

The controller therefore targets a breach rate against a latency budget, which is a statement of intent that can actually be met:

  • A breach is an attempt whose first token did not arrive within latency_budget_ms, or whose provider answered with an overload status.
  • Over a sliding window of at least min_samples completed observations, compute the observed breach rate.
  • If it exceeds breach_rate_target, decrease: f *= share_decay.
  • If it has stayed at or below recovery_breach_rate for dwell_ms, increase: f += share_step.
  • Between the two, hold.

The gap between the thresholds is hysteresis. A single threshold judged on a small sample decreases on noise: at a true breach rate exactly on a 5% target, 50 samples exceed it about half the time. Separating “clearly overloaded” from “clearly healthy” lets the controller decide on smaller windows without hunting.

Decrease fast, increase slowly. The asymmetry is intended to reduce oscillation and approach the largest share whose breach rate stays within the band.

A decrease clears the window: samples taken at the old share say nothing about the new one, and a stale overloaded window must not trigger a second decrease. An increase keeps the window, because a window that is still healthy after a step is evidence for the next step too. Each adjustment is still followed by a dwell before the next, so every increase is re-measured at the higher share before another is allowed.

Properties worth preserving

  • Never remove a provider from the pool. f biases which provider is tried first. A demoted provider is still reachable, and existing failover semantics are untouched.
  • Keep a floor on f. A small non-zero share preserves a live measurement of the preferred provider, so recovery is observed from real traffic rather than synthetic probes. This matters more than it looks — see the scoping rule below, which makes preferred-provider samples the only control input.
  • Disabled means today’s behaviour. With the controller off, selection and failover behave exactly as they do now.

Where it hooks into the code

Two kinds of observation

The controller needs two distinct inputs, and conflating them would corrupt it:

  • An uncensored sample — an observed first-token latency.
  • A breach — the deadline expired. This establishes only that first-token latency exceeded the deadline, a censored lower bound. It is not a latency measurement and must never be recorded as one; feeding the deadline value into a latency histogram would bias every statistic drawn from it.

Both feed the breach-rate calculation. Only uncensored samples feed the latency histogram.

The implemented breach counter counts failover deadline expiries. The controller’s latency budget is a separate threshold: a timeout earlier than the budget cannot establish a budget breach. Controller configuration must therefore require any armed failover deadline to be at least the latency budget. A slow observed frame can establish a budget breach without firing the failover deadline. Neither outcome may count twice in the controller’s denominator.

The exported observation metrics are not sufficient controller input: even strict streams can outlast the existing peek’s event or time limits and be forwarded unobserved. The controller therefore observes eligible frames after the peek and tracks unknown outcomes explicitly, rather than counting them as healthy or estimating a rate from the exported histogram.

Only the preferred provider’s attempts are control input

Once 1 - f of traffic is being sent to alternates, those alternates are also producing first-token observations. They must not update f. A slow alternate would otherwise demote a healthy preferred provider, and a fast alternate would ramp up a struggling one — in both cases the controller would be steering on a signal from the wrong upstream.

Observations are therefore attributed to the provider actually attempted, and only attempts against the preferred provider adjust the share. Observations from alternates are still exported, because they are useful for comparing providers, but they are inert as control input.

Where an uncensored sample can be taken

SiteWhat it establishesAvailable for
Response headers arriveHeaders only — not a first tokenAll responses
Lead-frame read (read_lead_frames)First decisive SSE frameStrict-mode 2xx SSE only
Deadline expiryBreach (censored)Wherever the deadline is armed

Header arrival is not a first-token sample. For a streamed response the headers can arrive long before the first token, so it cannot stand in for one.

The lead-frame read is where the exported histogram observes first frames, and it is gated: the enclosing branch requires (200..300).contains(&status) && state.targets.strict_mode. Non-strict SSE is forwarded without a lead-frame peek, deliberately — the pass-through path avoids forcing buffering and SSE re-framing onto streams that would otherwise stream straight through.

Consequences to accept explicitly:

  • The first-token histogram is populated only for strict-mode SSE traffic. For non-strict traffic there are no uncensored samples, so a controller there would have breaches and nothing else, and could never ramp back up.
  • Extending coverage to non-strict traffic needs a pass-through-safe observer that timestamps the first data: frame without re-framing or buffering the stream, and that ignores keep-alive comments. That is a separate opt-in step, not a free extension.

What counts as a first token

classify_sse_event distinguishes Data from the [DONE] sentinel’s Done variant. read_lead_frames sets saw_data for both, preserving the existing non-empty verdict, but sets saw_content only for Data. The histogram uses saw_content, so [DONE]-only streams neither produce samples nor become retryable empty responses. Keep-alive comments do not set either flag.

A sample measures the first non-sentinel data frame, which may contain metadata rather than generated text. Observing literal token content is not implemented.

Decision: scoped to the priority strategy

Provider selection lives in load_balancer.rs: select_iter yields providers lazily, and select_excluding dispatches to select_priority (definition order, first available) or select_least_connections (lowest active/weight, weighted-random tiebreak).

The controller is defined for LoadBalanceStrategy::Priority only. Under Priority the preferred provider is unambiguous — first in definition order — and skipping it genuinely hands the first attempt to the next provider.

WeightedRandom is out of scope, for two concrete reasons:

  • Leaving the preferred provider eligible does not make it first. select_least_connections ranks by lowest active/weight and consults weights only to break ties, so the realised preferred-first rate would sit below f by an amount that varies with load. Folding f into weights does not fix this — weights there are a least-connections normaliser, not a proportional splitter.
  • Seeding the shared exclusion set is unsafe. SelectIter::next only clears exclusions when select_excluding returns None, so while any alternate remains selectable a seeded exclusion persists and the preferred provider is unreachable for the rest of that request.

The bias must therefore be a first-attempt-only choice, expressed as an explicit override of the first provider rather than by mutating the exclusion set: with probability 1 - f, begin at the next provider, then let subsequent attempts proceed exactly as they do today, with the preferred provider still reachable. Supporting WeightedRandom would require that override to carry a provider identity into a freshly initialised iterator, and is deferred.

State and lifetime

ProviderPool is cloned per request out of the Targets map, and its fields are plain values, so shared mutable state must sit behind an Arc — exactly as Provider’s active-connection counter already does.

Config reloads rebuild pools. The watcher calls adopt_provider_state on the new pool before inserting it, which carries live state across the reload.

Controller state is carried the same way, but not unconditionally. A single pool-level share has no per-provider matching, so adopting blindly would apply a share learned about one upstream to whatever now sits first in definition order. Adoption therefore requires the same preferred provider identity, AIMD parameters and explicit deadline.

A pool can also lose eligibility without its configuration changing: an autoscaler disabling a self-hosted preferred provider leaves one provider. The controller is parked in that pool, carried through further reloads, and resumed when the same preferred provider returns. Retiring it instead would reset the share every time capacity is scaled down and back up, which for a frequently scaled model means the controller never keeps what it learned.

Configuration

Configure FallbackConfig.aimd alongside first_token_timeout_ms. AIMD defaults are applied to eligible pools; model/pool overrides use these fields:

OptionMeaning
latency_budget_msFirst-token latency defining a breach
breach_rate_targetBreach rate above which the share decreases
recovery_breach_rateBreach rate at or below which the share may increase
window_samplesCompleted outcomes over which the rate is measured
min_samplesCompleted outcomes required before a latency-driven decision
share_stepAdditive increase per healthy dwell or idle interval
share_decayMultiplicative decrease when the rate is exceeded
share_floorMinimum share retained for measurement
dwell_msMinimum time between adjustments, and healthy time before an increase
idle_recovery_msInterval that earns a step while samples are too few to judge; 0 disables
overload_statusesUpstream error statuses from the preferred provider counted as breaches

Dwell and window size matter more than they look: a first-token observation only exists once the attempt produces its first token or breaches, so decisions lag the traffic that caused them. And because only preferred-provider attempts are control input, a low share yields observations slowly — which is what the share floor and idle recovery protect.

Per-model values

Dwctl stores first_token_timeout_ms and nullable JSONB aimd on deployed models. Create/update/read and both standard/composite sync paths carry them. AIMD is accepted by the API only on priority composites with fallback enabled. Rows with null overrides use the defaults above, so default changes apply to them on the next routing reload; stored overrides keep their values, and members they omit use the defaults.

Observability

The metrics recorder deliberately runs with idle-timeout and eviction disabled, because the autoscaler reads an absent onwards_model_inflight series as “genuinely zero in-flight” and evicting a long-lived stream’s gauge would tear a worker down mid-stream. Every label combination therefore persists for the lifetime of the process, so every label must be bounded by configuration rather than by traffic.

That rules out provider URLs as labels. It does not allow alias-only labels either: a composite alias can carry several named ProviderPools, whose independent controllers would otherwise write the same series and aggregate unrelated observations.

MetricTypeLabels
onwards_provider_sharegaugemodel, pool
onwards_share_adjustments_totalcountermodel, pool, direction
onwards_aimd_activegaugemodel, pool
onwards_aimd_window_samplesgaugemodel, pool
onwards_aimd_window_breach_rategaugemodel, pool
onwards_aimd_in_flightgaugemodel, pool
onwards_aimd_unknown_totalcountermodel, pool
onwards_aimd_overload_breaches_totalcountermodel, pool, status
onwards_first_token_secondshistogrammodel, pool, role
onwards_first_token_breaches_totalcountermodel, pool, role

model is the alias, matching the existing onwards_model_inflight{model} convention. pool is the resolved pool name, bounded by configuration. role is preferred or alternate, and status is bounded by overload_statuses. The controller series carry no role, since one controller governs one pool.

onwards_first_token_seconds has real histogram buckets, including an exact 10-second edge, in both the onwards recorder and dwctl’s; without them it would render as a per-process summary whose quantiles cannot be aggregated across replicas. Embedders installing their own recorder can reuse onwards::FIRST_TOKEN_SECONDS_BUCKETS.

Reading the share

A share that settles well below 1.0 is an indicator of capacity shortfall, not proof of it. The controller’s inputs are classified to keep it meaningful:

  • Slow first frames, deadline expiries at or after the budget, and overload statuses are breaches — each says the preferred provider could not serve the request in time.
  • Other failures — connection errors, non-overload error statuses, cancellations — are unknown and excluded, so an outage of a different kind does not read as a capacity limit.

Use the window gauges to tell a healthy steady state from one that has too little signal. A share at 1.0 with a low onwards_aimd_window_samples and a rising onwards_aimd_unknown_total means the controller cannot see enough outcomes to judge, not that the provider is healthy. onwards_aimd_window_breach_rate shows where the pool sits relative to the two thresholds.

Named continuation pools keep their deterministic failover order and explicitly opt out of AIMD. Catalog provisioning preserves API overrides while routing remains compatible, and clears enabled AIMD overrides when the catalog changes to weighted routing or disables fallback.

The controller series are updated by eligible requests and observations. onwards_aimd_active drops to 0 when a controller is parked or retired. As with the other non-evicting metrics, a disabled or idle pool retains its last published share; read it together with onwards_aimd_active and recent adjustment activity.

Validation

Deterministic controller tests cover the hysteresis band, dwell, fresh generations after decreases, window retention after increases, idle recovery, parking and resumption, unknown exclusion, decisions with attempts in flight, overload statuses, share limits and a load-dependent capacity simulation. Selection tests cover alternate-first cascades, concurrency guards, reload adoption and parking across a disabled preferred provider. HTTP tests cover late frames, censored deadlines, overload statuses from the preferred provider, realtime-only failover statuses, exempt/nonstrict traffic, byte preservation and upstream errors. Database/API tests cover round trips, merged PATCH validation, clearing and onwards sync.