Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

API Usage

List available models

Get a list of all configured targets in the OpenAI models format:

curl http://localhost:3000/v1/models

Sending requests

Send requests to the gateway using the standard OpenAI API format:

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

The model field determines which target receives the request.

Model override header

Override the target using the model-override header. This routes the request to a different target regardless of the model field in the body:

curl -X POST http://localhost:3000/v1/chat/completions \
  -H "model-override: claude-3" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

This is also used for routing requests without bodies – for example, to get the embeddings usage for your organization:

curl -X GET http://localhost:3000/v1/organization/usage/embeddings \
  -H "model-override: claude-3"

Metrics

When the --metrics flag is enabled (the default), Prometheus metrics are exposed on a separate port:

curl http://localhost:9090/metrics

See Command Line Options for metrics configuration flags.

Rejected requests

Every client error that onwards decides on itself, such as an unknown model, a failed reasoning check, a strict-mode schema error or a rate limit, increments onwards_rejections_total{model, status, code, traffic}:

  • code is the error code returned to the client. Strict-mode body errors carry no code, so they are counted as:
    • invalid_json: malformed JSON;
    • schema_mismatch: valid JSON that doesn’t match the schema;
    • invalid_content_type: a /v1/responses body not sent as JSON;
    • invalid_body: a /v1/responses body that couldn’t be read.
  • model is the configured model the request names, without any serving-class suffix such as :interactive. It is empty when the request names no configured model.
  • traffic is dispatched for requests carrying the first-token-timeout exempt header and realtime otherwise.

Each rejection is also logged at info with its status, code, parameter, model, account and API key ID. Values and request bodies are never logged. A client error that reports an upstream’s response isn’t counted.