Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Configuration Reference

Reference guide for configuring the Control Layer using YAML and environment variables.

The Control Layer uses a YAML configuration file with environment variable overrides for all settings.

Configuration File Location

The configuration file is named config.yaml by default. The system checks:

  1. Path specified via --config CLI flag
  2. Path in DWCTL_CONFIG environment variable
  3. ./config.yaml in the current directory

For Docker deployments, mount your config at /app/config.yaml.

Environment Variable Overrides

Any setting can be overridden with environment variables prefixed with DWCTL_. Nested keys use double underscores:

DWCTL_PORT=8080
DWCTL_AUTH__NATIVE__ENABLED=false
DWCTL_DATABASE__POOL__MAX_CONNECTIONS=20

The special DATABASE_URL variable (no prefix) sets the database connection string.

Core Settings

Server

host: "0.0.0.0"
port: 3001
FieldTypeDefaultDescription
hoststring"0.0.0.0"Network interface to bind to. Use 127.0.0.1 for local-only access.
portinteger3001TCP port for the HTTP server.

Secret Key

secret_key: "your-secret-key-here"

Required when native authentication is enabled. Used for JWT signing. Generate with:

openssl rand -base64 32

Admin User

admin_email: "admin@example.com"
admin_password: "change-me-in-production"

Created on first startup if it doesn’t exist. The admin user has the PlatformManager role.

Danger

Change the default admin password immediately in production!

Model Provisioning

model_provisioning:
  enabled: true
  directory: /app/model-provisioning.d
  org_overlays_directory: /app/org-overlays.d
  account_limits_directory: /app/account-limits.d
FieldTypeDefaultDescription
enabledbooleanfalseApply the declarative model catalog during startup.
directorypath/app/model-provisioning.dDirectory containing one .yaml or .yml document per canonical model.
org_overlays_directorypath/app/org-overlays.dDirectory containing one .yaml or .yml document per organisation with per-model serving overrides (default_class or explicit targets, self_hosted_only). Applied after the model catalog when enabled is set. A missing directory is an empty catalog, so a deployment that does not mount one applies no overlays; an existing path that is not a directory fails startup.
account_limits_directorypath/app/account-limits.dDirectory containing one .yaml or .yml document per account with its own realtime in-flight limits on virtual models. Applied after the model catalog when enabled is set, replacing every stored per-account limit. A missing directory changes nothing; an empty one clears every per-account limit. See Account limits.

When enabled, the directory must exist. Startup fails before any provisioning writes if loading, validation, or a referenced endpoint/group lookup fails. See Model Provisioning for the document format and ownership rules.

Database Configuration

The Control Layer requires PostgreSQL. Two modes are available:

database:
  type: external
  url: "postgres://user:pass@localhost:5432/control_layer"
  replica_url: "postgres://user:pass@replica:5432/control_layer"  # Optional
  pool:
    max_connections: 10
    min_connections: 0
    acquire_timeout_secs: 30
    idle_timeout_secs: 600
    max_lifetime_secs: 1800
FieldTypeDefaultDescription
urlstring-PostgreSQL connection URL.
replica_urlstring-Optional read replica URL.
pool.max_connectionsinteger10Maximum connections in pool.
pool.min_connectionsinteger0Minimum idle connections.
pool.acquire_timeout_secsinteger30Max wait for a connection.
pool.idle_timeout_secsinteger600Close idle connections after N seconds.
pool.max_lifetime_secsinteger1800Maximum connection lifetime.

Embedded Database

For development or single-node deployments, use the embedded PostgreSQL database:

database:
  type: embedded
  data_dir: ".dwctl_data/postgres"
  persistent: false
FieldTypeDefaultDescription
data_dirstring-Directory for database files.
persistentbooleanfalsePersist data between restarts.

Component Databases

The batch processing system (Fusillade) and request logging (Outlet) can use separate databases or schemas:

database:
  # ... main database config ...
  fusillade:
    mode: schema      # Use schema in main database
    name: "fusillade"
    pool:
      max_connections: 20
      min_connections: 2
  outlet:
    mode: dedicated   # Use separate database
    url: "postgres://user:pass@localhost:5432/outlet"
    pool:
      max_connections: 5

Authentication Configuration

At least one authentication method must be enabled.

Native Authentication

Username/password authentication with session cookies:

auth:
  native:
    enabled: true
    allow_registration: false
    password:
      min_length: 8
      max_length: 64
      argon2_memory_kib: 19456
      argon2_iterations: 2
      argon2_parallelism: 1
    session:
      timeout: "24h"
      cookie_name: "dwctl_session"
      cookie_secure: true
      cookie_same_site: "strict"
FieldTypeDefaultDescription
enabledbooleantrueEnable native login.
allow_registrationbooleanfalseAllow self-registration.
password.min_lengthinteger8Minimum password length.
password.max_lengthinteger64Maximum password length.
password.argon2_*integer-Argon2 hashing parameters. Lower values speed up tests.
session.timeoutduration"24h"Session expiration.
session.cookie_securebooleantrueRequire HTTPS for cookies.
session.cookie_same_sitestring"strict"SameSite attribute: strict, lax, or none.

Email for Password Resets

Configure email transport for password reset functionality:

File transport (development):

auth:
  native:
    email:
      type: file
      path: "./emails"
      from_email: "noreply@example.com"
      from_name: "Control Layer"
      password_reset:
        token_expiry: "30m"
        base_url: "http://localhost:3001"

SMTP transport (production):

auth:
  native:
    email:
      type: smtp
      host: "smtp.example.com"
      port: 587
      username: "noreply@example.com"
      password: "smtp-password"
      use_tls: true
      from_email: "noreply@example.com"
      from_name: "Control Layer"
      password_reset:
        token_expiry: "30m"
        base_url: "https://app.example.com"

Proxy Header Authentication

For use with identity-aware proxies (SSO):

auth:
  proxy_header:
    enabled: false
    header_name: "x-doubleword-user"
    email_header_name: "x-doubleword-email"
    groups_field_name: "x-doubleword-user-groups"
    provider_field_name: "x-doubleword-sso-provider"
    auto_create_users: true
    import_idp_groups: false
    blacklisted_sso_groups:
      - "external-contractors"
FieldTypeDefaultDescription
enabledbooleanfalseEnable proxy header auth.
header_namestring"x-doubleword-user"Header with unique user ID.
email_header_namestring"x-doubleword-email"Header with user email.
groups_field_namestring"x-doubleword-user-groups"Header with comma-separated groups.
auto_create_usersbooleantrueCreate users automatically.
import_idp_groupsbooleanfalseSync groups from IdP.
blacklisted_sso_groupslist[]Groups to exclude from import.

Default User Roles

Roles assigned to new users (admin excluded):

auth:
  default_user_roles:
    - StandardUser
    - BatchAPIUser
    - BackgroundInferenceUser

Available roles:

  • StandardUser - Base access (always included, cannot be removed)
  • RequestViewer - Read-only access to request logs
  • BillingManager - Credit and billing management
  • BatchAPIUser - Batch file and job management
  • BackgroundInferenceUser - Spare-capacity background inference submission

Security Settings

auth:
  security:
    jwt_expiry: "24h"
    cors:
      allowed_origins:
        - "https://app.example.com"
      allow_credentials: true
      max_age: 3600
      exposed_headers:
        - "location"
FieldTypeDefaultDescription
jwt_expiryduration"24h"JWT token lifetime. Must be 5min-30days.
cors.allowed_originslist["http://localhost:3001"]Allowed CORS origins. Use "*" for any.
cors.allow_credentialsbooleantrueAllow credentials. Cannot use with wildcard origin.
cors.max_ageinteger3600Preflight cache duration (seconds).

Warning

In production, ensure your frontend URL is listed in allowed_origins.

Credits & Payments

Initial Credits

credits:
  initial_credits_for_standard_users: 10.00

Credits given to new users on creation. Set to 0 to disable.

Payment Provider

Stripe (production):

payment:
  stripe:
    api_key: "sk_live_..."
    webhook_secret: "whsec_..."
    price_id: "price_..."
    host_url: "https://app.example.com"
    enable_invoice_creation: false
FieldTypeDescription
api_keystringStripe secret key (starts with sk_).
webhook_secretstringWebhook signing secret (starts with whsec_).
price_idstringStripe price ID for credit purchases (starts with price_).
host_urlstringBase URL for success/cancel redirects.
enable_invoice_creationbooleanCreate invoices for sessions.

Dummy provider (testing):

payment:
  dummy:
    amount: 50.00
    host_url: "http://localhost:3001"

Adds a fixed amount without real payment processing.

Batches Configuration

Configure the batch inference API:

batches:
  enabled: true
  allowed_completion_windows:
    - "24h"
    - "1h"
    - "48h"
  files:
    max_file_size: 104857600        # 100 MB
    default_expiry_seconds: 86400   # 24 hours
    min_expiry_seconds: 3600        # 1 hour
    max_expiry_seconds: 2592000     # 30 days
    upload_buffer_size: 100
    download_buffer_size: 100
FieldTypeDefaultDescription
enabledbooleantrueEnable /ai/v1/files and /ai/v1/batches endpoints.
allowed_completion_windowslist["24h"]SLA options users can select.
files.max_file_sizeinteger104857600Maximum upload size in bytes.
files.default_expiry_secondsinteger86400Default file retention.

Background Services

Onwards Sync

Synchronizes configuration to the AI proxy routing layer:

background_services:
  onwards_sync:
    enabled: true

Note

Disable only if you’re not using the AI proxy functionality.

Probe Scheduler

Runs health checks against model endpoints:

background_services:
  probe_scheduler:
    enabled: true

Only runs on the leader instance when leader election is enabled.

Batch Daemon

Processes batch inference jobs:

background_services:
  batch_daemon:
    enabled: "leader"              # "always", "leader", or "never"
    claim_batch_size: 100
    default_model_concurrency: 10
    claim_interval_ms: 1000
    max_retries: 1000
    timeout_ms: 600000             # 10 minutes per request
    backoff_ms: 1000
    backoff_factor: 2
    max_backoff_ms: 10000
    stop_before_deadline_ms: 900000  # 15 min safety buffer
FieldTypeDefaultDescription
enabledstring"leader"When to run: always, leader, or never.
claim_batch_sizeinteger100Requests claimed per iteration.
default_model_concurrencyinteger10Concurrent requests per model.
max_retriesinteger1000Max retry attempts. null = unlimited until deadline.
timeout_msinteger600000Per-request timeout (10 min).
stop_before_deadline_msinteger900000Stop retrying before deadline (15 min buffer).

Model Escalation

Route requests to fallback models when approaching SLA deadlines:

background_services:
  batch_daemon:
    model_escalations:
      "llama-3.1-70b":
        escalation_model: "gpt-4o-mini"
        escalation_api_key: "OPENAI_API_KEY"  # Environment variable name
    sla_check_interval_seconds: 60
    sla_thresholds:
      - name: "warning"
        threshold_seconds: 3600
        action: "log"
        allowed_states: ["pending"]
      - name: "critical"
        threshold_seconds: 900
        action: "escalate"
        allowed_states: ["pending", "claimed"]

Task Retention

Deletes expired Underway background tasks in bounded batches. Every task carries a ttl (14 days by default); nothing else removes finished tasks, so without this daemon the task table grows for the life of the installation and slows every task claim.

background_services:
  task_retention:
    enabled: true
    interval_seconds: 300
    batch_size: 1000
    batch_pause_milliseconds: 2000
    min_age_days: 14
FieldTypeDefaultDescription
enabledbooleantrueRun the retention daemon.
interval_secondsinteger300Seconds between sweeps.
batch_sizeinteger1000Rows deleted per statement; each batch is its own short transaction.
batch_pause_millisecondsinteger2000Pause between batches of one sweep, to throttle a large backlog.
min_age_daysinteger14Tasks younger than this are never considered. A task’s own longer ttl is still honoured.

Every instance runs the daemon; an advisory lock collapses concurrent sweeps to one. Only succeeded and failed tasks are deleted; pending and in_progress tasks are preserved. Each batch has a 2-second lock timeout and a 30-second statement timeout. The daemon uses direct database connections for its session-level advisory lock. For a very large existing backlog, run scripts/purge_underway_tasks.sh once beforehand so the first sweeps stay short.

Prompt Cache Retention

Deletes prompt-cache entries that expired more than a grace period ago, in bounded batches. Expired entries serve no request, but nothing else removes them, so without this daemon the prefix table keeps every prefix ever written.

background_services:
  prompt_cache_retention:
    enabled: false
    interval_seconds: 300
    batch_size: 1000
    batch_pause_milliseconds: 2000
    grace_days: 7
FieldTypeDefaultDescription
enabledbooleanfalseRun the retention daemon. Off by default so a deployment chooses when to start clearing an existing history.
interval_secondsinteger300Seconds between sweeps.
batch_sizeinteger1000Rows deleted per statement; each batch is its own short transaction.
batch_pause_millisecondsinteger2000Pause between batches of one sweep, to throttle a large backlog.
grace_daysinteger7Days an entry is kept after it expires. Usage recompute re-derives a past request’s cache split from the entries live at the time, so values below 7 are rejected.

Every instance runs the daemon; an advisory lock collapses concurrent sweeps to one. Entries are deleted oldest-expiry first; an entry that a request is refreshing or re-writing at that moment is skipped. Each batch has a 2-second lock timeout and a 30-second statement timeout. At the default pace a sweep deletes at most about 1.8 million entries an hour, so an installation with a long existing history clears it over several sweeps without a burst of write and vacuum load.

Leader Election

For multi-instance deployments:

background_services:
  leader_election:
    enabled: true

Uses PostgreSQL advisory locks. Only the leader runs probe scheduler and batch daemon (when set to "leader" mode).

Model Sources

Seed model endpoints on first startup:

model_sources:
  - name: "openai"
    url: "https://api.openai.com"
    api_key: "sk-..."
    sync_interval: "30s"
    default_models:
      - name: "gpt-4o"
        add_to_everyone_group: true
      - name: "gpt-4o-mini"
        add_to_everyone_group: true
FieldTypeDefaultDescription
namestring-Identifier for the model source.
urlstring-Base URL of the OpenAI-compatible API.
api_keystring-API key for authentication.
sync_intervalduration"10s"How often to refresh model list.
default_modelslist-Models to auto-import on first run.

Note

Model sources are only seeded on first startup. After that, manage endpoints through the UI or API.

Metadata

UI display settings:

metadata:
  region: "UK South"
  organization: "ACME Corp"
  title: "ACME AI Gateway"
  docs_url: "https://docs.example.com"
  docs_jsonl_url: "https://docs.example.com/jsonl"
FieldTypeDefaultDescription
regionstring-Region displayed in UI header.
organizationstring-Organization name in UI header.
titlestring-Custom browser tab title.
docs_urlstring"https://doublewordai.github.io/control-layer/"Documentation link in header.
docs_jsonl_urlstring-JSONL docs link in batch upload modal.

Realtime In-Flight Limits

Each virtual model has a default number of realtime requests one account may have in flight on it (realtime_inflight_limit), and an account limit file can give one account a different limit. Replicas share their counts through Redis:

limits:
  realtime_inflight:
    enforce: true
    redis_url: rediss://:password@limits-redis.example:6379
SettingDefaultDescription
limits.realtime_inflight.enforcefalseRefuse requests over the limit. Off, every request is admitted, so defaults and account limits can be in place before the limit takes effect.
limits.realtime_inflight.redis_urlunsetRedis holding the shared counts. Unset, or unreachable, each replica counts only its own requests.

auth.rate_limits is no longer used and is ignored if present.

Batch In-Flight Limits

A virtual model’s batch_capacity also caps how many batch requests may be in flight on it at once. Unlike the realtime limit, this is a single global count for the model, not a per-account one: batch traffic is dispatched by worker pods, so the cap has to hold across all of them. It reuses the same Redis as the realtime limit but has its own switch:

limits:
  realtime_inflight:
    enforce: true
    redis_url: rediss://:password@limits-redis.example:6379
  batch_inflight:
    enforce: true
    default_capacity: 200
SettingDefaultDescription
limits.batch_inflight.enforcefalseRefuse batch requests that would exceed a virtual model’s effective batch cap. Off, every batch request is admitted and no shared count is touched, so no cap applies however it was configured.
limits.batch_inflight.default_capacity200Global batch cap for a virtual model that does not set its own positive batch_capacity. Set to 0 (or null) for no default. Must not be negative.

A virtual model’s effective cap is its own batch_capacity when that is positive, otherwise limits.batch_inflight.default_capacity. This default only feeds the onwards cap: it does not change fusillade’s per-daemon starting concurrency, which keeps using background_services.batch_daemon.default_model_concurrency for models without an explicit batch_capacity. Because an enforced default caps every virtual model that has not set its own value, raise default_capacity (or set it per model) wherever batch runs higher than 200, or the first requests past the default will be refused.

Why realtime and batch differ

Realtime has per-account limits and no per-model cap: when a model is full, the downstream answers 529, and realtime traffic may grow into whatever capacity exists (the per-account limit keeps one tenant from consuming it all). Batch instead has a per-model cap, because batch must never crowd out realtime and can always be processed later — a refused batch request is rescheduled, not lost.

batch_capacity means two things

The same settings.batch_capacity value now has two roles:

  • fusillade’s starting per-daemon concurrency for the model, which adaptive concurrency may grow beyond, and
  • the global onwards cap on batch requests in flight on the alias, once limits.batch_inflight.enforce is on.

Enabling enforcement therefore turns batch_capacity from a per-pod starting point into a global ceiling. Raise the value to the intended global cap before turning enforcement on, or the first batch requests will be refused against the old, smaller starting point.

The cap is shared by every kind of dispatched request: file batches, flex requests and background requests all count against the same per-model ceiling.

Behaviour over the cap

The Redis connection is shared with limits.realtime_inflight.redis_url; reachability and per-replica fallback behave the same way. The two limits are independent and use separate key spaces, so realtime realtime_inflight_limit and batch batch_capacity can be enforced in any combination. A batch request over the cap is refused with 529 and code batch_capacity_exceeded (never 429), naming the model and its cap, so the batch dispatcher backs off without spending a retry attempt and reschedules the request later.

Observability

Metrics

enable_metrics: true

Exposes Prometheus metrics at /internal/metrics.

dwctl_model_batch_inflight_limit reports each virtual model’s effective global batch cap: its own positive batch_capacity, or the configured limits.batch_inflight.default_capacity when it has none. It reports even when limits.batch_inflight.enforce is off, and a model with no effective cap (the default turned off and no own value) reports 0, so a zero does not by itself mean the cap is enforced — check the config switch.

Request Logging

enable_request_logging: true

Logs all AI proxy requests and responses to PostgreSQL. Disable if you have sensitive data.

OpenTelemetry

enable_otel_export: false

Exports traces via OTLP. Configure the exporter endpoint with standard OpenTelemetry environment variables (OTEL_EXPORTER_OTLP_ENDPOINT, etc.).

When request analytics and tracing are enabled, the pair http_analytics.trace_id and gateway_span_id identifies the gateway span for a captured request. Follow its descendant onwards.provider_attempt spans to the downstream serving spans. The gateway span’s doubleword.request_id attribute holds the logical request UUID for reverse lookup in request records.

Trace context must be propagated on every required forwarding hop. Propagation does not itself enable span export: correlation also depends on the connecting spans being sampled, exported, and retained. Stored IDs can therefore refer to unavailable spans, and older analytics records may have no gateway span ID.

Sample Files

Generate sample JSONL files for new users:

sample_files:
  enabled: true
  requests_per_file: 2000

Validation

The system validates configuration on startup and fails if:

  • Native auth is enabled but secret_key is missing
  • No authentication method is enabled
  • jwt_expiry is outside 5min-30day range
  • CORS uses wildcard origin with credentials enabled
  • Database URL is invalid or unreachable

Run validation without starting the server:

dwctl --config config.yaml --validate