> Raw Markdown twin (generated at build time from the source Markdown). Rendered page: https://docs.gatellm.io/en/reference/configuration · Doc index: https://docs.gatellm.io/en/llms.txt


# Environment variable reference

The official image already ships a built-in runtime configuration — you do **not** need to write or mount any config file; just override the items you want via environment variables. `docker run -e`, docker-compose's `environment:` section, and k8s container env all work. After changing them, `docker restart <container-name>` takes effect.

This page lists **all** environment variables the image accepts, their defaults, and value descriptions. The default is the image's built-in value (written in the Dockerfile's `ENV`); if unset, that default applies.

Business configuration (upstreams, models, access keys, key groups, load balancers, MCP, ACL, scripts, SSO) is **not** here — it lives in the database and is managed by the console, taking effect immediately upon change, unrelated to environment variables.

## Set these first

Under the zero-config main path the image bootstraps the console and admin; the following only need to be set when you have the corresponding requirement:

- [`ENCRYPTION_KEY`](#storage-multi-instance) — set before configuring upstream API keys or issuing access keys, otherwise saving errors with `encryption_key not set in config`; generate with `openssl rand -hex 32`
- [`CONSOLE_PASSWORD`](#console) — to preset the console admin password instead of the random one printed in the logs (**only takes effect on first startup**)
- [`LICENSE_KEY`](#image-level-vars) — activate the license to unlock multi-instance / Redis / PostgreSQL and higher memory limits
- [`STORAGE_MODE`](#storage-multi-instance) together with [`POSTGRES_URL`](#storage-multi-instance) and [`REDIS_URL`](#storage-multi-instance) — required for multi-instance deployment
- [`RESET_ADMIN`](#image-level-vars) — a one-time reset when you forget the console password

## Two rules for values and defaults

**1. Empty ≠ unset ≠ code default.** Setting a variable to an empty string (`VAR=`) means "off / none" and is interpreted as `None`. In the image you cannot "unset" a variable to fall back to some deeper code default — the image's `ENV` is the default. For example: the code's built-in default for `STREAM_IDLE_TIMEOUT_SECS` is 300 seconds, but the image ships 600 seconds; to get 300, write `STREAM_IDLE_TIMEOUT_SECS=300` explicitly rather than leaving it empty.

**2. List-type variables use comma-separated strings.** `CORS_ORIGINS`, `CORS_HEADERS`, `CORS_METHODS`, `CORS_EXPOSE_HEADERS`, `FALLBACK_DNS_SERVERS`, and `TRUSTED_PROXIES` all accept strings in the form `a,b,c` (auto-split on commas, trimmed, empty items dropped) and also accept true arrays. All other variables are single-valued.

## Listening and request limits {#listen-and-request-limits}

| Variable | Default | Purpose and values |
|------|--------|------|
| `HOST` | `0.0.0.0` | Listen address; `127.0.0.1` means localhost only |
| `PORT` | `7890` | Listen port |
| `MAX_REQUEST_SIZE_MB` | `50` | Maximum request body (MB), applies to all routes |
| `MAX_RESPONSE_BODY_MB` | `25` | Upstream response body limit (MB, non-streaming only) |
| `WORKER_THREADS` | `2` | Number of async runtime worker threads. **2 is the safe lower bound for 0.25 vCPU, not a recommendation** — increase it with available CPU (e.g. set 4 for 4 vCPUs) |
| `REQUEST_TIMEOUT_SECS` | empty (= unlimited) | Per-request total duration limit; empty means unenforced |
| `GRACEFUL_SHUTDOWN_TIMEOUT_SECS` | empty | Wait limit for graceful shutdown after receiving a signal |
| `PRE_STOP_DELAY_SECS` | empty | Delay before shutdown, giving the load balancer time to drain traffic |
| `DATA_PLANE_MAX_CONNECTIONS` | `0` | Data-plane connection budget: the maximum number of in-flight requests the proxy routes allow; requests exceeding the limit are rejected **before being written to disk** with an immediate 429 + Retry-After; console / health checks / monitoring are unaffected. `0` = derived automatically at startup from the fd (RLIMIT_NOFILE) limit; the effective value is printed in the startup log. The fd limit itself should be raised — see the [fd section of the Docker single-node tutorial](/en/quickstart/docker-single-node.md#raise-fd-limit) |
| `CONNECTION_BUDGET_RESERVE_FDS` | `0` | When deriving the connection budget automatically, the number of fds reserved for connections opened lazily after startup (PG pools, Redis, log queue); `0` = enumerate automatically |

## Storage and multi-instance {#storage-multi-instance}

| Variable | Default | Purpose and values |
|------|--------|------|
| `STORAGE_MODE` | `postgresql` | Persistence backend: `sqlite` (single-node) / `postgresql` (shared across instances). The image defaults to postgresql + empty `POSTGRES_URL`; the zero-config path only works because the license layer force-falls-back to sqlite when unlicensed |
| `SQLITE_PATH` | `/var/lib/protoflux/stats.sqlite` | SQLite file path (single-node mode) |
| `POSTGRES_URL` | empty | PostgreSQL connection string; the `sslmode` query parameter controls TLS (`disable`/`require`/`verify-full`). Required for multi-instance |
| `POSTGRES_POOL_SIZE` | `8` | Maximum concurrent PG connections |
| `POSTGRES_CONSOLE_POOL_SIZE` | `4` | A separate connection pool for console (login/user/audit) queries, isolated from the data plane — the management entry stays reachable even when the data plane saturates the main pool |
| `POSTGRES_WAIT_TIMEOUT_SECS` | `10` | Connection-pool acquisition timeout; `0` waits forever |
| `POSTGRES_STATEMENT_TIMEOUT_SECS` | `30` | Server-side single-statement timeout for the stats pool; `0` = unlimited |
| `REDIS_URL` | empty | Redis connection string (multi-instance: shared sessions / rate limits / log broadcast / IP bans); if unset, everything is in-memory |
| `REDIS_KEY_PREFIX` | `protoflux:` | Redis key prefix, to distinguish multiple gateway deployments sharing one Redis |
| `ENCRYPTION_KEY` | empty | Encryption key for sensitive data at rest (access keys, upstream API keys); generate with `openssl rand -hex 32` |

::: danger About ENCRYPTION_KEY
- Inject it only via `-e` / secret — do **not** write it into the image layer or compose plaintext
- **Be sure to back up** this key: losing it = already-encrypted data is **permanently unreadable**
- In a multi-instance deployment, all instances **must** use the same key
- The gateway can start with an empty value, but the moment it needs to read/write an encryption key it fail-closes with an error — set it before configuring upstream API keys
:::

## Console {#console}

| Variable | Default | Purpose and values |
|------|--------|------|
| `CONSOLE_ENABLED` | `true` | Whether the console is enabled; `false` means no admin is created and `/console` is unreachable |
| `CONSOLE_PASSWORD` | empty | Initial password for the console admin `protoflux`, **only takes effect on first startup (when the console user table is empty)**; later changes are ignored. If empty, a strong random password is generated on first startup and printed once to the log |
| `CONSOLE_SECRET_KEY` | empty | Bearer token for the console API (long-lived, for scripts/CI to call `/console/api/*` directly; matching it grants admin rights). **Empty = this auth method is disabled and not auto-generated** (what is auto-generated is the admin's initial password via `CONSOLE_PASSWORD`). When non-empty it must be ≥ 12 characters |
| `CONSOLE_ALLOW_REMOTE` | `true` | Whether remote access to the console is allowed. Set `false` and access from the host under Docker Desktop/bridged networking returns 403 |
| `CONSOLE_MAX_FAILURES` | `5` | Consecutive login-failure limit; reaching it bans the source IP |
| `CONSOLE_BAN_DURATION` | `300` | Login ban duration (seconds) |
| `CONSOLE_IP_BAN_ENABLED` | `true` | Whether IP banning is enabled |
| `CONSOLE_SESSION_AUTO_RENEW` | `true` | Whether sessions auto-renew |

## CORS and network trust

| Variable | Default | Purpose and values |
|------|--------|------|
| `CORS_ORIGINS` | empty | Allowed cross-origin origins (comma-separated); empty = browser CORS stays off |
| `CORS_HEADERS` | empty | Allowed request headers (comma-separated) |
| `CORS_METHODS` | empty | Allowed methods (comma-separated) |
| `CORS_EXPOSE_HEADERS` | empty | Response headers the frontend may read (comma-separated) |
| `CORS_MAX_AGE` | `7200` | Preflight result cache seconds |
| `CORS_CREDENTIALS` | `false` | Whether credentials are allowed |
| `TRUSTED_PROXIES` | empty | Trusted proxy IP/CIDR list (comma-separated, e.g. `10.0.0.0/8,192.168.0.0/16`); only when set is `X-Forwarded-For` trusted to resolve the real client IP |
| `IP_RATE_LIMIT_RPM` | empty (= unlimited) | Global per-minute request limit per IP; empty means no per-IP rate limiting |
| `METRICS_AUTH_TOKEN` | empty | Auth token for the `/metrics` endpoint; when set, `Authorization: Bearer <token>` is required |

## Upstream retry and routing affinity {#upstream-retry-affinity}

| Variable | Default | Purpose and values |
|------|--------|------|
| `UPSTREAM_IDLE_CONNECTIONS` | `16` | Connection pool size per upstream host |
| `UPSTREAM_SEND_TIMEOUT_SECS` | `180` | Timeout for sending the request body + waiting for response headers (TTFB) |
| `UPSTREAM_READ_TIMEOUT_SECS` | `600` | Per-read timeout for a single response chunk (reset each chunk); must be greater than `STREAM_IDLE_TIMEOUT_SECS` |
| `UPSTREAM_USER_AGENT` | `Protoflux` | Default User-Agent used toward upstreams |
| `FALLBACK_DNS_SERVERS` | `1.1.1.1,8.8.8.8,119.29.29.29,223.5.5.5` | Fallback DNS pool for upstreams without their own DNS config (comma-separated); empty `=` disables the fallback |
| `MAX_UPSTREAM_RETRIES` | `3` | Maximum upstream retries per request (before the first byte) |
| `MAX_KEY_ROTATIONS` | `0` | Maximum key rotations per request (when an upstream has multiple keys); `0` = unlimited (rotate through all available keys) |
| `LB_AFFINITY_TTL_SECS` | `300` | Load-balancing affinity binding lifetime in seconds; `0` = permanent binding |
| `LOAD_WINDOW_SECS` | `60` | Window seconds for computing upstream load |
| `KEY_BINDING_TTL_SECS` | `1800` | Lifetime in seconds of the API-key-to-upstream auto binding |
| `UPSTREAM_BAD_LINK_BUDGET` | empty (= auto) | Per-upstream **bad-link budget**: the maximum number of in-flight requests on that upstream that may pile up after exceeding the threshold without receiving a first response. Once the budget is reached the upstream is excluded from candidates; when all candidates are saturated the request is rejected immediately with 503 + Retry-After instead of queuing and dragging down healthy traffic. Empty = auto, half the initial permit count; `0` = disabled. Pure real-time in-flight counting — a request's slot is released the moment it finishes (any way), and the upstream is back to full capacity as soon as it recovers, with no failure memory |
| `UPSTREAM_BAD_LINK_THRESHOLD_SECS` | `60` | Bad-link threshold (seconds): an in-flight request that has received no first response beyond this duration counts as one bad link for that upstream |
| `UPSTREAM_DISPATCH_DEADLINE_SECS` | `300` | Total time budget for a single request's **entire outbound attempt chain** (retries + key rotations + connection-pool recovery). When exceeded, the attempt chain is terminated and 504 + Retry-After is returned — a hung request's held permits are guaranteed to be released within a bounded time. `0` = disabled (escape hatch only) |

## Streaming {#streaming}

| Variable | Default | Purpose and values |
|------|--------|------|
| `STREAMING_KEEPALIVE_SECONDS` | `15` | Interval for the SSE `:keep-alive` comment; must be shorter than the reverse proxy's idle timeout (Nginx defaults to 60s, so 15s is safe) |
| `STREAMING_BOOTSTRAP_RETRIES` | `1` | Retries before the first byte; no retries after the first chunk is sent |
| `STREAMING_LOG_TIMEOUT_SECS` | `3600` | Timeout for the SSE stream-log background task (releases zombie tasks whose "sender is gone") |
| `STREAM_IDLE_TIMEOUT_SECS` | `600` | Maximum idle gap between two data chunks; exceeding it drops the stream |
| `MAX_STREAM_DURATION_SECS` | `3600` | Hard cap on the total duration of a single SSE stream |

## Log retention and disk {#log-retention}

| Variable | Default | Purpose and values |
|------|--------|------|
| `LOG_LEVEL` | `info` | Log level (`error`/`warn`/`info`/`debug`/`trace`) |
| `LOG_FORMAT` | `json` | Log format (`json`/`plain`) |
| `LOG_MAX_BODY_SIZE_MB` | `25` | Per-log captured request/response body limit (MB) |
| `LOG_BODY_TO_TERMINAL` | `false` | Whether to also write the log body to stdout |
| `LOG_PERSIST_REQUEST_LOGS` | `false` | Whether to persist request logs to disk/database |
| `LOG_RETENTION_DAYS` | `7` | Request-log retention days; `0` = disable auto-cleanup |
| `OMS_RETENTION_DAYS` | `30` | OpenAI message-store retention days; `0` = disable auto-cleanup |
| `OTEL_ENDPOINT` | empty | OpenTelemetry export endpoint; empty means no reporting |
| `OTEL_SERVICE_NAME` | `protoflux` | Service name reported to OTEL |
| `SENTRY_DSN` | empty | **Backend** Sentry project DSN (Rust panic/error); empty means the backend doesn't report (compiled into the binary by default). Injected at image build time via `--build-arg SENTRY_DSN` (CI uses `secrets.SENTRY_DSN`), overridable at runtime via `-e` |
| `SENTRY_FRONTEND_DSN` | empty | **Frontend** Sentry project DSN (browser JS errors); empty means `client-config` falls back to `SENTRY_DSN` (same project for frontend and backend). Set a different project to isolate frontend noise |
| `SENTRY_ENVIRONMENT` | `production` | Sentry environment tag |
| `SENTRY_RELEASE` | empty | Sentry release identifier; empty defaults at runtime to `protoflux@<version>` |
| `SENTRY_TRACES_SAMPLE_RATE` | `0` | Sentry performance sampling rate (0.0–1.0); `0` = errors/crashes only |
| `LOG_QUEUE_DIR` | empty | Disk directory for the log queue; empty = fall back to the system temp directory. Ignored when `REDIS_URL` is set (multi-instance uses Redis Pub/Sub) |
| `LOG_QUEUE_SEGMENT_SIZE_MB` | `64` | Log-queue single-segment file size (MB) |
| `LOG_QUEUE_MAX_DISK_SIZE_MB` | `1024` | Log-queue total disk cap (MB) |
| `LOG_STREAM_BODY_MAX_DISK_MB` | `1024` | SSE stream-log body disk spill cap (MB) |
| `LOG_REQUEST_BODY_MAX_DISK_MB` | `1024` | Request-body disk storage cap (MB) |
| `LOG_ACCUMULATOR_MAX_DISK_MB` | `256` | Log-accumulator disk cap (MB) |
| `LOG_STREAM_BODY_BATCH_SIZE_KB` | `64` | SSE stream-log body batch write size (KB) |
| `LOG_STREAM_BODY_LINGER_MS` | `5` | SSE stream-log body linger milliseconds |
| `LOG_STREAM_BODY_CHANNEL_CAPACITY_CHUNKS` | `2048` | SSE stream-log body channel capacity (chunks) |

## Memory admission and spill {#memory-admission-spill}

Under memory pressure the gateway uses admission control + disk spill to keep the process from being killed. Defaults suffice in most cases; only tune under burst load or container OOM.

| Variable | Default | Purpose and values |
|------|--------|------|
| `MEMORY_SOFT_LIMIT_MB` | `0` | Soft memory limit (MB); `0` = no soft limit (the container cgroup is the backstop) |
| `MEMORY_QUEUE_TIMEOUT_SECS` | `300` | Timeout for a memory-queue item waiting for admission |
| `MEMORY_QUEUE_MAX_DEPTH` | empty (= auto) | Depth cap of the slow-path **waiting queue** (counts requests waiting for a processing permit, not concurrent processing — four states): empty = auto, initial permit count × 4 (bounded safe default); `-1` = unlimited (explicit escape hatch); `0` = zero wait (no queueing allowed; requests that can't get a permit are rejected immediately); `n` = cap of n. Requests beyond the cap are rejected immediately with 429 + Retry-After (reason `queue_full`), without waiting for `MEMORY_QUEUE_TIMEOUT_SECS`. Concurrent processing capacity is determined by processing permits, not limited by this item |
| `SPILL_WRITE_CONCURRENCY` | `4` | Concurrent write count for disk spill |
| `SPILL_WRITE_STALL_TIMEOUT_SECS` | `10` | Entry write-stall timeout (seconds); reaching it with zero write progress while a slot is waiting rejects with 503 + Retry-After (hung-volume defense), range 1–300 |
| `MEMORY_HOLD_SECS` | `5` | Hold seconds for a memory item |
| `MEMORY_RATE_ESCALATE_MB` | `40` | Memory rate-escalation threshold (MB) |
| `MEMORY_ADMISSION_ENABLED` | `true` | Whether admission control is enabled. Feed-forward admission makes per-request structural decisions (admit / spill-to-queue), closing the burst blind spot where a batch of requests all read a stale Normal RSS and then pass through with zero cost |
| `MEMORY_FORCE_SPILL_BODY_MB` | `0` | Force-spill body threshold (MB); `0` = don't force |

> Note: `MEMORY_ADMISSION_HANDLER_PERMIT` (per-handler MB budget per slot) has been removed — concurrency is no longer derived from a static budget but driven by the predictive `ConcurrencyController` (`target = budget / measured per-request cost`). To pin concurrency manually, use the TOML `memory_admission_handler_permits` (>0 disables the controller).

## Script sandbox

For script-sandbox limits see [Script runtime limits and configuration](/en/reference/scripting-config.md).

| Variable | Default | Purpose and values |
|------|--------|------|
| `SCRIPT_MAX_OPERATIONS` | `2000` | Maximum operations per QuickJS script execution |
| `SCRIPT_MEMORY_LIMIT_MB` | `64` | Interpreter heap cap (MB) |
| `SCRIPT_LAZY_BODY` | `true` | Whether to lazily project the request/response body |

## Per-key rate limiting

Field descriptions are in [Audit and security configuration](/en/reference/audit-and-security-config.md).

| Variable | Default | Purpose and values |
|------|--------|------|
| `RATE_LIMIT_ENABLED` | `false` | Whether per-key rate limiting is enabled |
| `RATE_LIMIT_RPM` | `120` | Per-key per-minute request limit |
| `RATE_LIMIT_WINDOW_SECS` | `60` | Rate-limit window seconds |

## Image-level variables {#image-level-vars}

These are read directly by the process and are not in the config-field table.

| Variable | Default | Purpose and values |
|------|--------|------|
| `LICENSE_KEY` | empty | License key (Ed25519). Empty = unlicensed: 512MB memory cap and Redis/PostgreSQL disabled. Activating unlocks multi-instance and higher memory limits |
| `RESET_ADMIN` | empty | One-time reset when you forget the console password: set it to a new password and restart to change the admin `protoflux` password to that value. **Takes effect only once** (the process writes a one-time marker) — see [Docker single node](/en/quickstart/docker-single-node.md) |
| `TZ` | `UTC` | Container timezone |
| `MALLOC_CONF` | `background_thread:true,narenas:1,...` | jemalloc config (background threads, arena count, dirty-page reclaim, profiling) |
| `DEV` | empty | Presence-only (setting it enables it, value irrelevant). Proxies `/console/*` to the Vite dev server — local development only |

## Config-file-only fields {#config-file-only-fields}

The following fields **can only be configured via `server.toml`**; the official image does not expose them as environment variables, so they are not in the env-var tables above. They can be set when deploying via config file (not the image env-var path):

| Field (TOML) | Default | Purpose |
|------|--------|------|
| `server.max_global_concurrency` | `1000` | Global concurrency cap for all proxy routes; reaching it returns 503 `overloaded_error`; set `0` to disable |
| `streaming.max_concurrent_streams` | `200` | Global concurrency cap for SSE streaming requests; reaching it returns 503 |
| `server.script_pool_size` | `4` | Number of concurrent script execution slots; each slot's heap is bounded by `SCRIPT_MEMORY_LIMIT_MB` |
| `sso_credentials.refresh_enabled` | `true` | Master switch for the upstream SSO credential background refresh task; hot-read each scan cycle, can be stopped without downtime |
| `sso_credentials.refresh_interval_secs` | `60` | Refresh scan interval seconds (lower bound 5) |
| `sso_credentials.refresh_lead_secs` | `300` | Refresh lead window: credentials expiring within this many seconds are refreshed immediately |
| `sso_credentials.refresh_max_parallel` | `4` | Global concurrent refresh-call cap per scan (per-credential concurrency is always 1, not configurable) |

These concurrency/slot caps are **not fixed values**; the sources of 503 are in [Error codes → Sources of 503](/en/reference/error-codes.md#source-of-503).

## Per-section reference pages

- **Logs**: [Logs and body storage](/en/reference/logs-and-body-storage.md) — body capture, SSE disk spill
- **Scripts**: [Script runtime limits and configuration](/en/reference/scripting-config.md) — script sandbox limits, error modes, script test entry
- **Security**: [Audit and security configuration](/en/reference/audit-and-security-config.md) — IP banning, password policy, security response headers
- **Pricing**: [Pricing and billing fields](/en/reference/pricing-and-billing-config.md) — pricing baseline, price-snapshot semantics, monthly billing fields
- **MCP**: [MCP configuration](/en/reference/mcp-config.md) — MCP configuration on the console/database side
- **Load balancing**: [Load-balancing fields](/en/reference/load-balancing-fields.md) — retry fields
- **SSO enterprise login**: [Configure SSO enterprise login](/en/howto/configure-sso.md) — console configuration; no environment variables needed

## FAQ

**Q: I want to change a deeper parameter that isn't on this page.**
This page is the full set of environment variables the image exposes. A few more fields can only be configured via `server.toml` (see [config-file-only fields](#config-file-only-fields)); deeper tuning items are not exposed — contact support if needed.

**Q: How do changes take effect?**
After changing environment variables, `docker restart <container-name>` suffices. Console-side business configuration (upstreams, models, keys, etc.) takes effect immediately.

**Q: Startup reports a config error.**
The image self-checks configuration at startup; errors in the log give the specific field and reason. Change the corresponding environment variable as prompted, then `docker restart <container-name>`.

**Next**: [Logs and body storage](/en/reference/logs-and-body-storage.md) / [Script runtime limits and configuration](/en/reference/scripting-config.md) / [Audit and security configuration](/en/reference/audit-and-security-config.md) / [Pricing and billing fields](/en/reference/pricing-and-billing-config.md) / [MCP configuration](/en/reference/mcp-config.md) / [Load-balancing fields](/en/reference/load-balancing-fields.md) for specific config references.
