load_balancer
On this page
Selects an upstream endpoint using the cluster’s configured strategy.
Configuration Notes
Supported strategies: - round_robin (default): cycles through endpoints in order, respecting weights via endpoint expansion. - least_connections: picks the endpoint with the fewest active in-flight requests; decrements the counter on on_response. - p2c: samples two random endpoints and picks the less loaded one. - random: picks a uniformly random endpoint, weighted by endpoint weight. - consistent_hash: hashes a configurable request header (or the URI path when the header is absent) to pin requests to a stable endpoint. - maglev: hashes a configurable request header (or the URI path) through a Maglev lookup table for even distribution and minimal disruption when endpoints change.
Configuration
| Field | Type | Required | Description |
|---|---|---|---|
clusters | Cluster[] | no | Cluster definitions. |
clusters[].name | string | yes | Unique name for the cluster. |
clusters[].http | ClusterHttpOptions | no | HTTP-specific cluster options. Grouped under http: so the shared cluster/upstream transport types stay protocol-agnostic; TCP load balancers ignore this block entirely. |
clusters[].http.authority | string | no | Override the upstream HTTP Host header. When set, the proxy rewrites the Host header sent to the upstream instead of forwarding the downstream value. The downstream HTTP/2 :authority pseudo-header is never forwarded upstream; on an HTTP/2 upstream leg Pingora rebuilds :authority from this Host value. TLS SNI remains independent — configure tls.sni separately when needed. Must be a valid HTTP authority: a hostname with an optional port, or a bracketed IPv6 address with an optional port. URI schemes, paths, userinfo, and fragments are rejected. |
clusters[].http.application_protocol | string | no | Opaque application protocol the upstream cluster expects. Declares the wire representation of the request body (for example openai_chat_completions or openai_responses). The value stays opaque to Praxis core — consuming filters interpret it; Praxis defines no enum of known protocols. Validated as a bounded, canonical identifier: 1–64 bytes of lowercase ASCII letters, digits, ., _, or -, starting and ending with a letter or digit. The open string type is deliberate: the protocol set is open-ended and owned by consuming filters, so keep this a string — do not convert it to an enum. |
clusters[].http.application_provider | string | no | Opaque application provider refining application_protocol. Distinguishes provider-specific semantics (for example openai or vllm) independent of the deployed cluster, whose identity is already the cluster name. Stays opaque to Praxis core and follows the same identifier rules — and the same enum-free rationale — as application_protocol. |
clusters[].http.version | h1 | h2 | auto | no | HTTP version used for upstream connections to this cluster. Defaults to [H1]. Set h2 for gRPC upstreams: response trailers, and therefore grpc-status, only exist on an HTTP/2 leg. |
clusters[].connection_timeout_ms | integer | no | TCP connection timeout in milliseconds. Applies to the TCP handshake only (before TLS). When exceeded, the connection attempt fails and the load balancer may retry on the next endpoint. None (the default) uses Pingora’s built-in timeout. |
clusters[].endpoints | (Simple | Weighted)[] | yes | List of endpoints for the cluster. Each entry is either a plain "host:port" string or a { address, weight } object. |
clusters[].endpoints[].address | string | yes | Socket address as host:port. |
clusters[].endpoints[].weight | integer | no | Relative forwarding weight. Higher values receive proportionally more traffic. Defaults to 1. |
clusters[].endpoints[].metadata | object<string, string> | no | Arbitrary key-value metadata for subset-based load balancing. |
clusters[].endpoints[].priority | integer | no | Priority tier (0 = primary, 1 = first failover, etc.). |
clusters[].endpoints[].zone | string | no | Locality zone identifier for zone-aware routing. |
clusters[].health_check | HealthCheckConfig | no | Active health check configuration for this cluster. |
clusters[].health_check.type | http | tcp | grpc | yes | Probe type: [Http], [Tcp], or [Grpc]. |
clusters[].health_check.expected_status | integer | no | Expected HTTP status code for a healthy response. |
clusters[].health_check.grpc_service | string | no | gRPC service name to check. Only used by the grpc probe. Empty (the default) asks for the server’s overall serving status, which is what the grpc.health.v1.Health contract defines an empty name to mean. |
clusters[].health_check.healthy_threshold | integer | no | Consecutive successes required to mark an endpoint healthy. |
clusters[].health_check.interval_ms | integer | no | Probe interval in milliseconds. |
clusters[].health_check.passive_healthy_threshold | integer | no | Consecutive successes to mark an endpoint healthy again via passive observation. None disables passive recovery (active checks must recover it). |
clusters[].health_check.passive_unhealthy_threshold | integer | no | Consecutive response failures (5xx or connect error) to mark an endpoint unhealthy via passive observation. None disables passive checking. |
clusters[].health_check.path | string | no | HTTP path to probe (only used for http type). |
clusters[].health_check.timeout_ms | integer | no | Probe timeout in milliseconds. Must be less than interval_ms. |
clusters[].health_check.unhealthy_threshold | integer | no | Consecutive failures required to mark an endpoint unhealthy. |
clusters[].idle_timeout_ms | integer | no | Idle connection timeout in milliseconds. Closes pooled upstream connections that have been idle longer than this duration. None uses Pingora’s default. |
clusters[].load_balancer_strategy | round_robin | least_connections | p2c | random | consistent_hash | maglev | ring_hash | subset | zone_aware | priority | no | Load-balancing algorithm for this cluster. Defaults to round_robin. |
clusters[].max_connections | integer | no | Maximum concurrent in-flight requests to this cluster. When set, excess requests receive 503. Prevents a single slow upstream from consuming all available capacity. |
clusters[].read_timeout_ms | integer | no | Per-read timeout in milliseconds. Applies to each individual read operation on an established upstream connection. A timeout fires a 502 response to the client. Use [total_connection_timeout_ms] to bound the entire exchange instead. |
clusters[].tls | ClusterTls | no | TLS settings for upstream connections. Presence implies TLS is enabled. Omit for plaintext HTTP. |
clusters[].tls.ca | CaConfig | no | Custom CA. |
clusters[].tls.ca.ca_path | string | yes | Path to the PEM CA certificate file. |
clusters[].tls.ca.crl_paths | string[] | no | Paths to PEM-encoded certificate revocation list (CRL) files. Applies only to listener client authentication: the client verifier checks presented client certificates against these CRLs and rejects revoked ones. Upstream (cluster) CRL checking is not implemented, so crl_paths under clusters[].tls.ca is rejected at config validation rather than being silently ignored. |
clusters[].tls.client_cert | CertKeyPair | no | Client certificate for upstream mTLS. |
clusters[].tls.client_cert.cert_path | string | yes | Path to the PEM certificate file. |
clusters[].tls.client_cert.default | bool | no | Whether this certificate is the default fallback for unmatched SNI. At most one certificate in a multi-cert config may set this to true. The default entry does not need server_names. |
clusters[].tls.client_cert.key_path | string | yes | Path to the PEM private key file. |
clusters[].tls.client_cert.server_names | string[] | no | SNI hostnames this certificate serves (listener only). |
clusters[].tls.sni | string | no | SNI hostname. |
clusters[].tls.verify | bool | no | Verify upstream certificate. |
clusters[].total_connection_timeout_ms | integer | no | Total connection timeout in milliseconds (TCP + TLS). Bounds the combined TCP handshake and TLS negotiation. When exceeded, the connection attempt fails with a 502 response. Prefer this over [connection_timeout_ms] for TLS-enabled clusters where the handshake dominates latency. |
clusters[].write_timeout_ms | integer | no | Per-write timeout in milliseconds. Applies to each individual write operation on an established upstream connection. A timeout fires a 502 response to the client. |
clusters[].retry_policy | RetryPolicy | no | Optional retry policy for this cluster. When unset, the proxy retains the legacy connect-failure retry behavior (3 attempts, idempotent methods, 64 KiB body). |
clusters[].retry_policy.max_retries | integer | no | Maximum number of retry attempts after the initial try. None means inherit from the merge parent / use the legacy default. |
clusters[].retry_policy.retriable_status_codes | HttpStatusCode[] | no | Explicit HTTP status codes that are retriable. |
clusters[].retry_policy.retriable_conditions | (connect_failure | reset | refused_stream | status5xx)[] | no | Named retriable conditions (connect failure, reset, 5xx, …). |
clusters[].retry_policy.per_try_timeout_ms | integer | no | Independent timeout for a single upstream attempt, in milliseconds. |
clusters[].retry_policy.request_timeout_ms | integer | no | Overall request deadline across all attempts, in milliseconds. When unset, the overall-timeout guard is skipped. |
clusters[].retry_policy.backoff | BackoffConfig | no | Exponential backoff between attempts. |
clusters[].retry_policy.backoff.base_interval_ms | integer | yes | Exponential base interval in milliseconds. |
clusters[].retry_policy.backoff.max_interval_ms | integer | yes | Maximum capped interval in milliseconds. |
clusters[].retry_policy.configured | bool | no | Whether this policy came from operator configuration rather than the built-in legacy default. Endpoint reselection on retry is enabled only for configured policies; the legacy default preserves the historical retry-same-endpoint semantics. |
clusters[].retry_policy.retry_budget | RetryBudgetConfig | no | Token-bucket retry budget. |
clusters[].retry_policy.retry_budget.percent | number | yes | Maximum retries as a percentage of active requests (0.0..=100.0). |
clusters[].retry_policy.retry_budget.min_retries_per_second | integer | no | Floor on tokens per second even at low traffic. |
clusters[].retry_policy.retry_body_limit_bytes | integer | no | Max request body size eligible for replay (bytes). Defaults to 64 KiB. |
clusters[].retry_policy.allow_non_idempotent | bool | no | Allow retries for non-idempotent methods (POST/PATCH) when true. |
cluster_source | router | bound_upstream | no | Where the target cluster name comes from. router (the default) uses the cluster a preceding router selected. bound_upstream (needs the upstream-binding build feature) uses the request’s logical binding, which lets a direct dispatch branch or an iterative_request_router step pick an endpoint with no router of its own; the bound cluster must be one of the clusters declared here, which startup validation checks. When an upstream was already selected, the filter does nothing. |
Example
filter: load_balancer
# cluster_source: router # the default; or bound_upstream (see below)
clusters:
- name: backend
endpoints: ["10.0.0.1:80"]