load_balancer

Selects an upstream endpoint using the cluster’s configured strategy.
On this page

Selects an upstream endpoint using the cluster’s configured strategy.

Configuration Notes

Supported strategies: - round_robin (default): cycles through endpoints in order, respecting weights via endpoint expansion. - least_connections: picks the endpoint with the fewest active in-flight requests; decrements the counter on on_response. - p2c: samples two random endpoints and picks the less loaded one. - random: picks a uniformly random endpoint, weighted by endpoint weight. - consistent_hash: hashes a configurable request header (or the URI path when the header is absent) to pin requests to a stable endpoint. - maglev: hashes a configurable request header (or the URI path) through a Maglev lookup table for even distribution and minimal disruption when endpoints change.

Configuration

FieldTypeRequiredDescription
clustersCluster[]noCluster definitions.
clusters[].namestringyesUnique name for the cluster.
clusters[].httpClusterHttpOptionsnoHTTP-specific cluster options. Grouped under http: so the shared cluster/upstream transport types stay protocol-agnostic; TCP load balancers ignore this block entirely.
clusters[].http.authoritystringnoOverride the upstream HTTP Host header. When set, the proxy rewrites the Host header sent to the upstream instead of forwarding the downstream value. The downstream HTTP/2 :authority pseudo-header is never forwarded upstream; on an HTTP/2 upstream leg Pingora rebuilds :authority from this Host value. TLS SNI remains independent — configure tls.sni separately when needed. Must be a valid HTTP authority: a hostname with an optional port, or a bracketed IPv6 address with an optional port. URI schemes, paths, userinfo, and fragments are rejected.
clusters[].http.application_protocolstringnoOpaque application protocol the upstream cluster expects. Declares the wire representation of the request body (for example openai_chat_completions or openai_responses). The value stays opaque to Praxis core — consuming filters interpret it; Praxis defines no enum of known protocols. Validated as a bounded, canonical identifier: 1–64 bytes of lowercase ASCII letters, digits, ., _, or -, starting and ending with a letter or digit. The open string type is deliberate: the protocol set is open-ended and owned by consuming filters, so keep this a string — do not convert it to an enum.
clusters[].http.application_providerstringnoOpaque application provider refining application_protocol. Distinguishes provider-specific semantics (for example openai or vllm) independent of the deployed cluster, whose identity is already the cluster name. Stays opaque to Praxis core and follows the same identifier rules — and the same enum-free rationale — as application_protocol.
clusters[].http.versionh1 | h2 | autonoHTTP version used for upstream connections to this cluster. Defaults to [H1]. Set h2 for gRPC upstreams: response trailers, and therefore grpc-status, only exist on an HTTP/2 leg.
clusters[].connection_timeout_msintegernoTCP connection timeout in milliseconds. Applies to the TCP handshake only (before TLS). When exceeded, the connection attempt fails and the load balancer may retry on the next endpoint. None (the default) uses Pingora’s built-in timeout.
clusters[].endpoints(Simple | Weighted)[]yesList of endpoints for the cluster. Each entry is either a plain "host:port" string or a { address, weight } object.
clusters[].endpoints[].addressstringyesSocket address as host:port.
clusters[].endpoints[].weightintegernoRelative forwarding weight. Higher values receive proportionally more traffic. Defaults to 1.
clusters[].endpoints[].metadataobject<string, string>noArbitrary key-value metadata for subset-based load balancing.
clusters[].endpoints[].priorityintegernoPriority tier (0 = primary, 1 = first failover, etc.).
clusters[].endpoints[].zonestringnoLocality zone identifier for zone-aware routing.
clusters[].health_checkHealthCheckConfignoActive health check configuration for this cluster.
clusters[].health_check.typehttp | tcp | grpcyesProbe type: [Http], [Tcp], or [Grpc].
clusters[].health_check.expected_statusintegernoExpected HTTP status code for a healthy response.
clusters[].health_check.grpc_servicestringnogRPC service name to check. Only used by the grpc probe. Empty (the default) asks for the server’s overall serving status, which is what the grpc.health.v1.Health contract defines an empty name to mean.
clusters[].health_check.healthy_thresholdintegernoConsecutive successes required to mark an endpoint healthy.
clusters[].health_check.interval_msintegernoProbe interval in milliseconds.
clusters[].health_check.passive_healthy_thresholdintegernoConsecutive successes to mark an endpoint healthy again via passive observation. None disables passive recovery (active checks must recover it).
clusters[].health_check.passive_unhealthy_thresholdintegernoConsecutive response failures (5xx or connect error) to mark an endpoint unhealthy via passive observation. None disables passive checking.
clusters[].health_check.pathstringnoHTTP path to probe (only used for http type).
clusters[].health_check.timeout_msintegernoProbe timeout in milliseconds. Must be less than interval_ms.
clusters[].health_check.unhealthy_thresholdintegernoConsecutive failures required to mark an endpoint unhealthy.
clusters[].idle_timeout_msintegernoIdle connection timeout in milliseconds. Closes pooled upstream connections that have been idle longer than this duration. None uses Pingora’s default.
clusters[].load_balancer_strategyround_robin | least_connections | p2c | random | consistent_hash | maglev | ring_hash | subset | zone_aware | prioritynoLoad-balancing algorithm for this cluster. Defaults to round_robin.
clusters[].max_connectionsintegernoMaximum concurrent in-flight requests to this cluster. When set, excess requests receive 503. Prevents a single slow upstream from consuming all available capacity.
clusters[].read_timeout_msintegernoPer-read timeout in milliseconds. Applies to each individual read operation on an established upstream connection. A timeout fires a 502 response to the client. Use [total_connection_timeout_ms] to bound the entire exchange instead.
clusters[].tlsClusterTlsnoTLS settings for upstream connections. Presence implies TLS is enabled. Omit for plaintext HTTP.
clusters[].tls.caCaConfignoCustom CA.
clusters[].tls.ca.ca_pathstringyesPath to the PEM CA certificate file.
clusters[].tls.ca.crl_pathsstring[]noPaths to PEM-encoded certificate revocation list (CRL) files. Applies only to listener client authentication: the client verifier checks presented client certificates against these CRLs and rejects revoked ones. Upstream (cluster) CRL checking is not implemented, so crl_paths under clusters[].tls.ca is rejected at config validation rather than being silently ignored.
clusters[].tls.client_certCertKeyPairnoClient certificate for upstream mTLS.
clusters[].tls.client_cert.cert_pathstringyesPath to the PEM certificate file.
clusters[].tls.client_cert.defaultboolnoWhether this certificate is the default fallback for unmatched SNI. At most one certificate in a multi-cert config may set this to true. The default entry does not need server_names.
clusters[].tls.client_cert.key_pathstringyesPath to the PEM private key file.
clusters[].tls.client_cert.server_namesstring[]noSNI hostnames this certificate serves (listener only).
clusters[].tls.snistringnoSNI hostname.
clusters[].tls.verifyboolnoVerify upstream certificate.
clusters[].total_connection_timeout_msintegernoTotal connection timeout in milliseconds (TCP + TLS). Bounds the combined TCP handshake and TLS negotiation. When exceeded, the connection attempt fails with a 502 response. Prefer this over [connection_timeout_ms] for TLS-enabled clusters where the handshake dominates latency.
clusters[].write_timeout_msintegernoPer-write timeout in milliseconds. Applies to each individual write operation on an established upstream connection. A timeout fires a 502 response to the client.
clusters[].retry_policyRetryPolicynoOptional retry policy for this cluster. When unset, the proxy retains the legacy connect-failure retry behavior (3 attempts, idempotent methods, 64 KiB body).
clusters[].retry_policy.max_retriesintegernoMaximum number of retry attempts after the initial try. None means inherit from the merge parent / use the legacy default.
clusters[].retry_policy.retriable_status_codesHttpStatusCode[]noExplicit HTTP status codes that are retriable.
clusters[].retry_policy.retriable_conditions(connect_failure | reset | refused_stream | status5xx)[]noNamed retriable conditions (connect failure, reset, 5xx, …).
clusters[].retry_policy.per_try_timeout_msintegernoIndependent timeout for a single upstream attempt, in milliseconds.
clusters[].retry_policy.request_timeout_msintegernoOverall request deadline across all attempts, in milliseconds. When unset, the overall-timeout guard is skipped.
clusters[].retry_policy.backoffBackoffConfignoExponential backoff between attempts.
clusters[].retry_policy.backoff.base_interval_msintegeryesExponential base interval in milliseconds.
clusters[].retry_policy.backoff.max_interval_msintegeryesMaximum capped interval in milliseconds.
clusters[].retry_policy.configuredboolnoWhether this policy came from operator configuration rather than the built-in legacy default. Endpoint reselection on retry is enabled only for configured policies; the legacy default preserves the historical retry-same-endpoint semantics.
clusters[].retry_policy.retry_budgetRetryBudgetConfignoToken-bucket retry budget.
clusters[].retry_policy.retry_budget.percentnumberyesMaximum retries as a percentage of active requests (0.0..=100.0).
clusters[].retry_policy.retry_budget.min_retries_per_secondintegernoFloor on tokens per second even at low traffic.
clusters[].retry_policy.retry_body_limit_bytesintegernoMax request body size eligible for replay (bytes). Defaults to 64 KiB.
clusters[].retry_policy.allow_non_idempotentboolnoAllow retries for non-idempotent methods (POST/PATCH) when true.
cluster_sourcerouter | bound_upstreamnoWhere the target cluster name comes from. router (the default) uses the cluster a preceding router selected. bound_upstream (needs the upstream-binding build feature) uses the request’s logical binding, which lets a direct dispatch branch or an iterative_request_router step pick an endpoint with no router of its own; the bound cluster must be one of the clusters declared here, which startup validation checks. When an upstream was already selected, the filter does nothing.

Example

filter: load_balancer
# cluster_source: router      # the default; or bound_upstream (see below)
clusters:
  - name: backend
    endpoints: ["10.0.0.1:80"]