Token Rate Limit Header Keys
Extends token-rate-limit.yaml with M5 header dimensions (ai#123 / ai#129): each distinct x-tenant-id value gets its own token budget under the same catch-all rule
Versions marked “overview” do not contain this page. Selecting one opens that version’s documentation overview.
Category: Setup-dependent integration
Task: Extends token-rate-limit.yaml with M5 header dimensions (ai#123 / ai#129): each distinct x-tenant-id value gets its own token budget under the same catch-all rule
Prerequisites: The external service, credentials, or certificates referenced by this configuration.
Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
Companion resources from the same snapshot:
# Token Rate Limiting -- Per-Header Bucket Keys
#
# Extends token-rate-limit.yaml with M5 header dimensions (ai#123 /
# ai#129): each distinct `x-tenant-id` value gets its own token budget
# under the same catch-all rule. Missing `x-tenant-id` is rejected (400)
# unless you set `missing: fallback`.
#
listeners:
- name: default
address: "127.0.0.1:8080"
filter_chains:
- main
filter_chains:
- name: main
filters:
- filter: router
routes:
- path: "/v1/chat/completions"
cluster: backend
- filter: token_rate_limit
key:
- header: x-tenant-id
default_weights:
input: 1.0
output: 1.0
cached_input: 0.1
cache_write: 1.25
reasoning: 0.9
rules:
- name: default
algorithm: sliding_window
window: 1h
capacity: 100000
reserved_tokens: 500
- filter: token_count
provider: openai
- filter: access_log
- filter: load_balancer
clusters:
- name: backend
endpoints:
- "127.0.0.1:3000"
insecure_options:
allow_private_endpoints: true