Token Rate Limit Mixed Algorithms
Extends token-rate-limit.yaml with per-rule algorithm choice (ai#789 / praxis#551): each rule in rules: independently picks sliding_window or token_bucket, matched by a static header value. team-alpha gets an exact trailing-window budget; team-beta gets a continuously-refilling bucket
Category: Setup-dependent integration
Task: Extends token-rate-limit.yaml with per-rule algorithm choice (ai#789 / praxis#551): each rule in rules: independently picks sliding_window or token_bucket, matched by a static header value. team-alpha gets an exact trailing-window budget; team-beta gets a continuously-refilling bucket
Prerequisites: The external service, credentials, or certificates referenced by this configuration.
Run it: Use ghcr.io/praxis-proxy/ai:0.4.1 and follow the container quickstart to mount and start the configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
Companion resources from the same snapshot:
# Token Rate Limiting -- Mixed Algorithms Per Rule
#
# Extends token-rate-limit.yaml with per-rule algorithm choice (ai#789 /
# praxis#551): each rule in `rules:` independently picks sliding_window
# or token_bucket, matched by a static header value. team-alpha gets an
# exact trailing-window budget; team-beta gets a continuously-refilling
# bucket. Requests matching neither rule are not rate limited by this
# filter instance.
#
listeners:
- name: default
address: "127.0.0.1:8080"
filter_chains:
- main
filter_chains:
- name: main
filters:
- filter: router
routes:
- path: "/v1/chat/completions"
cluster: backend
- filter: token_rate_limit
default_weights:
cached_input: 0.1
cache_write: 1.25
reasoning: 0.9
rules:
- name: team-alpha
match:
headers:
x-app-id: alpha
algorithm: sliding_window
window: 1h # exact trailing-window budget
capacity: 100000 # max tokens admitted within `window`
reserved_tokens: 500 # fixed cost reserved per request at admission
weights:
cached_input: 0.05 # this tenant gets a steeper cache discount
- name: team-beta
match:
headers:
x-app-id: beta
algorithm: token_bucket
capacity: 50000 # max tokens held at once
refill_rate: 50 # tokens refilled per second, up to `capacity`
reserved_tokens: 200
- filter: token_count
provider: openai # openai | anthropic | google | bedrock | azure
- filter: access_log
- filter: load_balancer
clusters:
- name: backend
endpoints:
- "127.0.0.1:3000"
insecure_options:
allow_private_endpoints: true # example proxies to a local backend