Token Rate Limit
Apply token-aware rate limiting to inference requests and reconcile reported usage.
Category: Setup-dependent integration
Task: Limit inference traffic by estimated token cost
Prerequisites: Praxis AI runtime; A compatible inference backend; Provider credentials when required
Expected outcome: Requests reserve token budget before forwarding and reconcile it with reported usage.
Run it: Use ghcr.io/praxis-proxy/ai:0.4.1 and follow the container quickstart to mount and start the configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
Companion resources from the same snapshot:
# Token Rate Limiting
#
# Reserves an estimated token cost at admission time and reconciles
# that reservation against actual provider-reported usage once the
# response completes. Rejects with 429 when the bucket can't cover
# the estimate.
#
listeners:
- name: default
address: "127.0.0.1:8080"
filter_chains:
- main
filter_chains:
- name: main
filters:
- filter: router
routes:
- path: "/v1/chat/completions"
cluster: backend
- filter: token_rate_limit
key: global
default_weights:
input: 1.0
output: 1.0
cached_input: 0.1 # prompt-cache hits (token.cache_read)
cache_write: 1.25 # prompt-cache writes (token.cache_write)
reasoning: 0.9 # thinking tokens (token.reasoning)
rules:
- name: default
algorithm: sliding_window
window: 1h # sliding window duration
capacity: 100000 # max tokens admitted within `window`
reserved_tokens: 500 # fixed cost reserved per request at admission
- filter: token_count
provider: openai # openai | anthropic | google | bedrock | azure
- filter: access_log
- filter: load_balancer
clusters:
- name: backend
endpoints:
- "127.0.0.1:3000"
insecure_options:
allow_private_endpoints: true # example proxies to a local backend