Token Rate Limit

Apply token-aware rate limiting to inference requests and reconcile reported usage.

Category: Setup-dependent integration
Task: Limit inference traffic by estimated token cost

Prerequisites: Praxis AI runtime; A compatible inference backend; Provider credentials when required

Expected outcome: Requests reserve token budget before forwarding and reconcile it with reported usage.

Run it: Use ghcr.io/praxis-proxy/ai:0.4.1 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

Companion resources from the same snapshot:

# Token Rate Limiting
#
# Reserves an estimated token cost at admission time and reconciles
# that reservation against actual provider-reported usage once the
# response completes. Rejects with 429 when the bucket can't cover
# the estimate.
#
listeners:
  - name: default
    address: "127.0.0.1:8080"
    filter_chains:
      - main

filter_chains:
  - name: main
    filters:
      - filter: router
        routes:
          - path: "/v1/chat/completions"
            cluster: backend

      - filter: token_rate_limit
        key: global
        default_weights:
          input: 1.0
          output: 1.0
          cached_input: 0.1    # prompt-cache hits (token.cache_read)
          cache_write: 1.25    # prompt-cache writes (token.cache_write)
          reasoning: 0.9       # thinking tokens (token.reasoning)
        rules:
          - name: default
            algorithm: sliding_window
            window: 1h             # sliding window duration
            capacity: 100000       # max tokens admitted within `window`
            reserved_tokens: 500   # fixed cost reserved per request at admission

      - filter: token_count
        provider: openai   # openai | anthropic | google | bedrock | azure

      - filter: access_log

      - filter: load_balancer
        clusters:
          - name: backend
            endpoints:
              - "127.0.0.1:3000"

insecure_options:
  allow_private_endpoints: true # example proxies to a local backend