Token Rate Limit Mixed Algorithms

Extends token-rate-limit.yaml with per-rule algorithm choice (ai#789 / praxis#551): each rule in rules: independently picks sliding_window or token_bucket, matched by a static header value. team-alpha gets an exact trailing-window budget; team-beta gets a continuously-refilling bucket

Category: Setup-dependent integration
Task: Extends token-rate-limit.yaml with per-rule algorithm choice (ai#789 / praxis#551): each rule in rules: independently picks sliding_window or token_bucket, matched by a static header value. team-alpha gets an exact trailing-window budget; team-beta gets a continuously-refilling bucket

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

Companion resources from the same snapshot:

# Token Rate Limiting -- Mixed Algorithms Per Rule
#
# Extends token-rate-limit.yaml with per-rule algorithm choice (ai#789 /
# praxis#551): each rule in `rules:` independently picks sliding_window
# or token_bucket, matched by a static header value. team-alpha gets an
# exact trailing-window budget; team-beta gets a continuously-refilling
# bucket. Requests matching neither rule are not rate limited by this
# filter instance.
#
listeners:
  - name: default
    address: "127.0.0.1:8080"
    filter_chains:
      - main

filter_chains:
  - name: main
    filters:
      - filter: router
        routes:
          - path: "/v1/chat/completions"
            cluster: backend

      - filter: token_rate_limit
        default_weights:
          cached_input: 0.1
          cache_write: 1.25
          reasoning: 0.9
        rules:
          - name: team-alpha
            match:
              headers:
                x-app-id: alpha
            algorithm: sliding_window
            window: 1h                    # exact trailing-window budget
            capacity: 100000              # max tokens admitted within `window`
            reserved_tokens: 500          # fixed cost reserved per request at admission
            weights:
              cached_input: 0.05          # this tenant gets a steeper cache discount
          - name: team-beta
            match:
              headers:
                x-app-id: beta
            algorithm: token_bucket
            capacity: 50000               # max tokens held at once
            refill_rate: 50               # tokens refilled per second, up to `capacity`
            reserved_tokens: 200

      - filter: token_count
        provider: openai   # openai | anthropic | google | bedrock | azure

      - filter: access_log

      - filter: load_balancer
        clusters:
          - name: backend
            endpoints:
              - "127.0.0.1:3000"

insecure_options:
  allow_private_endpoints: true # example proxies to a local backend