token_rate_limit
On this page
Token-denominated rate limiter: reserves an estimated cost at admission, reconciles against actual usage after the response completes. Evaluates an ordered list of rules, each with its own optional match condition, algorithm choice, and budget.
Configuration Notes
Experimental: requires the token-rate-limit-filter cargo feature, which is off by default and activates the experimental marker. This filter delivers the agreed M1/M2/M6 milestone scope, but its parent proposal is not yet accepted and open questions remain (HA/clustered-Valkey failure modes, and the relationship to Kuadrant’s TokenRateLimitPolicy – see ai#127). The configuration surface may change between releases.
Mirrors the rules:/match: shape from the 00121_token-rate-limiting proposal in praxis-proxy/enhancements, scoped to this milestone’s static header-value matchers, per-rule algorithm choice, configurable estimation strategies (M3, see [EstimationConfig]), and M4 token-type weights (default_weights / per-rule weights). CEL matchers and soft-limit tiers are still out of scope (see the module doc comment) – upstream itself defers those.
Assumes request identity has already been resolved upstream (this filter doesn’t authenticate callers) – a catch-all rule (no match:) reserves quota for every request that reaches it, including probes and health checks. Scope rules with explicit match: conditions, or place an identity/auth filter earlier in the pipeline. Tracked as follow-on integration work in grid#101.
Configuration
| Field | Type | Required | Description |
|---|---|---|---|
rules | RuleConfig[] | yes | Evaluated in order; the first rule whose match is satisfied (or which has no match at all) applies to a given request. A request satisfying no rule’s match is not rate limited by this filter instance – add a trailing rule with no match to enforce a catch-all budget instead. |
rules[].name | string | yes | Human-readable rule identifier, folded into Valkey key namespacing so distinct rules sharing one backend never collide. Renaming a live valkey-backed rule is therefore not a no-op for operators: it changes the Valkey key hash, so the old name’s tracked budget is orphaned (left to expire on its own TTL) and the new name starts with a fresh budget. There’s no migration/rename path today – routine config hygiene (e.g. renaming "gold" to "gold-tier") silently resets that rule’s state. |
rules[].match | MatchConfig | no | Static header-value match condition. Every listed header must be present on the request with an exact value match (ANDed) for this rule to apply. Omit entirely for a catch-all rule. |
rules[].match.headers | object<string, string> | yes | Every header must be present on the request with this exact value for the rule to match (ANDed across all entries). |
rules[].algorithm | sliding_window | token_bucket | yes | Which admission algorithm this rule enforces, and that algorithm’s own parameters. |
rules[].window | string | one of | Sliding window duration (e.g. "1h", "60s"). |
rules[].capacity | integer | one of | Maximum tokens admitted within window. |
rules[].capacity | integer | one of | Maximum tokens held at once (the bucket’s ceiling). |
rules[].refill_rate | number | one of | Tokens refilled per second, up to capacity. |
rules[].reserved_tokens | integer | no | Fixed token cost reserved at admission time, before actual usage is known. Legacy field, retained for backward compatibility: a bare reserved_tokens: N is equivalent to estimation: { strategy: fixed, fallback_estimate: N }. Mutually exclusive with estimation – specifying both on the same rule is a config error. |
rules[].estimation | EstimationConfig | no | Configurable estimation strategy for computing the token cost reserved at admission time. Replaces the legacy reserved_tokens field with request-metadata-aware strategies. Mutually exclusive with reserved_tokens – specifying both on the same rule is a config error. Omitting both is also an error. |
rules[].estimation.strategy | fixed | max_tokens | input_plus_max_tokens | model_scaled | yes | Which strategy to use for this rule’s cost estimation. |
rules[].estimation.multiplier | number | no | Safety-margin multiplier applied to the computed estimate. Defaults to 1.0 (no margin). Must be positive and finite. |
rules[].estimation.fallback_estimate | integer | no | Token count to use when max_tokens is absent from the request. Required for fixed; optional for body-dependent strategies (if unset and the strategy can’t extract a value, the request is admitted without a reservation). |
rules[].estimation.model_multipliers | object<string, number> | no | Per-model multiplier map for model_scaled strategy. |
rules[].estimation.default_multiplier | number | no | Default multiplier for models not listed in model_multipliers. |
rules[].estimation.bytes_per_token | number | no | Approximate bytes-per-token ratio for input_plus_max_tokens. Defaults to 4.0. |
rules[].reservation_timeout | string | no | How long an admitted-but-never-reconciled reservation (lost request: timeout, connection reset, upstream crash) is tracked as active before that already-reserved-at-admission charge against its estimate becomes irreversibly locked in (sliding-window: folded into the settled total so it survives the window’s normal aging-out; token-bucket: the tokens were already decremented at reserve time regardless, this only bounds how long the reservation is tracked as pending). This does not defer when the charge first applies – it applies immediately at admission, same as any other reservation. Answers the proposal’s still-open “lost request handling” question for this milestone. Defaults to [DEFAULT_RESERVATION_TIMEOUT] when unset. |
rules[].weights | TokenTypeWeightsConfig | no | Optional per-rule overlay on [TokenRateLimitConfig::default_weights]. Omitted types inherit the filter defaults (then 1.0). |
rules[].weights.input | number | no | Weight for uncached input tokens (the residual of token.input after subtracting cache read/write). Defaults to 1.0 when omitted. |
rules[].weights.output | number | no | Weight for visible output tokens (the residual of token.output after subtracting nested reasoning). Defaults to 1.0 when omitted. |
rules[].weights.cached_input | number | no | Weight for prompt-cache hits (token.cache_read). A value below 1.0 cheapens cached input; the proposal’s example is 0.1. |
rules[].weights.cache_write | number | no | Weight for prompt-cache writes (token.cache_write). Anthropic cache creation is typically priced above uncached input; omit to keep 1.0. |
rules[].weights.reasoning | number | no | Weight for reasoning / thinking tokens (token.reasoning). |
rules[].tiers | TierConfig[] | no | Graduated enforcement tiers (proposal S1). Each tier defines a usage threshold and an action (inject or deny). When the backend admits a request, every tier whose capacity is at or below the current usage level fires: - inject: the request continues and the tier’s headers are set on the upstream request. - deny: hard-reject with 429 (same as M6; must be the last tier). Tiers must have strictly ascending capacity values. At most one deny tier is allowed, and it must be the last. Its capacity must equal the algorithm’s own capacity. When omitted, the rule behaves as before: a single hard deny at the algorithm’s capacity. |
rules[].tiers[].capacity | integer | yes | Usage threshold at which this tier activates. |
rules[].tiers[].action | ActionConfig | yes | What happens when usage crosses this tier’s threshold. |
rules[].tiers[].action.type | inject | deny | yes | Whether to continue with injected headers or hard-reject. |
rules[].tiers[].action.headers | object<string, string> | no | Headers to inject on the upstream request (required for inject, ignored for deny). |
key | global | authenticated_subject | no | Trusted request identity used to partition each rule’s budget. The default preserves the historical single global bucket. |
backend | BackendConfig | no | Where every rule’s admission state lives: in-process (default, one budget per gateway instance) or a shared Valkey backend (one budget shared across every gateway instance/replica). One backend for the whole filter, not per rule – rules already share Valkey key-space isolation via namespace/rule-name hashing, so per-rule backend selection bought no isolation benefit, only a separate Valkey connection per rule pointed at the same URL. Revisit if a real deployment ever needs to mix in-process and Valkey rules in one filter instance. |
backend.kind | memory | valkey | no | Which backend implementation to use. |
backend.url | string | no | Backend connection URL. Supports one ${ENV_VAR} reference, so credentials/hostnames don’t need to be committed to config. Required when kind: valkey, ignored otherwise. |
backend.namespace | string | no | Key namespace prefix, so multiple filter rules or deployments can share one Valkey instance without colliding. Ignored for kind: memory. Defaults to "praxis:token_rate_limit" when unset. |
default_weights | TokenTypeWeightsConfig | no | Filter-wide default per-type weights applied at reconciliation (proposal M4). Omitted types default to 1.0. Rules may overlay individual types via [RuleConfig::weights]. Admission still reserves the estimation/reserved_tokens cost unweighted. |
default_weights.input | number | no | Weight for uncached input tokens (the residual of token.input after subtracting cache read/write). Defaults to 1.0 when omitted. |
default_weights.output | number | no | Weight for visible output tokens (the residual of token.output after subtracting nested reasoning). Defaults to 1.0 when omitted. |
default_weights.cached_input | number | no | Weight for prompt-cache hits (token.cache_read). A value below 1.0 cheapens cached input; the proposal’s example is 0.1. |
default_weights.cache_write | number | no | Weight for prompt-cache writes (token.cache_write). Anthropic cache creation is typically priced above uncached input; omit to keep 1.0. |
default_weights.reasoning | number | no | Weight for reasoning / thinking tokens (token.reasoning). |
Example
filter: token_rate_limit
backend: # optional: defaults to in-process state, shared by every rule
kind: valkey # memory (default) | valkey
url: "${TOKEN_RATE_LIMIT_VALKEY_URL}"
namespace: praxis:token_rate_limit
default_weights: # optional: omitted types default to 1.0
input: 1.0
output: 1.0
cached_input: 0.1 # prompt-cache hits (token.cache_read)
cache_write: 1.25 # prompt-cache writes (token.cache_write)
reasoning: 0.9 # thinking tokens (token.reasoning)
rules:
- name: team-alpha # human-readable, unique per filter instance
match: # optional: omit for a catch-all rule
headers:
x-app-id: alpha
algorithm: sliding_window # sliding_window | token_bucket
window: 1h # sliding_window only: window duration
capacity: 100000 # max tokens admitted (sliding_window) or held (token_bucket)
estimation: # M3: configurable estimation strategy
strategy: max_tokens # fixed | max_tokens | input_plus_max_tokens | model_scaled
multiplier: 1.2 # optional safety margin (default: 1.0)
fallback_estimate: 500 # used when max_tokens absent from request
weights: # optional per-rule overlay on default_weights
cached_input: 0.05
- name: team-beta
match:
headers:
x-app-id: beta
algorithm: token_bucket
capacity: 50000
refill_rate: 50 # token_bucket only: tokens refilled per second
reserved_tokens: 200 # legacy: equivalent to estimation: { strategy: fixed, fallback_estimate: 200 }