token_rate_limit

Token-denominated rate limiter: reserves an estimated cost at admission, reconciles against actual usage after the response completes.
On this page

Token-denominated rate limiter: reserves an estimated cost at admission, reconciles against actual usage after the response completes. Evaluates an ordered list of rules, each with its own optional match condition, algorithm choice, and budget.

Configuration Notes

Experimental: requires the token-rate-limit-filter cargo feature, which is off by default and activates the experimental marker. This filter delivers the agreed M1/M2/M6 milestone scope, but its parent proposal is not yet accepted and open questions remain (HA/clustered-Valkey failure modes, and the relationship to Kuadrant’s TokenRateLimitPolicy – see ai#127). The configuration surface may change between releases.

Mirrors the rules:/match: shape from the 00121_token-rate-limiting proposal in praxis-proxy/enhancements, scoped to this milestone’s static header-value matchers, per-rule algorithm choice, configurable estimation strategies (M3, see [EstimationConfig]), and M4 token-type weights (default_weights / per-rule weights). CEL matchers and soft-limit tiers are still out of scope (see the module doc comment) – upstream itself defers those.

Assumes request identity has already been resolved upstream (this filter doesn’t authenticate callers) – a catch-all rule (no match:) reserves quota for every request that reaches it, including probes and health checks. Scope rules with explicit match: conditions, or place an identity/auth filter earlier in the pipeline. Tracked as follow-on integration work in grid#101.

Configuration

FieldTypeRequiredDescription
rulesRuleConfig[]yesEvaluated in order; the first rule whose match is satisfied (or which has no match at all) applies to a given request. A request satisfying no rule’s match is not rate limited by this filter instance – add a trailing rule with no match to enforce a catch-all budget instead.
rules[].namestringyesHuman-readable rule identifier, folded into Valkey key namespacing so distinct rules sharing one backend never collide. Renaming a live valkey-backed rule is therefore not a no-op for operators: it changes the Valkey key hash, so the old name’s tracked budget is orphaned (left to expire on its own TTL) and the new name starts with a fresh budget. There’s no migration/rename path today – routine config hygiene (e.g. renaming "gold" to "gold-tier") silently resets that rule’s state.
rules[].matchMatchConfignoStatic header-value match condition. Every listed header must be present on the request with an exact value match (ANDed) for this rule to apply. Omit entirely for a catch-all rule.
rules[].match.headersobject<string, string>yesEvery header must be present on the request with this exact value for the rule to match (ANDed across all entries).
rules[].algorithmsliding_window | token_bucketyesWhich admission algorithm this rule enforces, and that algorithm’s own parameters.
rules[].windowstringone ofSliding window duration (e.g. "1h", "60s").
rules[].capacityintegerone ofMaximum tokens admitted within window.
rules[].capacityintegerone ofMaximum tokens held at once (the bucket’s ceiling).
rules[].refill_ratenumberone ofTokens refilled per second, up to capacity.
rules[].reserved_tokensintegernoFixed token cost reserved at admission time, before actual usage is known. Legacy field, retained for backward compatibility: a bare reserved_tokens: N is equivalent to estimation: { strategy: fixed, fallback_estimate: N }. Mutually exclusive with estimation – specifying both on the same rule is a config error.
rules[].estimationEstimationConfignoConfigurable estimation strategy for computing the token cost reserved at admission time. Replaces the legacy reserved_tokens field with request-metadata-aware strategies. Mutually exclusive with reserved_tokens – specifying both on the same rule is a config error. Omitting both is also an error.
rules[].estimation.strategyfixed | max_tokens | input_plus_max_tokens | model_scaledyesWhich strategy to use for this rule’s cost estimation.
rules[].estimation.multipliernumbernoSafety-margin multiplier applied to the computed estimate. Defaults to 1.0 (no margin). Must be positive and finite.
rules[].estimation.fallback_estimateintegernoToken count to use when max_tokens is absent from the request. Required for fixed; optional for body-dependent strategies (if unset and the strategy can’t extract a value, the request is admitted without a reservation).
rules[].estimation.model_multipliersobject<string, number>noPer-model multiplier map for model_scaled strategy.
rules[].estimation.default_multipliernumbernoDefault multiplier for models not listed in model_multipliers.
rules[].estimation.bytes_per_tokennumbernoApproximate bytes-per-token ratio for input_plus_max_tokens. Defaults to 4.0.
rules[].reservation_timeoutstringnoHow long an admitted-but-never-reconciled reservation (lost request: timeout, connection reset, upstream crash) is tracked as active before that already-reserved-at-admission charge against its estimate becomes irreversibly locked in (sliding-window: folded into the settled total so it survives the window’s normal aging-out; token-bucket: the tokens were already decremented at reserve time regardless, this only bounds how long the reservation is tracked as pending). This does not defer when the charge first applies – it applies immediately at admission, same as any other reservation. Answers the proposal’s still-open “lost request handling” question for this milestone. Defaults to [DEFAULT_RESERVATION_TIMEOUT] when unset.
rules[].weightsTokenTypeWeightsConfignoOptional per-rule overlay on [TokenRateLimitConfig::default_weights]. Omitted types inherit the filter defaults (then 1.0).
rules[].weights.inputnumbernoWeight for uncached input tokens (the residual of token.input after subtracting cache read/write). Defaults to 1.0 when omitted.
rules[].weights.outputnumbernoWeight for visible output tokens (the residual of token.output after subtracting nested reasoning). Defaults to 1.0 when omitted.
rules[].weights.cached_inputnumbernoWeight for prompt-cache hits (token.cache_read). A value below 1.0 cheapens cached input; the proposal’s example is 0.1.
rules[].weights.cache_writenumbernoWeight for prompt-cache writes (token.cache_write). Anthropic cache creation is typically priced above uncached input; omit to keep 1.0.
rules[].weights.reasoningnumbernoWeight for reasoning / thinking tokens (token.reasoning).
rules[].tiersTierConfig[]noGraduated enforcement tiers (proposal S1). Each tier defines a usage threshold and an action (inject or deny). When the backend admits a request, every tier whose capacity is at or below the current usage level fires: - inject: the request continues and the tier’s headers are set on the upstream request. - deny: hard-reject with 429 (same as M6; must be the last tier). Tiers must have strictly ascending capacity values. At most one deny tier is allowed, and it must be the last. Its capacity must equal the algorithm’s own capacity. When omitted, the rule behaves as before: a single hard deny at the algorithm’s capacity.
rules[].tiers[].capacityintegeryesUsage threshold at which this tier activates.
rules[].tiers[].actionActionConfigyesWhat happens when usage crosses this tier’s threshold.
rules[].tiers[].action.typeinject | denyyesWhether to continue with injected headers or hard-reject.
rules[].tiers[].action.headersobject<string, string>noHeaders to inject on the upstream request (required for inject, ignored for deny).
keyglobal | authenticated_subjectnoTrusted request identity used to partition each rule’s budget. The default preserves the historical single global bucket.
backendBackendConfignoWhere every rule’s admission state lives: in-process (default, one budget per gateway instance) or a shared Valkey backend (one budget shared across every gateway instance/replica). One backend for the whole filter, not per rule – rules already share Valkey key-space isolation via namespace/rule-name hashing, so per-rule backend selection bought no isolation benefit, only a separate Valkey connection per rule pointed at the same URL. Revisit if a real deployment ever needs to mix in-process and Valkey rules in one filter instance.
backend.kindmemory | valkeynoWhich backend implementation to use.
backend.urlstringnoBackend connection URL. Supports one ${ENV_VAR} reference, so credentials/hostnames don’t need to be committed to config. Required when kind: valkey, ignored otherwise.
backend.namespacestringnoKey namespace prefix, so multiple filter rules or deployments can share one Valkey instance without colliding. Ignored for kind: memory. Defaults to "praxis:token_rate_limit" when unset.
default_weightsTokenTypeWeightsConfignoFilter-wide default per-type weights applied at reconciliation (proposal M4). Omitted types default to 1.0. Rules may overlay individual types via [RuleConfig::weights]. Admission still reserves the estimation/reserved_tokens cost unweighted.
default_weights.inputnumbernoWeight for uncached input tokens (the residual of token.input after subtracting cache read/write). Defaults to 1.0 when omitted.
default_weights.outputnumbernoWeight for visible output tokens (the residual of token.output after subtracting nested reasoning). Defaults to 1.0 when omitted.
default_weights.cached_inputnumbernoWeight for prompt-cache hits (token.cache_read). A value below 1.0 cheapens cached input; the proposal’s example is 0.1.
default_weights.cache_writenumbernoWeight for prompt-cache writes (token.cache_write). Anthropic cache creation is typically priced above uncached input; omit to keep 1.0.
default_weights.reasoningnumbernoWeight for reasoning / thinking tokens (token.reasoning).

Example

filter: token_rate_limit
backend:                           # optional: defaults to in-process state, shared by every rule
  kind: valkey                      # memory (default) | valkey
  url: "${TOKEN_RATE_LIMIT_VALKEY_URL}"
  namespace: praxis:token_rate_limit
default_weights:                   # optional: omitted types default to 1.0
  input: 1.0
  output: 1.0
  cached_input: 0.1                # prompt-cache hits (token.cache_read)
  cache_write: 1.25                # prompt-cache writes (token.cache_write)
  reasoning: 0.9                   # thinking tokens (token.reasoning)
rules:
  - name: team-alpha                 # human-readable, unique per filter instance
    match:                           # optional: omit for a catch-all rule
      headers:
        x-app-id: alpha
    algorithm: sliding_window        # sliding_window | token_bucket
    window: 1h                       # sliding_window only: window duration
    capacity: 100000                 # max tokens admitted (sliding_window) or held (token_bucket)
    estimation:                      # M3: configurable estimation strategy
      strategy: max_tokens           # fixed | max_tokens | input_plus_max_tokens | model_scaled
      multiplier: 1.2                # optional safety margin (default: 1.0)
      fallback_estimate: 500         # used when max_tokens absent from request
    weights:                         # optional per-rule overlay on default_weights
      cached_input: 0.05
  - name: team-beta
    match:
      headers:
        x-app-id: beta
    algorithm: token_bucket
    capacity: 50000
    refill_rate: 50                  # token_bucket only: tokens refilled per second
    reserved_tokens: 200             # legacy: equivalent to estimation: { strategy: fixed, fallback_estimate: 200 }