token_rate_limit

Token-denominated rate limiter: reserves an estimated cost at admission, reconciles against actual usage after the response completes.
On this page

Token-denominated rate limiter: reserves an estimated cost at admission, reconciles against actual usage after the response completes. Evaluates an ordered list of rules, each with its own optional match condition, algorithm choice, and budget.

Configuration Notes

Experimental: requires the token-rate-limit-filter cargo feature, which is off by default and activates the experimental marker. This filter delivers the agreed M1/M2/M6/M7 milestone scope, but its parent proposal is not yet accepted and open questions remain (HA/clustered-Valkey failure modes, and the relationship to Kuadrant’s TokenRateLimitPolicy – see ai#127). The configuration surface may change between releases.

Mirrors the rules:/match: shape from the 00121_token-rate-limiting proposal in praxis-proxy/enhancements, scoped to this milestone’s static header-value matchers, per-rule algorithm choice, configurable estimation strategies (M3, see [EstimationConfig]), and M4 token-type weights (default_weights / per-rule weights). CEL matchers and soft-limit tiers are still out of scope (see the module doc comment) – upstream itself defers those.

Assumes request identity has already been resolved upstream (this filter doesn’t authenticate callers) – a catch-all rule (no match:) reserves quota for every request that reaches it, including probes and health checks. Scope rules with explicit match: conditions, or place an identity/auth filter earlier in the pipeline. Tracked as follow-on integration work in grid#101.

Observability is group-level by rule, never by user. Metrics carry only bounded rule, algorithm, backend, result, and capacity labels; accounting logs and optional OpenTelemetry spans likewise omit raw subject and bucket-key values. The Prometheus contract is:

  • praxis_trl_requests_total{rule,result} (admitted or denied): budget decisions only. Requests rejected before a decision are counted by praxis_trl_unauthenticated_total (401, no trusted subject) and praxis_trl_backend_errors_total (503, fail closed) instead.

  • praxis_trl_unauthenticated_total{rule}

  • praxis_trl_tokens_reserved_total{rule}

  • praxis_trl_tokens_reconciled_total{rule}

  • praxis_trl_tokens_refunded_total{rule}

  • praxis_trl_tokens_overage_total{rule}

  • praxis_trl_reservations_total{rule,result} (reconciled or orphaned)

  • praxis_trl_soft_tier_activations_total{rule,capacity}

  • praxis_trl_backend_errors_total{rule,backend}: failed reservations (the 503 path) and reconciliations abandoned after their retries.

  • praxis_trl_backend_reconciliation_total{rule,backend,result}: reconciliations completed by a Valkey worker.

  • praxis_trl_budget_remaining{rule,algorithm}

  • praxis_trl_reservations_active{rule}

  • praxis_trl_active_keys{rule}

Every previous praxis_ai_token_rate_limit_* name has moved to this prefix; no compatibility aliases are emitted.

budget_remaining is the remaining budget for the key of the most recent admission decision on this replica (admitted or denied), and active_keys is how many keys the backend currently retains. Both are snapshots taken as decisions happen, not continuously refreshed values. Like all Prometheus gauges they are f64 and saturate at the largest exactly representable integer (2^53 - 1).

Gauge scope depends on the backend. With the memory backend every gauge describes this process only, so aggregate replicas with sum. With the valkey backend, reservations_active is scoped to the namespace and algorithm, so summing it over rules double-counts; aggregate with max across replicas and rules. active_keys is scoped per rule, so aggregate with max across replicas for each rule. Each replica exports the value it last observed from the shared store. budget_remaining stays per replica and per last decision on either backend: it describes whichever key that replica decided last, so max or sum across replicas says little beyond “some key had this much left”. A replica that stops seeing traffic for a rule keeps exporting its last observation until it does.

The valkey backend requires Valkey or Redis 7.0+ (PEXPIRE NX/GT is used). The valkey backend keeps sliding-window usage in 60 fixed sub-windows per window (one per second for windows under a minute); usage leaves the window up to one sub-window late, never early. Changing a window’s length changes its sub-window width and so starts that window’s usage from zero. On the sliding window, concurrent admissions on one key are not serialised, so they can overshoot the budget by their combined estimates for one round trip. Usage written is never lost.

The valkey token bucket, by contrast, serialises admissions per key through an optimistic transaction: one key admits at most about one request per two Valkey round trips across the whole fleet, and contention shows up first as added latency, up to the 500 ms Valkey timeout, then as 503s. Use a non-global key for high-throughput token-bucket rules so the load spreads over many buckets.

During a rolling upgrade from the earlier scripted valkey backend, replicas on the old and new versions keep separate state, so for one window (and until old token buckets have drained) combined admissions can reach about twice the budget. All valkey timestamps come from the proxy replicas’ clocks, not Valkey’s: skew between replicas can under-count usage at window edges by up to the skew, and a replica whose clock runs fast trims other replicas’ live reservations and keys from the caps early.

Admissions, denials, reconciliations, and backend failures also emit structured records on the praxis_ai::token_rate_limit::accounting tracing target: INFO for admissions and settlements, WARN for failures. They are on by default at INFO, so every admitted or denied request produces one line in the operational log stream; keep only failures with runtime.log_overrides: {"praxis_ai::token_rate_limit::accounting": "warn"}, and separate them from other operational logs by filtering on the target field. The records contain bounded policy and token-count fields only. They are best-effort operational audit records, not a durable billing source.

Configuration

FieldTypeRequiredDescription
rulesRuleConfig[]yesEvaluated in order; the first rule whose match is satisfied (or which has no match at all) applies to a given request. A request satisfying no rule’s match is not rate limited by this filter instance – add a trailing rule with no match to enforce a catch-all budget instead.
rules[].namestringyesHuman-readable rule identifier, folded into Valkey key namespacing so distinct rules sharing one backend never collide. Renaming a live valkey-backed rule is therefore not a no-op for operators: it changes the Valkey key hash, so the old name’s tracked budget is orphaned (left to expire on its own TTL) and the new name starts with a fresh budget. There’s no migration/rename path today – routine config hygiene (e.g. renaming "gold" to "gold-tier") silently resets that rule’s state.
rules[].matchMatchConfignoStatic header-value match condition. Every listed header must be present on the request with an exact value match (ANDed) for this rule to apply. Omit entirely for a catch-all rule.
rules[].match.headersobject<string, string>yesEvery header must be present on the request with this exact value for the rule to match (ANDed across all entries).
rules[].algorithmsliding_window | token_bucketyesWhich admission algorithm this rule enforces, and that algorithm’s own parameters.
rules[].windowstringone ofSliding window duration (e.g. "1h", "60s").
rules[].capacityintegerone ofMaximum tokens admitted within window.
rules[].capacityintegerone ofMaximum tokens held at once (the bucket’s ceiling).
rules[].refill_ratenumberone ofTokens refilled per second, up to capacity.
rules[].reserved_tokensintegernoFixed token cost reserved at admission time, before actual usage is known. Legacy field, retained for backward compatibility: a bare reserved_tokens: N is equivalent to estimation: { strategy: fixed, fallback_estimate: N }. Mutually exclusive with estimation – specifying both on the same rule is a config error.
rules[].estimationEstimationConfignoConfigurable estimation strategy for computing the token cost reserved at admission time. Replaces the legacy reserved_tokens field with request-metadata-aware strategies. Mutually exclusive with reserved_tokens – specifying both on the same rule is a config error. Omitting both is also an error.
rules[].estimation.strategyfixed | max_tokens | input_plus_max_tokens | model_scaledyesWhich strategy to use for this rule’s cost estimation.
rules[].estimation.multipliernumbernoSafety-margin multiplier applied to the computed estimate. Defaults to 1.0 (no margin). Must be positive and finite.
rules[].estimation.fallback_estimateintegernoToken count to use when max_tokens is absent from the request. Required for fixed; optional for body-dependent strategies (if unset and the strategy can’t extract a value, the request is admitted without a reservation).
rules[].estimation.model_multipliersobject<string, number>noPer-model multiplier map for model_scaled strategy.
rules[].estimation.default_multipliernumbernoDefault multiplier for models not listed in model_multipliers.
rules[].estimation.bytes_per_tokennumbernoApproximate bytes-per-token ratio for input_plus_max_tokens. Defaults to 4.0.
rules[].reservation_timeoutstringnoHow long an admitted-but-never-reconciled reservation (lost request: timeout, connection reset, upstream crash) is tracked as active before that already-reserved-at-admission charge against its estimate becomes irreversibly locked in (sliding-window: folded into the settled total so it survives the window’s normal aging-out; token-bucket: the tokens were already decremented at reserve time regardless, this only bounds how long the reservation is tracked as pending). This does not defer when the charge first applies – it applies immediately at admission, same as any other reservation. Answers the proposal’s still-open “lost request handling” question for this milestone. Defaults to [DEFAULT_RESERVATION_TIMEOUT] when unset.
rules[].weightsTokenTypeWeightsConfignoOptional per-rule overlay on [TokenRateLimitConfig::default_weights]. Omitted types inherit the filter defaults (then 1.0).
rules[].weights.inputnumbernoWeight for uncached input tokens (the residual of token.input after subtracting cache read/write). Defaults to 1.0 when omitted.
rules[].weights.outputnumbernoWeight for visible output tokens (the residual of token.output after subtracting nested reasoning). Defaults to 1.0 when omitted.
rules[].weights.cached_inputnumbernoWeight for prompt-cache hits (token.cache_read). A value below 1.0 cheapens cached input; the proposal’s example is 0.1.
rules[].weights.cache_writenumbernoWeight for prompt-cache writes (token.cache_write). Anthropic cache creation is typically priced above uncached input; omit to keep 1.0.
rules[].weights.reasoningnumbernoWeight for reasoning / thinking tokens (token.reasoning).
rules[].tiersTierConfig[]noGraduated enforcement tiers (proposal S1). Each tier defines a usage threshold and an action (inject or deny). When the backend admits a request, every tier whose capacity is at or below the current usage level fires: - inject: the request continues and the tier’s headers are set on the upstream request. - deny: hard-reject with 429 (same as M6; must be the last tier). Tiers must have strictly ascending capacity values. At most one deny tier is allowed, and it must be the last. Its capacity must equal the algorithm’s own capacity. When omitted, the rule behaves as before: a single hard deny at the algorithm’s capacity.
rules[].tiers[].capacityintegeryesUsage threshold at which this tier activates.
rules[].tiers[].actionActionConfigyesWhat happens when usage crosses this tier’s threshold.
rules[].tiers[].action.typeinject | denyyesWhether to continue with injected headers or hard-reject.
rules[].tiers[].action.headersobject<string, string>noHeaders to inject on the upstream request (required for inject, ignored for deny).
keyKeySpecnoHow this filter partitions each matched rule’s token budget. Accepts a scalar (global, authenticated_subject, ip, model), a list of dimensions (composite keys), a single dimension mapping (header: x-tenant-id), or a full spec with dimensions and missing. Defaults to one shared global bucket. Composite dimension order does not matter: compiled dimensions are sorted into a canonical order so reordering a list cannot silently reset live budgets. ip and model are as caller-controlled as header when they come from a forwarding header or a client-supplied model string. Pair them with max_keys so one client cannot fill the table.
max_keysintegernoSoft cap on distinct budget keys retained at once, per rule. Bounds cardinality from per-header, per-IP, and composite keying. Defaults to [super::MAX_KEYS]. A new distinct key past this cap is denied (429, accounting outcome key_capacity) rather than growing without bound. In-process ledgers enforce the cap per rule. Valkey enforces it against the per-rule retained-key set ({namespace}:v2:keys:{rule_hash}, or the token-bucket equivalent). Idle in-process keys are reaped by ledger cleanup, which walks a bounded number of entries per request (including busy ones) so a single in-window key cannot pin the table at this cap.
backendBackendConfignoWhere every rule’s admission state lives: in-process (default, one budget per gateway instance) or a shared Valkey backend (one budget shared across every gateway instance/replica). One backend for the whole filter, not per rule – rules already share Valkey key-space isolation via namespace/rule-name hashing, so per-rule backend selection bought no isolation benefit, only a separate Valkey connection per rule pointed at the same URL. Revisit if a real deployment ever needs to mix in-process and Valkey rules in one filter instance.
backend.kindmemory | valkeynoWhich backend implementation to use.
backend.urlstringnoBackend connection URL. Supports one ${ENV_VAR} reference, so credentials/hostnames don’t need to be committed to config. Required when kind: valkey, ignored otherwise.
backend.namespacestringnoKey namespace prefix, so multiple filter rules or deployments can share one Valkey instance without colliding. Ignored for kind: memory. Defaults to "praxis:token_rate_limit" when unset.
default_weightsTokenTypeWeightsConfignoFilter-wide default per-type weights applied at reconciliation (proposal M4). Omitted types default to 1.0. Rules may overlay individual types via [RuleConfig::weights]. Admission still reserves the estimation/reserved_tokens cost unweighted.
default_weights.inputnumbernoWeight for uncached input tokens (the residual of token.input after subtracting cache read/write). Defaults to 1.0 when omitted.
default_weights.outputnumbernoWeight for visible output tokens (the residual of token.output after subtracting nested reasoning). Defaults to 1.0 when omitted.
default_weights.cached_inputnumbernoWeight for prompt-cache hits (token.cache_read). A value below 1.0 cheapens cached input; the proposal’s example is 0.1.
default_weights.cache_writenumbernoWeight for prompt-cache writes (token.cache_write). Anthropic cache creation is typically priced above uncached input; omit to keep 1.0.
default_weights.reasoningnumbernoWeight for reasoning / thinking tokens (token.reasoning).

Example

filter: token_rate_limit
key:                               # optional: defaults to one shared bucket per rule
  - authenticated_subject          # global | authenticated_subject | ip | model | header: NAME
  - model                          # header first (x-model); body only if already buffered
# key:                             # IP via a forwarding header (right-most hop after trusted_hops)
#   ip:
#     header: x-forwarded-for
#     trusted_hops: 1
#     ipv6_prefix: 64
# key:                             # named header, fail-open when absent
#   header: x-tenant-id
#   missing: fallback
max_keys: 100000                   # optional: per-rule cap on distinct budget keys
backend:                           # optional: defaults to in-process state, shared by every rule
  kind: valkey                      # memory (default) | valkey
  url: "${TOKEN_RATE_LIMIT_VALKEY_URL}"
  namespace: praxis:token_rate_limit
default_weights:                   # optional: omitted types default to 1.0
  input: 1.0
  output: 1.0
  cached_input: 0.1                # prompt-cache hits (token.cache_read)
  cache_write: 1.25                # prompt-cache writes (token.cache_write)
  reasoning: 0.9                   # thinking tokens (token.reasoning)
rules:
  - name: team-alpha                 # human-readable, unique per filter instance
    match:                           # optional: omit for a catch-all rule
      headers:
        x-app-id: alpha
    algorithm: sliding_window        # sliding_window | token_bucket
    window: 1h                       # sliding_window only: window duration
    capacity: 100000                 # max tokens admitted (sliding_window) or held (token_bucket)
    estimation:                      # M3: configurable estimation strategy
      strategy: max_tokens           # fixed | max_tokens | input_plus_max_tokens | model_scaled
      multiplier: 1.2                # optional safety margin (default: 1.0)
      fallback_estimate: 500         # used when max_tokens absent from request
    weights:                         # optional per-rule overlay on default_weights
      cached_input: 0.05
  - name: team-beta
    match:
      headers:
        x-app-id: beta
    algorithm: token_bucket
    capacity: 50000
    refill_rate: 50                  # token_bucket only: tokens refilled per second
    reserved_tokens: 200             # legacy: equivalent to estimation: { strategy: fixed, fallback_estimate: 200 }