Token Counting

Extracts token usage from AI inference responses (streaming and non-streaming) and makes counts available to downstream filters via filter metadata as token.input, token.output, and token.total

Category: Setup-dependent integration
Task: Extracts token usage from AI inference responses (streaming and non-streaming) and makes counts available to downstream filters via filter metadata as token.input, token.output, and token.total

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

Run it: Use ghcr.io/praxis-proxy/ai:0.4.1 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# Token Counting
#
# Extracts token usage from AI inference responses (streaming and
# non-streaming) and makes counts available to downstream filters
# via filter metadata as token.input, token.output, and token.total.
#
# For providers that support prompt caching, the cached portion of the
# input is also reported as token.cache_read and token.cache_write.
# Both are a breakdown of token.input, not an addition to it, so
# summing them with token.input would double-count.
#
# Each cache key is set only when the provider reported that count. An
# absent key means the response carried no cache information; a zero
# means the provider reported a cache miss. Providers whose API has no
# cache write concept, such as Google, never set token.cache_write;
# OpenAI reports cache writes as cache_write_tokens on both Chat
# Completions and Responses API usage.
#
# Reasoning / thinking tokens are reported as token.reasoning when the
# provider sends them:
#   - OpenAI / Azure Chat Completions:
#     usage.completion_tokens_details.reasoning_tokens
#     (a breakdown of token.output, not an addition)
#   - OpenAI / Azure Responses API:
#     usage.output_tokens_details.reasoning_tokens
#     (a breakdown of token.output, not an addition)
#   - Anthropic: usage.output_tokens_details.thinking_tokens
#     (a breakdown of token.output; on streams, the final message_delta)
#   - Google: usageMetadata.thoughtsTokenCount
#     (separate from candidatesTokenCount / token.output; already in
#     the provider total when totalTokenCount is present, otherwise
#     included in the computed fallback total)
#   - Bedrock Converse: not reported (no documented field)
# An absent key means the provider did not report reasoning; a zero
# means it reported none.
#
listeners:
  - name: default
    address: "127.0.0.1:8080"
    filter_chains:
      - main

filter_chains:
  - name: main
    filters:
      - filter: router
        routes:
          - path_prefix: "/"
            cluster: backend

      # token_usage_headers is declared before token_count: response
      # hooks run in reverse declared order, so token_count's
      # on_response_body (which sets filter_metadata) runs before
      # token_usage_headers reads it.
      - filter: token_usage_headers

      - filter: token_count
        provider: openai   # openai | anthropic | google | bedrock | azure

      - filter: access_log

      - filter: load_balancer
        clusters:
          - name: backend
            endpoints:
              - "127.0.0.1:3000"

insecure_options:
  allow_private_endpoints: true # example proxies to local backends