Token Counting
Extracts token usage from AI inference responses (streaming and non-streaming) and makes counts available to downstream filters via filter metadata as token.input, token.output, and token.total
Category: Setup-dependent integration
Task: Extracts token usage from AI inference responses (streaming and non-streaming) and makes counts available to downstream filters via filter metadata as token.input, token.output, and token.total
Prerequisites: The external service, credentials, or certificates referenced by this configuration.
Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
# Token Counting
#
# Extracts token usage from AI inference responses (streaming and
# non-streaming) and makes counts available to downstream filters
# via filter metadata as token.input, token.output, and token.total.
#
# For providers that support prompt caching, the cached portion of the
# input is also reported as token.cache_read and token.cache_write.
# Both are a breakdown of token.input, not an addition to it, so
# summing them with token.input would double-count.
#
# Each cache key is set only when the provider reported that count. An
# absent key means the response carried no cache information; a zero
# means the provider reported a cache miss. Providers whose API has no
# cache write concept, such as Google, never set token.cache_write;
# OpenAI reports cache writes as cache_write_tokens on both Chat
# Completions and Responses API usage.
#
# Reasoning / thinking tokens are reported as token.reasoning when the
# provider sends them:
# - OpenAI / Azure Chat Completions:
# usage.completion_tokens_details.reasoning_tokens
# (a breakdown of token.output, not an addition)
# - OpenAI / Azure Responses API:
# usage.output_tokens_details.reasoning_tokens
# (a breakdown of token.output, not an addition)
# - Anthropic: usage.output_tokens_details.thinking_tokens
# (a breakdown of token.output; on streams, the final message_delta)
# - Google: usageMetadata.thoughtsTokenCount
# (separate from candidatesTokenCount / token.output; already in
# the provider total when totalTokenCount is present, otherwise
# included in the computed fallback total)
# - Bedrock Converse: not reported (no documented field)
# An absent key means the provider did not report reasoning; a zero
# means it reported none.
#
listeners:
- name: default
address: "127.0.0.1:8080"
filter_chains:
- main
filter_chains:
- name: main
filters:
- filter: router
routes:
- path_prefix: "/"
cluster: backend
# token_usage_headers is declared before token_count: response
# hooks run in reverse declared order, so token_count's
# on_response_body (which sets filter_metadata) runs before
# token_usage_headers reads it.
- filter: token_usage_headers
- filter: token_count
provider: openai # openai | anthropic | google | bedrock | azure
- filter: access_log
- filter: load_balancer
clusters:
- name: backend
endpoints:
- "127.0.0.1:3000"
insecure_options:
allow_private_endpoints: true # example proxies to local backends