Token Ceiling
Place token_ceiling after any request translation or prompt enrichment so it evaluates the final provider-bound JSON body
Versions marked “overview” do not contain this page. Selecting one opens that version’s documentation overview.
Category: Setup-dependent integration
Task: Place token_ceiling after any request translation or prompt enrichment so it evaluates the final provider-bound JSON body
Prerequisites: The external service, credentials, or certificates referenced by this configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
# Pre-forward token ceiling guard.
#
# Enable with:
# cargo run -p praxis-ai-proxy --features token-ceiling-filter -- \
# -c examples/configs/token-ceiling.yaml
#
# Place token_ceiling after any request translation or prompt enrichment so it
# evaluates the final provider-bound JSON body. Missing output limits are
# rejected when max_output_tokens is configured; values are never clamped.
listeners:
- name: default
address: "127.0.0.1:8080"
filter_chains:
- main
filter_chains:
- name: main
filters:
- filter: router
routes:
- path: "/v1/chat/completions"
cluster: upstream
- filter: token_ceiling
max_input_tokens: 4000
max_output_tokens: 1024
tokenizer: cl100k_base
- filter: load_balancer
clusters:
- name: upstream
endpoints:
- "127.0.0.1:9000"
insecure_options:
allow_private_endpoints: true # example proxies to a local backend