Time To First Token
Measures the elapsed time from request receipt to the first non-empty SSE body chunk and records a praxis_ai_ttft_seconds Prometheus histogram labeled by model
Category: Setup-dependent integration
Task: Measures the elapsed time from request receipt to the first non-empty SSE body chunk and records a praxis_ai_ttft_seconds Prometheus histogram labeled by model
Prerequisites: The external service, credentials, or certificates referenced by this configuration.
Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
# Time to First Token (TTFT)
#
# Measures the elapsed time from request receipt to the first non-empty
# SSE body chunk and records a praxis_ai_ttft_seconds Prometheus
# histogram labeled by model.
#
listeners:
- name: default
address: "127.0.0.1:8080"
filter_chains:
- main
filter_chains:
- name: main
filters:
- filter: openai_responses_format
- filter: router
routes:
- path_prefix: "/"
cluster: backend
- filter: time_to_first_token
- filter: load_balancer
clusters:
- name: backend
endpoints:
- "127.0.0.1:3000"
insecure_options:
allow_private_endpoints: true # example proxies to local backends