Time To First Token

Measures the elapsed time from request receipt to the first non-empty SSE body chunk and records a praxis_ai_ttft_seconds Prometheus histogram labeled by model

Category: Setup-dependent integration
Task: Measures the elapsed time from request receipt to the first non-empty SSE body chunk and records a praxis_ai_ttft_seconds Prometheus histogram labeled by model

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# Time to First Token (TTFT)
#
# Measures the elapsed time from request receipt to the first non-empty
# SSE body chunk and records a praxis_ai_ttft_seconds Prometheus
# histogram labeled by model.
#
listeners:
  - name: default
    address: "127.0.0.1:8080"
    filter_chains:
      - main

filter_chains:
  - name: main
    filters:
      - filter: openai_responses_format

      - filter: router
        routes:
          - path_prefix: "/"
            cluster: backend

      - filter: time_to_first_token

      - filter: load_balancer
        clusters:
          - name: backend
            endpoints:
              - "127.0.0.1:3000"

insecure_options:
  allow_private_endpoints: true # example proxies to local backends