Responses Routing

Routes Responses API traffic by detected mode

Category: Setup-dependent integration
Task: Routes Responses API traffic by detected mode

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

Run it: Use ghcr.io/praxis-proxy/ai:0.4.1 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# Responses API Mode Routing
#
# Routes Responses API traffic by detected mode. The `openai_responses_format`
# filter classifies requests as `stateless` (simple inference, no
# state management) or `stateful` (needs conversation history, tools,
# or persistence).
#
# Both modes use the same inference backend. The difference is whether
# the request passes through the stateful filter chain (orchestration,
# persistence, tool dispatch) or bypasses it entirely.
#
# A request is stateful if any of:
# - previous_response_id is set
# - tools array is non-empty
# - store is true (the OpenAI spec default when omitted)
# - conversation is set
# - prompt.id is set
#
# Stateless requires store=false and no other stateful markers.

listeners:
  - name: ai-gateway
    address: "127.0.0.1:8080"
    filter_chains: [responses-routing]

filter_chains:
  - name: responses-routing
    filters:
      - filter: openai_responses_format
        on_invalid: reject
        branch_chains:
          - name: stateful
            on_result:
              filter: openai_responses_format
              key: mode
              result: stateful
            chains:
              - name: stateful-chain
                filters:
                  # Placeholder: the full stateful chain would include
                  # request_validate, response_store, rehydrate,
                  # openai_tool_parse, openai_responses_proxy, stream_events,
                  # tool_dispatch, reasoning, etc.
                  #
                  # For now, proxy to the same backend. Future
                  # filters will be inserted here as they land.
                  - filter: router
                    routes:
                      - path_prefix: "/"
                        cluster: "inference-backend"
            rejoin: shared_load_balancer

      # Stateless: no branch matched — proxy directly to backend
      # with zero additional filter processing.
      - filter: router
        routes:
          - path_prefix: "/"
            cluster: "inference-backend"
      - name: shared_load_balancer
        filter: load_balancer
        clusters:
          - name: "inference-backend"
            endpoints:
              - "127.0.0.1:3001"

insecure_options:
  allow_private_endpoints: true # example proxies to local backends