Client Tool Compat Chat Completions

Lets a rich Codex-style Responses client (POST /v1/responses with custom, namespace, local shell, and client-executed tool_search tools) reach a function-only Chat Completions backend (POST /v1/chat/completions) by composing openai_client_tool_compat with responses_to_chat_completions in one iterative-router step (GitHub issue #1206)

Category: Setup-dependent integration
Task: Lets a rich Codex-style Responses client (POST /v1/responses with custom, namespace, local shell, and client-executed tool_search tools) reach a function-only Chat Completions backend (POST /v1/chat/completions) by composing openai_client_tool_compat with responses_to_chat_completions in one iterative-router step (GitHub issue #1206)

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# Client Tool Compatibility over a Chat Completions backend
# Requires `--features store-sqlite` because these filters are opt-in.
#
# Lets a rich Codex-style Responses client (POST /v1/responses with custom,
# namespace, local shell, and client-executed tool_search tools) reach a
# function-only Chat Completions backend (POST /v1/chat/completions) by composing
# openai_client_tool_compat with responses_to_chat_completions in one
# iterative-router step (GitHub issue #1206).
#
# Client boundary:
#   POST /v1/responses
#   Authorization: Bearer <gateway credential>
#   OpenAI Responses JSON and SSE (rich typed items)
#
# Provider boundary:
#   POST /v1/chat/completions
#   Authorization: Bearer <provider credential>
#   OpenAI Chat Completions JSON and SSE (function tools only)
#
# Request-path order inside the inference step is significant and test-locked:
#
#   openai_stream_events      arms streaming restoration (request-phase marker)
#   openai_agentic_loop       prepares history / drives continuation rounds
#   openai_client_tool_compat lowers rich client tools to private function tools
#                             in request_body ONLY (canonical state.tools stays
#                             rich for restore); fails closed on a client tool
#                             named a reserved hosted call (file_search/web_search)
#   responses_to_chat_completions  translates the lowered Responses request into a
#                             Chat Completions request (reads the lowered
#                             request_body tools via request_tools/request_tool_choice)
#   path_rewrite              /v1/responses -> /v1/chat/completions
#   router / credential_injection / load_balancer
#
# Response filters run in reverse, so responses_to_chat_completions rebuilds the
# Responses object first, then openai_client_tool_compat restores the private
# function_call items to their canonical typed items BEFORE openai_agentic_loop
# observes them (this ordering keeps a client tool lowered to a private function
# name from being misclassified as a hosted call). No private lowered name leaks
# to the client. When the request declares no rich client tools the compat filter
# is a transparent passthrough, so plain-function traffic is unchanged.

listeners:
  - name: ai-gateway
    address: "127.0.0.1:8080"
    filter_chains: [client-tool-compat-chat]

filter_chains:
  - name: client-tool-compat-chat
    filters:
      - filter: openai_responses_format
        on_invalid: reject
        headers:
          format: x-praxis-ai-format
          model: x-praxis-ai-model
          stream: x-praxis-ai-stream

      - filter: openai_responses_validate

      # Persisted Responses/Conversations state is scoped to an immutable owner.
      # This single-tenant demo installs a fixed owner directly (no auth gateway
      # in front); a multi-tenant deployment would use mode: trusted_headers so
      # an upstream gateway asserts the tenant/subject instead.
      - filter: state_owner
        mode: single_tenant
        tenant_id: default

      - filter: openai_response_store
        backend: sqlite
        database_url: "sqlite://responses.db?mode=rwc"
        responses_table: openai_responses
        conversations_table: openai_conversations

      - filter: openai_responses_rehydrate

      - filter: iterative_request_router
        initial_step: inference
        max_iterations: 4
        timeout_ms: 30000
        steps:
          - name: inference
            filters:
              - filter: openai_stream_events
                # Must exceed the translator's complete-frame ceiling below.
                max_buffer_bytes: 4194304
                max_events: 1000
                timeout_secs: 30

              - filter: openai_agentic_loop
                max_infer_iters: 4

              - filter: openai_client_tool_compat
                max_rewritten_body_bytes: 67108864
                max_client_tools: 512

              - filter: responses_to_chat_completions
                max_rewritten_body_bytes: 1048576
                max_emitted_sse_frame_bytes: 2097152
                max_stream_events: 500
                stream_timeout_secs: 30

              - filter: path_rewrite
                replace:
                  pattern: "^/v1/responses/?$"
                  replacement: "/v1/chat/completions"
                conditions:
                  - when:
                      path_prefix: "/v1/responses"
                      methods: [POST]

              - filter: router
                routes:
                  - path_prefix: "/"
                    cluster: chat-provider

              - filter: credential_injection
                clusters:
                  - name: chat-provider
                    header: Authorization
                    value: "synthetic-provider-key"
                    header_prefix: "Bearer "
                    strip_client_credential: true

              - filter: load_balancer
                clusters:
                  - name: chat-provider
                    endpoints:
                      - "127.0.0.1:3001"
            on_result:
              - default: true
                done: true

insecure_options:
  allow_private_endpoints: true # example proxies to a local scripted backend