Responses To Chat Completions

Translate the Responses API shape for an OpenAI-compatible Chat Completions backend.

Category: Setup-dependent integration
Task: Translate Responses API requests for a Chat Completions backend

Prerequisites: Praxis AI runtime; An OpenAI-compatible Chat Completions backend; Provider credentials when required

Expected outcome: Responses API requests are translated for the configured Chat Completions backend.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# Responses API to Chat Completions
# Requires `--features store-sqlite` because these filters are opt-in.
#
# Accepts OpenAI Responses create requests, including finite stored
# continuations, while targeting a backend that only implements
# /v1/chat/completions.
#
# Request order:
#   1. openai_responses_format classifies the client request.
#   2. openai_responses_validate performs the canonical JSON parse and
#      creates ResponsesState for downstream enrichment and translation.
#   3. openai_response_store initializes the shared response store.
#   4. openai_responses_rehydrate resolves previous_response_id into the
#      complete message history required by Chat Completions.
#   5. iterative_request_router runs one inference step (max_iterations: 1)
#      that composes the translated stream and reaches the backend:
#      a. openai_stream_events accumulates the translated terminal Responses
#         resource from the SSE stream so the store can persist streamed turns.
#      b. responses_to_chat_completions converts the canonical request to
#         Chat Completions, converts finite JSON responses back to a Responses
#         resource, and translates streaming Chat Completions SSE into Responses
#         SSE events incrementally.
#      c. path_rewrite changes the upstream endpoint explicitly.
#
# `openai_stream_events` always composes the current IRR execution into one
# logical Responses stream, so it must run inside an `iterative_request_router`
# step; a `stream: true` request placed outside IRR fails closed at request
# time. A single inference round is just a one-round logical stream.
#
# Response filters run in reverse order. Within the IRR step:
#
#   responses_to_chat_completions -> openai_stream_events
#
# then the pre-IRR openai_response_store runs last. The Chat response (finite
# or streamed) is therefore translated into a client-visible Responses
# resource, its terminal resource is accumulated by openai_stream_events, and
# openai_response_store then decides whether to persist it.

listeners:
  - name: responses-chat-gateway
    address: "127.0.0.1:8080"
    filter_chains: [responses-chat-translation]

filter_chains:
  - name: responses-chat-translation
    filters:
      - filter: openai_responses_format
        on_invalid: reject
        headers:
          format: x-praxis-ai-format
          model: x-praxis-ai-model
          stream: x-praxis-ai-stream

      - filter: openai_responses_validate

      - filter: state_owner
        mode: single_tenant
        tenant_id: default

      - filter: openai_response_store
        backend: sqlite
        database_url: "sqlite://responses.db?mode=rwc"
        responses_table: openai_responses
        conversations_table: openai_conversations

      - filter: openai_responses_rehydrate

      - filter: iterative_request_router
        initial_step: inference
        max_iterations: 1
        # The OpenResponses suite drives concurrent requests through a CPU
        # backend, and production model streams may also exceed the IRR's 30s
        # default. This single step inherits an 11-minute total deadline for
        # time-to-first-byte plus response streaming; the accumulator below
        # independently caps its portion of the stream at 10 minutes.
        timeout_ms: 660000
        steps:
          - name: inference
            filters:
              - filter: openai_stream_events
                # Accumulates every Responses SSE event the converter emits below
                # so the store can persist streamed turns. Always composes the
                # current IRR execution into one logical Responses stream, so it
                # must run inside this step. Its limits must never be tighter than
                # the converter's: if this filter aborted a stream the client had
                # already received, that turn would silently skip persistence.
                #
                # - max_events sits at or above the converter's max_stream_events,
                #   which bounds the converter's TOTAL emitted events (terminal
                #   included), so an accepted stream never exceeds this event budget.
                # - max_buffer_bytes must hold the largest single frame the converter
                #   can emit. The converter refuses to emit any frame larger than its
                #   max_emitted_sse_frame_bytes, so this buffer is set strictly above
                #   that ceiling (leaving SSE-framing headroom) so no frame the client
                #   receives is ever rejected here and dropped from the store.
                # - timeout_secs is no shorter than the converter's stream_timeout_secs.
                max_buffer_bytes: 67108864
                max_events: 200000
                timeout_secs: 600

              - filter: responses_to_chat_completions
                # The converter is the authority on stream bounds and fails closed
                # when it exceeds them. Keep every limit at or below
                # openai_stream_events' above so any stream the client receives is
                # always within the accumulator's capacity and is never silently
                # dropped from the store.
                #
                # The size ceilings form a chain, smallest to largest:
                #   max_rewritten_body_bytes <= max_emitted_sse_frame_bytes
                #     < accumulator max_buffer_bytes
                # - max_rewritten_body_bytes bounds the response resource a
                #   terminal/lifecycle frame carries.
                # - max_emitted_sse_frame_bytes bounds the COMPLETE encoded frame
                #   (resource plus SSE framing). Kept at or above
                #   max_rewritten_body_bytes so legitimate frames are not rejected,
                #   and strictly below the accumulator's max_buffer_bytes so every
                #   emitted frame fits the buffer it is reassembled in. A frame that
                #   would exceed this ceiling fails closed with a minimal
                #   response.failed instead of being emitted and later dropped
                #   downstream.
                max_rewritten_body_bytes: 33554432
                max_emitted_sse_frame_bytes: 50331648
                max_stream_events: 100000
                stream_timeout_secs: 300

              - filter: path_rewrite
                replace:
                  pattern: "^/v1/responses/?$"
                  replacement: "/v1/chat/completions"
                conditions:
                  - when:
                      path_prefix: "/v1/responses"
                      methods: [POST]

              - filter: router
                routes:
                  - path_prefix: "/"
                    cluster: "chat-completions-backend"

              - filter: load_balancer
                clusters:
                  - name: "chat-completions-backend"
                    endpoints:
                      - "127.0.0.1:3001"
            on_result:
              - default: true
                done: true

insecure_options:
  allow_private_endpoints: true # example proxies to local backends