Responses To Chat Completions Reasoning

Same pipeline as responses-to-chat-completions.yaml, but targets a vLLM backend that returns raw reasoning in choices[].message.reasoning

Category: Setup-dependent integration
Task: Same pipeline as responses-to-chat-completions.yaml, but targets a vLLM backend that returns raw reasoning in choices[].message.reasoning

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

Companion resources from the same snapshot:

# Responses API to Chat Completions with reasoning translation
# Requires `--no-default-features --features store-sqlite` because SQLite is opt-in.
#
# Same pipeline as responses-to-chat-completions.yaml, but targets a vLLM
# backend that returns raw reasoning in choices[].message.reasoning. The reasoning
# dialect promotes that raw reasoning into a Responses reasoning output item.
#
# Reasoning dialect contract (see the responses_to_chat_completions filter):
#   - dialect: vllm reads message.reasoning (the current vLLM field) and
#     falls back to the deprecated message.reasoning_content alias
#     for older backends, emitting it as a reasoning item whose content is
#     reasoning_text. Raw reasoning is never placed in the item summary, which
#     is reserved for safe summaries.
#   - vLLM has no safe-summary capability, so a client that requests
#     reasoning.summary (or the deprecated reasoning.generate_summary) is
#     rejected with a 400 before the request is forwarded.
#   - max_reasoning_bytes bounds the raw reasoning preserved per
#     response; oversized reasoning fails closed rather than truncating.
#   - On a continuation, raw reasoning is replayed in the following
#     assistant turn's reasoning field, preserving its ordinary content.
#   - Reasoning-only output is replayed as its own assistant message before
#     the next non-assistant item or at the end of input.
#   - Reasoning input requires an enabled dialect and non-empty raw reasoning_text
#     content. Encrypted, summary-only, and malformed items are rejected with 400.
#
# Streaming reasoning translation is out of scope for this base filter and is
# tracked in https://github.com/praxis-proxy/ai/issues/36.

listeners:
  - name: responses-chat-reasoning-gateway
    address: "127.0.0.1:8080"
    filter_chains: [responses-chat-reasoning-translation]

filter_chains:
  - name: responses-chat-reasoning-translation
    filters:
      - filter: openai_responses_format
        on_invalid: reject
        headers:
          format: x-praxis-ai-format
          model: x-praxis-ai-model
          stream: x-praxis-ai-stream

      - filter: openai_responses_validate

      - filter: state_owner
        mode: single_tenant
        tenant_id: default

      - filter: openai_response_store
        backend: sqlite
        database_url: "sqlite://responses.db?mode=rwc"
        responses_table: openai_responses
        conversations_table: openai_conversations

      - filter: openai_responses_rehydrate

      - filter: responses_to_chat_completions
        max_rewritten_body_bytes: 67108864
        reasoning:
          dialect: vllm
          max_reasoning_bytes: 65536

      - filter: path_rewrite
        replace:
          pattern: "^/v1/responses/?$"
          replacement: "/v1/chat/completions"
        conditions:
          - when:
              path_prefix: "/v1/responses"
              methods: [POST]

      - filter: router
        routes:
          - path_prefix: "/"
            cluster: "chat-completions-backend"

      - filter: load_balancer
        clusters:
          - name: "chat-completions-backend"
            endpoints:
              - "127.0.0.1:3001"

insecure_options:
  allow_private_endpoints: true # example proxies to local backends