Responses To Chat Completions Reasoning
Same pipeline as responses-to-chat-completions.yaml, but targets a vLLM backend that returns raw reasoning in choices[].message.reasoning
Category: Setup-dependent integration
Task: Same pipeline as responses-to-chat-completions.yaml, but targets a vLLM backend that returns raw reasoning in choices[].message.reasoning
Prerequisites: The external service, credentials, or certificates referenced by this configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
Companion resources from the same snapshot:
# Responses API to Chat Completions with reasoning translation
# Requires `--no-default-features --features store-sqlite` because SQLite is opt-in.
#
# Same pipeline as responses-to-chat-completions.yaml, but targets a vLLM
# backend that returns raw reasoning in choices[].message.reasoning. The reasoning
# dialect promotes that raw reasoning into a Responses reasoning output item.
#
# Reasoning dialect contract (see the responses_to_chat_completions filter):
# - dialect: vllm reads message.reasoning (the current vLLM field) and
# falls back to the deprecated message.reasoning_content alias
# for older backends, emitting it as a reasoning item whose content is
# reasoning_text. Raw reasoning is never placed in the item summary, which
# is reserved for safe summaries.
# - vLLM has no safe-summary capability, so a client that requests
# reasoning.summary (or the deprecated reasoning.generate_summary) is
# rejected with a 400 before the request is forwarded.
# - max_reasoning_bytes bounds the raw reasoning preserved per
# response; oversized reasoning fails closed rather than truncating.
# - On a continuation, raw reasoning is replayed in the following
# assistant turn's reasoning field, preserving its ordinary content.
# - Reasoning-only output is replayed as its own assistant message before
# the next non-assistant item or at the end of input.
# - Reasoning input requires an enabled dialect and non-empty raw reasoning_text
# content. Encrypted, summary-only, and malformed items are rejected with 400.
#
# Streaming reasoning translation is out of scope for this base filter and is
# tracked in https://github.com/praxis-proxy/ai/issues/36.
listeners:
- name: responses-chat-reasoning-gateway
address: "127.0.0.1:8080"
filter_chains: [responses-chat-reasoning-translation]
filter_chains:
- name: responses-chat-reasoning-translation
filters:
- filter: openai_responses_format
on_invalid: reject
headers:
format: x-praxis-ai-format
model: x-praxis-ai-model
stream: x-praxis-ai-stream
- filter: openai_responses_validate
- filter: state_owner
mode: single_tenant
tenant_id: default
- filter: openai_response_store
backend: sqlite
database_url: "sqlite://responses.db?mode=rwc"
responses_table: openai_responses
conversations_table: openai_conversations
- filter: openai_responses_rehydrate
- filter: responses_to_chat_completions
max_rewritten_body_bytes: 67108864
reasoning:
dialect: vllm
max_reasoning_bytes: 65536
- filter: path_rewrite
replace:
pattern: "^/v1/responses/?$"
replacement: "/v1/chat/completions"
conditions:
- when:
path_prefix: "/v1/responses"
methods: [POST]
- filter: router
routes:
- path_prefix: "/"
cluster: "chat-completions-backend"
- filter: load_balancer
clusters:
- name: "chat-completions-backend"
endpoints:
- "127.0.0.1:3001"
insecure_options:
allow_private_endpoints: true # example proxies to local backends