responses_to_chat_completions
On this page
Translates canonical Responses create requests for a Chat Completions backend.
Configuration Notes
The filter consumes the classification metadata and ResponsesState produced by openai_responses_format and openai_responses_validate. It converts the enriched request to Chat Completions wire format, converts finite successful Chat responses back to Responses resources, and normalizes finite provider errors while preserving their HTTP status. OpenAI-managed prompt template references fail closed because Chat Completions has no equivalent field and silently dropping them would change the requested prompt. Supported hosted web-search tools are exposed to the Chat backend as a private, bounded web_search function. Returned calls are restored to canonical web_search_call output before downstream agentic filters run. Streaming Chat Completions SSE responses are translated incrementally into Responses SSE events by an internal state machine that reuses the finite translation builders for the terminal snapshot.
Configure path_rewrite after this filter when the upstream endpoint must change from /v1/responses to /v1/chat/completions.
Requests using previous_response_id require openai_response_store and openai_responses_rehydrate earlier in the request pipeline. The filter fails closed if stored history has not been resolved, preventing a continuation from silently losing prior turns. For finite web-search loops, place openai_web_search and openai_agentic_loop before this filter in an iterative-router step. The reverse response order then restores the hosted call before those filters inspect it.
The optional reasoning block selects a dialect that promotes raw chain-of-thought returned by the backend into a Responses reasoning output item. The default dialect none performs no extraction and preserves only portable Chat Completions fields. The vllm dialect reads the current message.reasoning field (falling back to the deprecated message.reasoning_content alias) and emits it as a reasoning item whose content is reasoning_text. Raw reasoning is never placed in the item summary, which is reserved for safe summaries. No current dialect can generate a safe summary, so a client that requests reasoning.summary (or the deprecated reasoning.generate_summary) is rejected. Streaming reasoning translation is not yet implemented, so a streaming request is rejected when valid reasoning dialect is configured. On continuation, raw reasoning is replayed into the following assistant turn’s reasoning field, preserving its ordinary content. Reasoning-only output becomes a standalone assistant message at a turn boundary or end of input. Reasoning input requires an enabled dialect and non-empty raw reasoning_text content; encrypted, summary-only, and malformed items are rejected before forwarding.
To emit translated SSE events incrementally, this filter forces the reconciled response body mode to Stream for the entire filter chain. The protocol layer reconciles a single chain-wide response body mode with no per-filter provenance, so this downgrade cannot be scoped to one neighbor: it overrides every response filter’s StreamBuffer requirement, not only openai_response_store’s. Only compose response-body filters after this one that tolerate incremental fragments; openai_stream_events and openai_response_store are compatible because they persist streamed turns from the accumulator, not a buffered body, whereas any other response-body rewriter needing the complete buffered body would instead receive fragments.
Examples
Example 1
filter: responses_to_chat_completions
Example 2
filter: responses_to_chat_completions
max_rewritten_body_bytes: 67108864
reasoning:
dialect: vllm
max_reasoning_bytes: 65536