Responses

41 configurations in Responses for Praxis AI.

Configurations in Responses.

Configurations

  • Agentic Loop Deferred Mcp Fixture — Minimal agentic-loop pipeline that sanitizes deferred MCP connectors on the first inference round. defer_loading: true skips tools/list, so replay never opens an independent MCP callout
  • Agentic Loop Fixture — Minimal agentic loop pipeline for inference fixture replay
  • Agentic Loop — Demonstrates the openai_agentic_loop filter with iterative_request_router for step-based model-tool-model looping in the Responses API
  • Body Size Limits — Demonstrates how raw request body size is enforced across a chain of OpenAI Responses filters that each buffer the request body
  • Client Tool Compat Chat Completions — Lets a rich Codex-style Responses client (POST /v1/responses with custom, namespace, local shell, and client-executed tool_search tools) reach a function-only Chat Completions backend (POST /v1/chat/completions) by composing openai_client_tool_compat with responses_to_chat_completions in one iterative-router step (GitHub issue #1206)
  • Client Tool Compat — Lets a rich Codex-style Responses client talk to a function-only Responses backend (for example vLLM at POST /v1/responses) without routing through /v1/chat/completions and without executing client-owned tools in Praxis
  • Codex Http Chat Translation — Release-acceptance configuration for GitHub issue #870. A Codex client speaks the OpenAI Responses API over HTTP while Praxis selects a Chat Completions-only provider, translates both request and streaming response, rewrites the provider path, and replaces the client credential
  • Compact — Demonstrates compaction after rehydrate, file resolve, and document extract so rewritten current-turn content survives history replacement
  • Doc Extract — Converts input_file content parts to input_text for inference backends that do not natively support input_file (e.g. vLLM, llm-d)
  • File Resolve — Resolves file_id and file_url references in Responses API input by fetching file metadata and content, then inlining base64 content as file_data or image_url before forwarding
  • File Search Callout — Demonstrates hosted file_search execution under the unified agentic loop (#1046). openai_agentic_loop is the sole loop owner: it parses each model response, records the file_search_call items it sees as assignments, and publishes the single continuation signal (action=loop|done). openai_file_search_callout is a pure request-phase dispatcher: at request- body EOS on each IRR re-entry it executes the assigned calls against the vector store and reconciles each item in place, then the owner prepares the next inference request
  • File Search Chat Completions Fixture — Single-upstream fixture configuration for recording the private Chat Completions function representation of a Responses file_search tool
  • File Search Chat Completions — Accepts finite OpenAI Responses requests with hosted file search while targeting a backend that only implements /v1/chat/completions
  • File Search Streaming — Demonstrates streaming hosted file_search through the iterative_request_router
  • Format Routing — Routes AI API traffic by request-head operation identity and body format
  • Full Flow Agentic — Runs the complete Responses API pipeline through an agentic iterative_request_router that executes hosted file_search, web_search, and MCP tool calls in a model-tool-model loop, persisting both buffered and streaming (stream: true) responses
  • Http Passthrough — The strict-HTTP acceptance test for pinned Codex CLI for GitHub issue #870 drives the proxy over POST /v1/responses and any other Responses API paths the client may probe
  • Irr Terminal Streaming — Demonstrates a single-step iterative_request_router pipeline that exposes a native OpenAI Responses SSE body incrementally. openai_responses_proxy always advertises the streaming capability and selects Praxis’s typed streaming transport automatically for an effective “stream”: true request; there is no operator opt-in
  • Mcp Dispatch — Demonstrates the openai_mcp_dispatch filter configuration
  • Mcp Outbound Chain — Demonstrates binding an operator outbound_chain onto the outbound MCP callout made by openai_mcp_tool_resolve (the tools/list discovery request)
  • Mcp Streaming — Demonstrates MCP tool calls over the filtered-subrequest transport with SSE streaming support
  • Mcp Tool Resolve — Demonstrates the openai_mcp_tool_resolve filter, which resolves MCP tool entries in the Responses API tools array into concrete tool definitions by calling tools/list on each upstream MCP server
  • Model Rewrite — Rewrites or injects the top-level model field in Responses API and Chat Completions request bodies before forwarding to the inference backend
  • Rehydrate Fixture — Minimal native OpenAI Responses pipeline that stores a first turn, rehydrates a stored previous_response_id into the outbound input history, and proxies to a native /v1/responses backend
  • Rehydrate — Validates previous_response_id by fetching the stored response, confirming its status is completed, and promoting the ID to filter metadata
  • Request Validate — Validates Responses API JSON and enriches request metadata
  • Response Store Postgres Mtls — Persists non-streaming Responses API responses to PostgreSQL over a TLS-verified connection that authenticates with a client certificate instead of a password
  • Response Store — Persists non-streaming Responses API responses to a database and serves stored data via GET endpoints and handles DELETE /v1/responses/{id} locally
  • Responses Proxy — Forward Responses API traffic through a Praxis AI listener to a compatible backend.
  • Responses Routing — Routes Responses API traffic by detected mode
  • Responses To Chat Completions Reasoning — Same pipeline as responses-to-chat-completions.yaml, but targets a vLLM backend that returns raw reasoning in choices[].message.reasoning
  • Responses To Chat Completions — Translate the Responses API shape for an OpenAI-compatible Chat Completions backend.
  • State Ownership — Provider-neutral owner isolation for persisted OpenAI Responses and Conversations
  • Stream Events — Demonstrates the openai_stream_events filter, which composes the current iterative-request-router (IRR) execution into one logical Responses SSE stream: it parses each backend SSE chunk, accumulates state (response object, output items, tool calls, usage) into ResponsesState, normalizes the per-round lifecycle, and preserves parser state through stream completion
  • Tool Routing — Demonstrates using openai_tool_parse to route Responses API requests by their tool composition
  • Vector Stores Routing — Routes /v1/vector_stores traffic and all its subresources to a dedicated backend (any server compatible with the OpenAI Files / Vector Stores API), while sending everything else to a default backend
  • Vllm Agentic Api — vLLM Agentic API: https://github.com/vllm-project/agentic-api
  • Web Search Chat Completions Fixture — Single-upstream fixture configuration for recording the private Chat Completions function representation of a Responses web_search tool
  • Web Search Chat Completions — Accepts OpenAI Responses requests with hosted web search while targeting a backend that only implements /v1/chat/completions
  • Web Search Scoped Credentials — A scoped-credentials variant of web-search-chat-completions.yaml
  • Web Search — Demonstrates the openai_web_search filter configuration

Agentic Loop

Demonstrates the openai_agentic_loop filter with iterative_request_router for step-based model-tool-model looping in the Responses API

Agentic Loop Deferred Mcp Fixture

Minimal agentic-loop pipeline that sanitizes deferred MCP connectors on the first inference round. defer_loading: true skips tools/list, so replay never opens an independent MCP callout

Agentic Loop Fixture

Minimal agentic loop pipeline for inference fixture replay

Body Size Limits

Demonstrates how raw request body size is enforced across a chain of OpenAI Responses filters that each buffer the request body

Client Tool Compat

Lets a rich Codex-style Responses client talk to a function-only Responses backend (for example vLLM at POST /v1/responses) without routing through /v1/chat/completions and without executing client-owned tools in Praxis

Client Tool Compat Chat Completions

Lets a rich Codex-style Responses client (POST /v1/responses with custom, namespace, local shell, and client-executed tool_search tools) reach a function-only Chat Completions backend (POST /v1/chat/completions) by composing openai_client_tool_compat with responses_to_chat_completions in one iterative-router step (GitHub issue #1206)

Codex Http Chat Translation

Release-acceptance configuration for GitHub issue #870. A Codex client speaks the OpenAI Responses API over HTTP while Praxis selects a Chat Completions-only provider, translates both request and streaming response, rewrites the provider path, and replaces the client credential

Compact

Demonstrates compaction after rehydrate, file resolve, and document extract so rewritten current-turn content survives history replacement

Doc Extract

Converts input_file content parts to input_text for inference backends that do not natively support input_file (e.g. vLLM, llm-d)

File Resolve

Resolves file_id and file_url references in Responses API input by fetching file metadata and content, then inlining base64 content as file_data or image_url before forwarding

File Search Callout

Demonstrates hosted file_search execution under the unified agentic loop (#1046). openai_agentic_loop is the sole loop owner: it parses each model response, records the file_search_call items it sees as assignments, and publishes the single continuation signal (action=loop|done). openai_file_search_callout is a pure request-phase dispatcher: at request- body EOS on each IRR re-entry it executes the assigned calls against the vector store and reconciles each item in place, then the owner prepares the next inference request

File Search Chat Completions

Accepts finite OpenAI Responses requests with hosted file search while targeting a backend that only implements /v1/chat/completions

File Search Chat Completions Fixture

Single-upstream fixture configuration for recording the private Chat Completions function representation of a Responses file_search tool

File Search Streaming

Demonstrates streaming hosted file_search through the iterative_request_router

Format Routing

Routes AI API traffic by request-head operation identity and body format

Full Flow Agentic

Runs the complete Responses API pipeline through an agentic iterative_request_router that executes hosted file_search, web_search, and MCP tool calls in a model-tool-model loop, persisting both buffered and streaming (stream: true) responses

Http Passthrough

The strict-HTTP acceptance test for pinned Codex CLI for GitHub issue #870 drives the proxy over POST /v1/responses and any other Responses API paths the client may probe

Irr Terminal Streaming

Demonstrates a single-step iterative_request_router pipeline that exposes a native OpenAI Responses SSE body incrementally. openai_responses_proxy always advertises the streaming capability and selects Praxis’s typed streaming transport automatically for an effective “stream”: true request; there is no operator opt-in

Mcp Dispatch

Demonstrates the openai_mcp_dispatch filter configuration

Mcp Outbound Chain

Demonstrates binding an operator outbound_chain onto the outbound MCP callout made by openai_mcp_tool_resolve (the tools/list discovery request)

Mcp Streaming

Demonstrates MCP tool calls over the filtered-subrequest transport with SSE streaming support

Mcp Tool Resolve

Demonstrates the openai_mcp_tool_resolve filter, which resolves MCP tool entries in the Responses API tools array into concrete tool definitions by calling tools/list on each upstream MCP server

Model Rewrite

Rewrites or injects the top-level model field in Responses API and Chat Completions request bodies before forwarding to the inference backend

Rehydrate

Validates previous_response_id by fetching the stored response, confirming its status is completed, and promoting the ID to filter metadata

Rehydrate Fixture

Minimal native OpenAI Responses pipeline that stores a first turn, rehydrates a stored previous_response_id into the outbound input history, and proxies to a native /v1/responses backend

Request Validate

Validates Responses API JSON and enriches request metadata

Response Store

Persists non-streaming Responses API responses to a database and serves stored data via GET endpoints and handles DELETE /v1/responses/{id} locally

Response Store Postgres Mtls

Persists non-streaming Responses API responses to PostgreSQL over a TLS-verified connection that authenticates with a client certificate instead of a password

Responses Proxy

Forward Responses API traffic through a Praxis AI listener to a compatible backend.

Responses Routing

Routes Responses API traffic by detected mode

Responses To Chat Completions

Translate the Responses API shape for an OpenAI-compatible Chat Completions backend.

Responses To Chat Completions Reasoning

Same pipeline as responses-to-chat-completions.yaml, but targets a vLLM backend that returns raw reasoning in choices[].message.reasoning

State Ownership

Provider-neutral owner isolation for persisted OpenAI Responses and Conversations

Stream Events

Demonstrates the openai_stream_events filter, which composes the current iterative-request-router (IRR) execution into one logical Responses SSE stream: it parses each backend SSE chunk, accumulates state (response object, output items, tool calls, usage) into ResponsesState, normalizes the per-round lifecycle, and preserves parser state through stream completion

Tool Routing

Demonstrates using openai_tool_parse to route Responses API requests by their tool composition

Vector Stores Routing

Routes /v1/vector_stores traffic and all its subresources to a dedicated backend (any server compatible with the OpenAI Files / Vector Stores API), while sending everything else to a default backend

Vllm Agentic Api

vLLM Agentic API: https://github.com/vllm-project/agentic-api

Web Search

Demonstrates the openai_web_search filter configuration

Web Search Chat Completions

Accepts OpenAI Responses requests with hosted web search while targeting a backend that only implements /v1/chat/completions

Web Search Chat Completions Fixture

Single-upstream fixture configuration for recording the private Chat Completions function representation of a Responses web_search tool

Web Search Scoped Credentials

A scoped-credentials variant of web-search-chat-completions.yaml