Full Flow Agentic
Runs the complete Responses API pipeline through an agentic iterative_request_router that executes hosted file_search, web_search, and MCP tool calls in a model-tool-model loop, persisting both buffered and streaming (stream: true) responses
Category: Setup-dependent integration
Task: Runs the complete Responses API pipeline through an agentic iterative_request_router that executes hosted file_search, web_search, and MCP tool calls in a model-tool-model loop, persisting both buffered and streaming (stream: true) responses
Prerequisites: The external service, credentials, or certificates referenced by this configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
Companion resources from the same snapshot:
# Responses API Full Flow — Unified Agentic Gateway
# Requires `--features openai-file-resolve-filter,openai-conversations,openai-mcp-tools,store-sqlite` because these filters are opt-in.
#
# Runs the complete Responses API pipeline through an agentic
# iterative_request_router that executes hosted file_search, web_search,
# and MCP tool calls in a model-tool-model loop, persisting both buffered
# and streaming (`stream: true`) responses. This is the single config an
# operator would deploy for a Responses API gateway that must serve the
# full agentic loop, the Conversations API, and the dedicated
# prompts/embeddings/files/vector_stores services behind one listener. The
# request trace is continued through every nested web-search, OGX, and MCP
# callout with a fresh span ID per hop.
#
# The request is classified once, then a top-level router freezes one logical
# cluster plus its application_provider/application_protocol metadata. That
# binding drives two mutually exclusive paths:
#
# - application_provider=openai: a terminal branch forwards the original
# Responses request directly. Managed validation, response state,
# persistence/rehydration, hosted-tool resolution, and the IRR never run.
# - every managed provider: the bound-body barrier runs provider-aware body
# filters once, then the IRR executes the agentic loop.
#
# ResponsesState created on the managed path persists across all IRR
# iterations via extension swapping:
#
# 0. state_owner maps deployment-specific trusted ingress headers into one
# normalized tenant/issuer/subject context for owner-aware filters.
# callout_credentials captures per-user Brave and OGX provider keys from
# trusted x-user-brave-key and x-user-ogx-key ingress headers into the
# `brave_search` and `ogx_files` slots, then strips both source headers
# before inference or IRR processing. The trusted boundary MUST
# unconditionally delete-then-set those source headers on every request;
# append-on-success is insufficient because this filter cannot authenticate
# header provenance. The IRR carries the resulting typed request extension
# across iterations, while each callout explicitly selects its slot for its
# exact-authority provider destination.
# project_state_owner_headers then projects that context as x-tenant-id and
# x-user-id for OGX and explicitly allowlisted callouts. The raw ingress
# names are stripped, and the projection is repeated inside the IRR step
# because it owns a separate destination request lifecycle.
# These adapters do not themselves change persistence predicates: this
# context-only configuration must be combined with owner-scoped store
# enforcement before it is exposed as a multi-owner state service.
#
# 0. openai_operation publishes the typed registry match consumed by
# openai_conversations. It must remain earlier in this filter chain.
#
# 1. openai_conversations handles the Conversations API only for managed
# providers. The router binds /v1/conversations first: an OpenAI binding
# takes the direct terminal branch, while a managed binding runs the
# Praxis CRUD implementation. On the managed Responses response path the
# filter also appends input+output items back to the conversation after a
# completed buffered or streamed response. For streams it writes from the
# canonical terminal state before releasing response.completed, including
# when the Response itself has store:false.
#
# 2. openai_responses_format classifies the request and promotes
# format, model, stream, and mode to internal routing headers.
# on_invalid is left at the default (passthrough) so that
# Conversations API traffic is not rejected before the conversations
# filter can handle it in the body phase.
#
# 3. The binding router selects the logical provider/protocol. The filters
# below that declare bound-body access then run once against the buffered
# request. Their `unless application_provider=openai` conditions make the
# OpenAI route true passthrough. openai_responses_validate checks JSON,
# initializes ResponsesState, and rejects gateway-unsupported fields such
# as background=true only on managed routes.
#
# 4. On managed routes, openai_tool_parse parses the tools array and
# tool_choice from Responses API requests and promotes summary facts
# (has_tools, has_web_search, tool_choice, function_count, etc.) to
# metadata and filter results. It does not mutate the request body and is
# skipped with the other gateway-owned body work on direct OpenAI routes.
#
# 5. openai_response_store persists responses to SQLite and registers
# the store backend so downstream filters (rehydrate, compact) can
# read from it. It runs once at the post-binding body barrier before
# the iterative_request_router, so on the response path — which runs in reverse config
# order — it persists AFTER the IRR step has composed the terminal
# event. This is what makes streaming persistence work: the store
# reads the accumulated response object that openai_stream_events
# builds inside the IRR.
#
# 6. openai_responses_rehydrate loads conversation context from
# previous_response_id or conversation field by fetching stored
# responses/conversations and prepending their message history to the
# current request. previous_response_id takes precedence when both
# are present.
#
# 7. openai_file_resolve resolves file_id references in the current
# input and rehydrated history through a configured Files API. Runs
# after rehydrate so stored history is available, and before
# responses_proxy so the rebuilt upstream body contains file_data or
# image_url rather than stale file_id references.
#
# 8. openai_doc_extract converts input_file content parts to input_text
# for backends that do not natively support input_file (e.g. vLLM,
# llm-d). Text-safe content (text/*, application/json,
# application/xml) is decoded from base64, validated as UTF-8, and
# forwarded as plain text. Unsupported formats are left unchanged
# (on_unsupported: continue). This filter is optional — omit it when
# the inference backend natively supports input_file.
#
# 9. openai_mcp_tool_resolve resolves MCP tool entries from the tools
# array by calling tools/list on each upstream MCP server. Runs after
# rehydrate so previous_tools is available for cross-request caching.
# Writes mcp_tool_map to ResponsesState for downstream tool dispatch.
# Projected identity headers are forwarded only to operator-configured
# connector_id targets, never arbitrary client-selected server_url values.
#
# 10. Two terminal branches consume the frozen binding before the IRR.
# One forwards provider-owned OpenAI Responses traffic unchanged; the
# other preserves the direct service/WebSocket routes. Both use
# `cluster_source: bound_upstream`, so no second routing decision can
# disagree with the provider selected for the request.
#
# 11. The iterative_request_router owns the agentic loop for classified
# openai_responses create requests. For a native Responses upstream,
# openai_responses_proxy rebuilds the outbound body from ResponsesState
# and automatically selects the typed streaming transport for an
# effective `stream: true` request, so each backend SSE chunk reaches
# openai_stream_events incrementally; for a Chat Completions upstream,
# responses_to_chat_completions does the equivalent build and translates
# Chat responses back to Responses. Both now run AFTER the step load
# balancer, each gated on the selected upstream's protocol (see the
# "Protocol-adaptive translation" note below).
# The step load balancer consumes the request-level binding directly;
# there is no router inside the IRR. openai_agentic_loop is the sole loop owner, sole
# response parser, and sole transition authority: it extracts
# completed function_call, web_search_call, file_search_call, and MCP
# items, records the assignment each request-phase dispatcher must
# execute next, and publishes the single loop/done decision.
#
# IRR inference filter chain (canonical #1046 order with
# protocol-adaptive translation):
# Request phase (forward):
# openai_stream_events → openai_web_search → openai_mcp_dispatch
# → openai_file_search_callout → openai_agentic_loop
# → headers → load_balancer(cluster_source=bound_upstream)
# → openai_responses_proxy (when selected upstream = Responses)
# → responses_to_chat_completions (when selected upstream = Chat)
# → path_rewrite (when selected upstream = Chat)
# Response phase (reverse):
# path_rewrite → responses_to_chat_completions → openai_responses_proxy
# → load_balancer → headers → openai_agentic_loop
# → (dispatchers have no response phase) → openai_stream_events
# The protocol-adaptive body builders run AFTER the load balancer because
# they carry selected_upstream conditions, which are rejected
# at build time unless an unconditional load balancer is guaranteed to run
# before them. On the response path they therefore run FIRST, converting a
# Chat Completions response into Responses before openai_agentic_loop and
# openai_stream_events observe it. Exactly one body builder fires per
# request, keyed to the selected cluster's declared application_protocol;
# the match fails closed when the cluster declares none.
# The three dispatchers (openai_web_search, openai_mcp_dispatch,
# openai_file_search_callout) have no response phase: each executes the
# loop owner's assigned calls at request-body EOS on the next iteration,
# reconciling results in place inside ResponsesState.accumulated_output.
# Each is inert unless the model emits the matching hosted tool call.
#
# Transition rules (evaluated after each step's response-body):
# openai_agentic_loop.action = "loop" → next: inference
# default → done (exit to client)
# The loop owner is the sole transition authority (issue #1046); the
# dispatch filters never drive the IRR transition.
#
# Config knobs:
# max_infer_iters: application-level iteration cap on openai_agentic_loop
# max_iterations: infrastructure-level safety cap on the IRR. It must be
# at least max_infer_iters + 1 (initial inference plus
# the allowed tool-backed inference continuations). Here
# 7 + 1 = 8.
#
# Streaming persistence:
# Removing openai_stream_events (or moving openai_response_store after
# the IRR) silently breaks streaming persistence: the store persists
# ResponsesState.response_object, and only openai_stream_events
# populates it for `stream: true` responses. openai_stream_events is
# also REQUIRED for loop correctness — openai_responses_proxy selects
# typed streaming automatically for `stream: true`, which commits
# `response.completed` to the client as it arrives, so a loop-terminal
# error detected by openai_agentic_loop can only reach the client
# through this filter's logical-stream finalizer. When it is absent from
# a streaming step, openai_agentic_loop fails closed with a 500 before
# any backend request. Keep the store before the IRR and openai_stream_events
# INSIDE the IRR step so both buffered and streaming responses are
# persisted and retrievable via GET /v1/responses/{id} (served before the IRR
# by openai_response_store).
#
# Protocol-adaptive translation:
# This config serves ONE Responses gateway pipeline regardless of whether
# the inference backend speaks the Responses API natively or only Chat
# Completions. The IRR step's load balancer publishes the selected
# cluster's typed application metadata (http.application_protocol /
# http.application_provider), and three filters carry selected_upstream
# conditions that read it:
# - openai_responses_proxy when application_protocol=openai_responses
# - responses_to_chat_completions when application_protocol=openai_chat_completions
# - path_rewrite when application_protocol=openai_chat_completions
# The load balancer declares both supported protocol variants because core
# validates selected_upstream matcher values against the clusters that can be
# selected. The top-level binding router maps `vllm-chat` to
# chat-completions-backend and other managed models to inference-backend;
# deployments can replace that illustrative model policy. selected_upstream fails
# closed: declare application_protocol on every selectable cluster, or neither
# branch fires and no outbound body is built. The gated filters must sit AFTER
# the unconditional load balancer, so build-time validation can prove a load
# balancer always runs before them.
#
# Security: StreamBuffer body callouts run before this listener's
# header-phase filters. Deploy this example behind an outer authentication
# and authorization boundary before enabling the required
# allow_pre_security_callout acknowledgement below.
#
# A request is stateful when any of these hold:
# - previous_response_id is set
# - tools array is non-empty
# - store is true (the OpenAI spec default when omitted)
# - background is true
# - conversation is set
# - prompt.id is set
#
# Example requests:
#
# # Stateful (store defaults to true) — validated, looped, persisted
# curl -X POST http://localhost:8080/v1/responses \
# -H "Content-Type: application/json" \
# -d '{"model":"gpt-4.1","input":"Hello, world!"}'
#
# # Streaming with persistence (store defaults to true)
# curl -N -X POST http://localhost:8080/v1/responses \
# -H "Content-Type: application/json" \
# -d '{"model":"gpt-4.1","input":"Hello","stream":true}'
#
# # Stateless (store=false, no stateful markers) — validated and routed
# curl -X POST http://localhost:8080/v1/responses \
# -H "Content-Type: application/json" \
# -d '{"model":"gpt-4.1","input":"Hello","store":false}'
#
# # Multi-turn with previous_response_id
# curl -X POST http://localhost:8080/v1/responses \
# -H "Content-Type: application/json" \
# -d '{"model":"gpt-4.1","input":"What next?","previous_response_id":"resp_abc"}'
#
# # File search — model issues file_search_call, executed via vector store
# curl -X POST http://localhost:8080/v1/responses \
# -H "Content-Type: application/json" \
# -H "Authorization: Bearer token" \
# -d '{
# "model": "gpt-4.1",
# "input": "What does the README say about deployment?",
# "tools": [{"type": "file_search", "vector_store_ids": ["vs_abc"]}]
# }'
#
# # MCP tool loop — model issues an MCP call, executed via the server URL
# curl -X POST http://localhost:8080/v1/responses \
# -H "Content-Type: application/json" \
# -d '{
# "model": "gpt-4.1",
# "input": "What is the weather in SF?",
# "tools": [{"type": "mcp", "server_label": "weather",
# "server_url": "http://mcp-weather:8001/mcp",
# "allowed_tools": ["get_weather"],
# "require_approval": "never"}]
# }'
#
# # Retrieve a stored response (buffered or streamed)
# curl http://localhost:8080/v1/responses/<response-id>
#
# # Responses WebSocket (requires a WebSocket-capable backend) —
# # takes the bypass branch, never the IRR
# websocat -H="Authorization: Bearer $OPENAI_API_KEY" \
# ws://localhost:8080/v1/responses
#
# Web search requires the trusted boundary to provide x-user-brave-key. The
# WEB_SEARCH_API_KEY field remains required by the provider configuration, so
# set it before starting the proxy; when `user_credential: brave_search` is
# configured, a missing per-user slot fails closed instead of falling back to
# that shared key. openai_web_search stays inert until the model emits a
# web_search_call.
#
# Listener and cluster transport timeouts are intentionally omitted.
# Configured downstream read and upstream idle/read/write timeouts remain
# active after a 101 upgrade and can terminate long-lived WebSockets.
#
# Direct OpenAI upstream variant:
# The example routes model `gpt-5` to openai-responses-backend and tags that
# binding with application_provider=openai. Configure that cluster's endpoint
# as api.openai.com:443 with tls.sni/Host api.openai.com, and either preserve
# the caller's Authorization header or inject an operator-owned credential in
# the direct-openai branch. Do not put the OpenAI cluster in the IRR: the
# provider tag intentionally terminates through direct-openai-dispatch before
# validation, local persistence/rehydration, or agentic iteration.
#
# Build:
# cargo build -p praxis-ai-proxy --features openai-file-resolve-filter,openai-conversations,openai-mcp-tools,store-sqlite
listeners:
- name: ai-gateway
address: "127.0.0.1:8080"
filter_chains: [full-flow-agentic-pipeline]
filter_chains:
- name: full-flow-agentic-pipeline
filters:
- filter: trace_context
# Establish typed correlation before IRR snapshots request extensions.
# Child callouts retain this request/trace identity while Praxis core
# creates a fresh span ID for each outbound hop.
- filter: state_owner
# These source names are deployment-specific. The authentication
# boundary must overwrite them and prevent clients from bypassing it.
mode: trusted_headers
tenant:
header: x-auth-tenant
issuer:
static: urn:example:gateway
subject:
header: x-auth-user
- filter: callout_credentials
# The trusted authentication boundary must delete any client-supplied
# instance and then set exactly one authenticated value on every
# request. This filter establishes typed request context and strips the
# ingress header; it does not authenticate the header's provenance.
credentials:
- slot: brave_search
source_header: x-user-brave-key
- slot: ogx_files
source_header: x-user-ogx-key
- slot: mcp_gateway
source_header: x-user-mcp-key
assertions:
# The same establishing filter keeps assertions in a separate typed
# slot map. The trusted boundary MUST delete any client-supplied
# x-mcp-authorized value and then set the verified, short-lived
# assertion. Praxis forwards it unchanged only to configured
# connector_id destinations; the MCP Gateway decides authorization.
- slot: mcp_gateway
source_header: x-mcp-authorized
# Emit the normalized owner in the contract expected by OGX and by the
# explicitly configured downstream callouts below. Raw x-auth-* inputs
# are stripped and never used directly as destination assertions.
- filter: project_state_owner_headers
tenant_header: x-tenant-id
subject_header: x-user-id
# Publishes the typed operation consumed by openai_conversations.
- filter: openai_operation
- filter: openai_responses_format
on_invalid: continue
headers:
format: x-praxis-ai-format
model: x-praxis-ai-model
stream: x-praxis-ai-stream
mode: x-praxis-responses-mode
# Bind one logical destination before any provider-owned body work. The
# classifier overwrites the reserved x-praxis-ai-model header from the
# parsed body, so clients cannot spoof a different promoted value. This
# example uses `gpt-5` for the direct OpenAI path, `vllm-chat` for a
# managed Chat Completions backend, and any other Responses model for a
# managed native-Responses backend. Deployments may replace this model
# policy with another trusted routing fact.
- filter: router
routes:
- path_prefix: "/v1/prompts"
cluster: "prompts-api"
- path_prefix: "/v1/embeddings"
cluster: "embeddings-api"
- path_prefix: "/v1/files"
cluster: "files-api"
# Conversations requests carry no model, so this route is the
# deployment's explicit ownership policy. The managed default below
# selects the Praxis implementation. Point it at
# openai-responses-backend to pass the API through to OpenAI instead.
- path_prefix: "/v1/conversations"
cluster: "inference-backend"
- path: "/v1/responses"
headers:
upgrade: "websocket"
cluster: "responses-websocket-backend"
- path: "/v1/responses"
headers:
x-praxis-ai-format: "openai_responses"
x-praxis-ai-model: "gpt-5"
cluster: "openai-responses-backend"
- path: "/v1/responses"
headers:
x-praxis-ai-format: "openai_responses"
x-praxis-ai-model: "vllm-chat"
cluster: "chat-completions-backend"
- path_prefix: "/v1/responses"
headers:
x-praxis-ai-format: "openai_responses"
cluster: "inference-backend"
- path_prefix: "/v1/vector_stores"
cluster: "vector-stores-backend"
# Core runs the bound-body barrier immediately after the router and
# before evaluating this branch. Each gateway-owned body filter below
# therefore still needs its `unless application_provider: openai` gate.
# This terminal branch then dispatches the original request directly and
# prevents the remaining header chain, including IRR, from running. Both
# layers are necessary. Endpoint selection consumes the frozen logical
# binding rather than making a second routing decision.
#
# `headers` is used here only as a branch anchor; it does not modify any
# headers. Its condition makes the entry run only when the router bound
# the request to a cluster whose application_provider is `openai`.
- filter: headers
conditions:
- when:
bound_upstream:
application_provider: openai
# Branches are evaluated only when their host filter runs. This branch
# has no `on_result`, so it is unconditional once the OpenAI condition
# above matched.
branch_chains:
- name: direct-openai
# `terminal` means: run this branch, dispatch its selected upstream,
# and stop the parent chain. In particular, do not continue to the
# iterative_request_router near the end of this pipeline.
rejoin: terminal
chains:
- name: direct-openai-dispatch
filters:
- filter: load_balancer
# Reuse the cluster already chosen by the router. This must
# not perform another routing decision based on a rewritten
# request later in the pipeline.
cluster_source: bound_upstream
clusters:
# This name must match the router-selected cluster. The
# load balancer resolves that logical binding to a concrete
# endpoint and forwards the original request there.
- name: "openai-responses-backend"
http:
application_protocol: openai_responses
application_provider: openai
endpoints:
- "127.0.0.1:3001"
# Non-agentic service and WebSocket routes retain the historical bypass,
# now driven by the same binding router rather than a router nested in a
# branch (binding pipelines permit exactly one top-level router).
- filter: headers
conditions:
- when:
bound_upstream:
application_provider: direct
branch_chains:
- name: direct-services
rejoin: terminal
chains:
- name: direct-services-dispatch
filters:
- filter: load_balancer
cluster_source: bound_upstream
clusters:
- name: "prompts-api"
http:
application_provider: direct
endpoints:
- "127.0.0.1:9998"
- name: "embeddings-api"
http:
application_provider: direct
endpoints:
- "127.0.0.1:9997"
- name: "files-api"
http:
application_provider: direct
endpoints:
- "127.0.0.1:9999"
- name: "responses-websocket-backend"
http:
application_provider: direct
endpoints:
- "127.0.0.1:3001"
- name: "vector-stores-backend"
http:
application_provider: direct
endpoints:
- "127.0.0.1:3002"
# The filters below participate in the once-per-request bound-body phase.
# OpenAI owns these capabilities on its direct path, so the condition
# suppresses their request, response, and response-body lifecycle hooks.
- filter: openai_responses_validate
conditions:
- unless:
bound_upstream:
application_provider: openai
- filter: openai_tool_parse
conditions:
- unless:
bound_upstream:
application_provider: openai
- filter: openai_conversations
backend: sqlite
database_url: "sqlite://responses.db?mode=rwc"
conversations_table: openai_conversations
items_table: openai_conversation_items
conditions:
- unless:
bound_upstream:
application_provider: openai
- filter: openai_response_store
backend: sqlite
# In-memory:
# database_url: "sqlite::memory:"
# File-backed:
database_url: "sqlite://responses.db?mode=rwc"
responses_table: openai_responses
conversations_table: openai_conversations
conditions:
- unless:
bound_upstream:
application_provider: openai
- filter: openai_responses_rehydrate
conditions:
- unless:
bound_upstream:
application_provider: openai
- filter: openai_file_resolve
files_api_url: "http://127.0.0.1:9999"
# Require the caller-scoped OGX credential captured above. It is
# injected only on configured file_id callouts; file_url downloads
# remain on the separate credential-free resolver.
user_credential: ogx_files
allow_pre_security_callout: true
# file_id callouts run through this outbound filter chain via the
# filtered subrequest executor, which enforces destination authority,
# DNS/SSRF (from the pipeline's allow_private_upstreams), TLS/SNI, and
# Host centrally — this holds even with an empty or omitted chain. The
# chain only re-projects the trusted tenant/subject identity from the
# StateOwner extension; client file_url downloads never traverse it.
outbound_chain:
name: files-api-outbound
filters:
- filter: project_state_owner_headers
tenant_header: x-tenant-id
subject_header: x-user-id
on_missing: reject
timeout_ms: 10000
conditions:
- unless:
bound_upstream:
application_provider: openai
- filter: openai_doc_extract
allow_pre_security_callout: true
on_unsupported: continue
conditions:
- unless:
bound_upstream:
application_provider: openai
- filter: openai_mcp_tool_resolve
user_credential: mcp_gateway
authorization_assertion: mcp_gateway
forward_headers:
- x-tenant-id
- x-user-id
# Ambient trusted headers are forwarded only for connector-backed MCP
# entries; arbitrary client-supplied server_url targets never receive
# them. Clients reference this destination with connector_id.
connectors:
- id: trusted-mcp
server_url: http://127.0.0.1:8001/mcp
conditions:
- unless:
bound_upstream:
application_provider: openai
# The router needs a /v1/responses prefix route so managed GET/DELETE
# resource endpoints can reach the local store. That prefix also binds
# unknown subpaths. After recognized store operations have terminated,
# reject any remaining managed subpath here so it cannot enter the IRR.
# Provider-owned OpenAI traffic already left through direct-openai.
- filter: static_response
status: 404
body: '{"error":{"message":"unsupported managed Responses operation","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}'
headers:
- name: Content-Type
value: application/json
conditions:
- when:
path_prefix: "/v1/responses"
- unless:
path: "/v1/responses"
# Gateway-owned Responses requests flow here after both terminal direct
# branches have been considered. OpenAI-bound requests skip the IRR
# structurally because direct-openai rejoins at terminal above. Do not add
# a bound_upstream condition to IRR itself: IRR owns an ordinary pre-read
# body hook, which cannot be gated on metadata published later by the
# binding router. The IRR runs the agentic model-tool-model loop: each round
# streams through openai_stream_events, the three request-phase
# dispatchers execute the loop owner's assigned hosted-tool calls, and
# openai_agentic_loop parses each response and publishes the single
# loop/done transition. The bound-body openai_response_store persists the
# composed object on the response path (reverse order). The inference
# step consumes the frozen logical binding directly; it must not route
# again or attempt to replace the request-level provider decision.
- filter: iterative_request_router
initial_step: inference
# Safety cap: at least max_infer_iters + 1 (7 + 1 = 8).
max_iterations: 8
# The IRR's 30s default is too short for production model streams.
# Allow six minutes total for time-to-first-byte and logical
# streaming across the loop; each step inherits a five-minute
# deadline (the vLLM integration tests already run at 300s).
timeout_ms: 360000
step_timeout_ms: 300000
max_response_bytes: 67108864
max_stream_response_bytes: 67108864
max_state_bytes: 136314880
steps:
- name: inference
filters:
# Re-project inside the destination-owned subrequest. Core moves
# the normalized StateOwner extension into each IRR step; this
# strips any still-uncommitted raw ingress identity headers and
# recreates only the OGX/callout contract.
- filter: project_state_owner_headers
tenant_header: x-tenant-id
subject_header: x-user-id
# Composes every inference SSE stream into one logical
# Responses lifecycle: preserves one response identity across
# rounds and withholds per-round terminal events until the IRR
# transition is known. A single inference round is just a
# one-round logical stream, so this filter always composes and
# must run inside this IRR step. REQUIRED for both streaming
# persistence and loop-terminal error delivery (see the
# "Streaming persistence" note above); openai_agentic_loop
# fails closed with a 500 on an effective stream:true request
# when it is absent.
#
# Its limits are set at or above responses_to_chat_completions'
# converter bounds below so no frame the client receives is ever
# rejected here and dropped from the store (the accumulator must
# never be tighter than the converter): max_events (200000) >=
# converter max_stream_events (100000); max_buffer_bytes (64 MiB)
# > converter max_emitted_sse_frame_bytes (48 MiB); timeout_secs
# (600) >= converter stream_timeout_secs (300). These apply on the
# native path too and stay within the IRR's byte/time caps above.
- filter: openai_stream_events
max_buffer_bytes: 67108864
max_events: 200000
timeout_secs: 600
# Request-phase dispatcher: runs first on re-entry to execute
# web_search_call assignments prepared by the loop owner. Inert
# until the model emits a web_search_call.
- filter: openai_web_search
provider: brave
api_key: ${WEB_SEARCH_API_KEY}
# Require the per-user key captured by the outer
# callout_credentials filter. The secret is injected only at
# the exact resolved provider authority.
user_credential: brave_search
max_calls_per_round: 32
# Request-phase dispatcher: runs second on re-entry to execute
# MCP calls prepared and classified by the loop owner. Inert
# until the model emits an MCP tool call.
- filter: openai_mcp_dispatch
user_credential: mcp_gateway
authorization_assertion: mcp_gateway
forward_headers:
- x-tenant-id
- x-user-id
max_calls_per_round: 32
max_parallel_calls: 8
max_result_bytes: 1048576
max_total_result_bytes: 8388608
# Request-phase dispatcher: at request-body EOS on each IRR
# re-entry it executes the file_search_call items the loop owner
# assigned in the prior response, reconciling each in place
# inside ResponsesState.accumulated_output. It never parses the
# response body and never decides whether another round runs,
# so it is inert unless the model emits a hosted file_search_call.
- filter: openai_file_search_callout
vector_store_url: http://127.0.0.1:3002
user_credential: ogx_files
# Every vector-store sub-request runs through the filtered
# subrequest executor, which enforces destination authority,
# DNS/SSRF, TLS/SNI, and Host centrally (destination from
# vector_store_url) — this holds even with an empty or omitted
# chain. The outbound chain only re-projects the trusted
# tenant/subject identity from the StateOwner extension.
outbound_chain:
name: vector-store-outbound
filters:
- filter: project_state_owner_headers
tenant_header: x-tenant-id
subject_header: x-user-id
timeout_ms: 5000
max_response_bytes: 10485760
max_total_response_bytes: 67108864
max_state_bytes: 136314880
on_failure: closed
# Sole loop owner: parses each model response, records the
# hosted-tool assignments for the dispatchers, and publishes the
# single continuation signal (action=loop|done). max_infer_iters
# must stay below the IRR max_iterations safety cap (7 + 1 = 8).
- filter: openai_agentic_loop
max_infer_iters: 7
# Protocol adapters run after the load balancer publishes the
# selected cluster's typed application metadata.
- filter: headers
request_set:
- name: Content-Type
value: application/json
# Select an endpoint from the request-level binding and publish
# exchange-local application metadata. Every selected_upstream
# condition below reads this fresh selection, so the matching
# protocol adapter runs independently on every IRR round.
- filter: load_balancer
cluster_source: bound_upstream
clusters:
- name: "inference-backend"
# Gateway-owned native Responses backend.
http:
application_protocol: openai_responses
application_provider: vllm
endpoints:
- "127.0.0.1:3001"
- name: "chat-completions-backend"
# Gateway-owned Chat Completions backend.
http:
application_protocol: openai_chat_completions
application_provider: vllm
endpoints:
- "127.0.0.1:3001"
# ---- Protocol-adaptive body builders (run AFTER selection) ----
# Exactly one branch fires, chosen by the selected upstream's
# declared application_protocol. selected_upstream fails closed:
# always declare http.application_protocol on the cluster above.
# Native Responses upstream: enforce provider-aware prompt
# support, rebuild the outbound body from ResponsesState, and
# select the typed streaming transport for `stream: true`.
- filter: openai_responses_proxy
conditions:
- when:
selected_upstream:
application_protocol: openai_responses
# Chat Completions upstream: translate the canonical Responses
# request into Chat Completions on the request path, and
# translate the Chat Completions response (finite JSON or streamed
# SSE) back into a Responses resource / Responses SSE on the
# response path. Because response filters run in reverse, this
# converts the raw Chat response into Responses BEFORE
# openai_agentic_loop and openai_stream_events observe it (both
# require Responses). It fully owns outbound body build for Chat
# backends. It also rejects OpenAI prompt templates, which have no
# Chat representation. Stream bounds stay at/under
# openai_stream_events'
# limits so no frame the client receives is dropped from the store
# (cf. responses-to-chat-completions.yaml).
- filter: responses_to_chat_completions
max_rewritten_body_bytes: 33554432
max_emitted_sse_frame_bytes: 50331648
max_stream_events: 100000
stream_timeout_secs: 300
conditions:
- when:
selected_upstream:
application_protocol: openai_chat_completions
# The translator owns body/response conversion but deliberately
# leaves the request URI unchanged. A Chat Completions backend
# still expects POST /v1/chat/completions, so rewrite that transport
# detail separately. The identical protocol gate plus create
# verb/path guards ensure only translated requests are rewritten.
- filter: path_rewrite
replace:
pattern: "^/v1/responses/?$"
replacement: "/v1/chat/completions"
conditions:
- when:
selected_upstream:
application_protocol: openai_chat_completions
path_prefix: "/v1/responses"
methods: [POST]
on_result:
# openai_agentic_loop is the sole loop authority (issue #1046):
# it parses each model response, records assignments for the
# request-phase dispatchers (web_search, mcp_dispatch,
# file_search_callout), and publishes the single continuation
# signal. The IRR transitions only on the owner's action, never
# on a dispatcher's.
- filter: openai_agentic_loop
key: action
value: loop
next: inference
- default: true
done: true
insecure_options:
allow_private_endpoints: true # example proxies to local backends
# Central SSRF control: permits resolved callout addresses (vector_store_url,
# file_id Files API) to be private/loopback at connect time (needed for
# local/private vector stores and Files API backends).
allow_private_upstreams: true