File Search Streaming
Demonstrates streaming hosted file_search through the iterative_request_router
Category: Setup-dependent integration
Task: Demonstrates streaming hosted file_search through the iterative_request_router
Prerequisites: The external service, credentials, or certificates referenced by this configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
# File Search Callout Filter — Streaming (stream: true)
# Requires `--features openai-responses` because these filters are opt-in.
#
# Demonstrates streaming hosted file_search through the iterative_request_router.
# The IRR sends the request to the model, executes any returned file_search_call
# through the vector store, and streams the enriched context back through a second
# model inference before finalizing the SSE lifecycle with citation annotations.
#
# This config uses the streaming step order mandated by §5 of the #313 spec:
# `openai_stream_events` is REQUIRED ahead of `openai_file_search_callout` and
# `openai_responses_proxy`. It always composes the current IRR execution into
# one logical Responses stream. The proxy derives
# the streaming transport automatically from the effective outbound `stream:
# true`; there is no operator opt-in. Typed streaming commits response.completed
# to the client as it arrives, so a loop-terminal error detected after that event
# must reach the client through the logical-stream finalizer, which defers each
# round's terminal event and can replace it with an error frame. Without the
# stream-events finalizer that error is silently dropped after the success is
# already on the wire.
#
# openai_agentic_loop is the sole loop owner: it parses each streamed
# model round, records file-search assignments, and publishes the single
# continuation signal (action=loop|done). openai_file_search_callout is a
# request-phase dispatcher that executes those assignments at request-body
# EOS on IRR re-entry (#1046). openai_stream_events composes every round
# into one logical Responses stream.
#
# Configuration:
# vector_store_url: Vector store API base URL. Structural validation at
# startup rejects a non-http(s) scheme, userinfo, or a
# query/fragment. Private/loopback targets (literal or
# resolved) are gated at connect time and permitted
# only when insecure_options.allow_private_upstreams
# is set.
# outbound_chain: Filter chain every vector-store sub-request runs
# through (cross-cutting concerns such as credential
# injection). Must be inline ({ name, filters }); a
# named filter_chains reference cannot resolve inside
# the iterative_request_router step this filter runs
# in. The destination comes from vector_store_url, so
# the chain never selects an upstream.
# timeout_ms: Whole-call timeout in milliseconds
# max_response_bytes: Maximum response bytes retained per callout
# max_total_response_bytes: Maximum successful bytes across one fan-out
# on_failure: closed (fail closed) or open (fail open)
# forward_headers: Headers to forward from the client request
# (e.g. Authorization) to the vector store API
#
# Shared sub-request connector settings under `runtime` are restart-required.
# When `runtime.subrequest_circuit_breaker` is set, client-aware filters such
# as this one inherit peer-level failure isolation from the shared client:
#
# runtime:
# subrequest_circuit_breaker:
# consecutive_failures: 5
# recovery_window_secs: 30
listeners:
- name: ai-gateway
address: "127.0.0.1:8080"
filter_chains: [file-search-streaming-pipeline]
filter_chains:
- name: file-search-streaming-pipeline
filters:
- filter: openai_responses_request
- filter: openai_tool_parse
- filter: iterative_request_router
initial_step: inference
max_iterations: 8
timeout_ms: 120000
step_timeout_ms: 60000
max_stream_response_bytes: 67108864
# Retains one 64 MiB request plus one 64 MiB response and metadata.
max_state_bytes: 136314880
steps:
- name: inference
filters:
- filter: openai_stream_events
# Request-phase dispatcher: at request-body EOS on each IRR
# re-entry it executes the file_search_call items the loop owner
# assigned in the prior streamed round, reconciling each in place
# inside ResponsesState.accumulated_output. It never parses the
# response body and never decides whether another round runs.
- filter: openai_file_search_callout
vector_store_url: http://127.0.0.1:8001
# Every vector-store sub-request runs through this chain. It must
# be inline ({ name, filters }): this filter runs inside an
# iterative_request_router step, whose pipeline is built with an
# empty named-chain map, so a top-level filter_chains reference
# cannot resolve here. The destination comes from
# vector_store_url, so the chain never selects an upstream.
outbound_chain:
name: vector-store-outbound
filters:
- filter: headers
request_set:
- name: X-Vector-Store-Client
value: praxis-ai-gateway
timeout_ms: 5000
max_response_bytes: 10485760
max_total_response_bytes: 67108864
# The filter and the enclosing iterative router may use
# different values; the smaller limit wins at runtime.
max_state_bytes: 136314880
on_failure: closed
forward_headers:
- authorization
# Sole loop owner: parses each streamed model round, records
# file-search assignments for the dispatcher, and publishes the
# single continuation signal (action=loop|done). max_infer_iters
# must stay below the IRR max_iterations safety cap (7 + 1 = 8).
- filter: openai_agentic_loop
max_infer_iters: 7
- filter: openai_responses_proxy
- filter: headers
request_set:
- name: Content-Type
value: application/json
- filter: router
routes:
- path: "/v1/responses"
cluster: "inference"
- filter: load_balancer
clusters:
- name: "inference"
endpoints:
- "127.0.0.1:3001"
on_result:
- filter: openai_agentic_loop
key: action
value: loop
next: inference
- default: true
done: true
insecure_options:
allow_private_endpoints: true # example proxies to local backends
# Central SSRF control for outbound callouts: permits the vector_store_url's
# resolved address to be private/loopback at connect time. Required only when
# the vector store runs on a private network (e.g. local development).
allow_private_upstreams: true