File Search Streaming

Demonstrates streaming hosted file_search through the iterative_request_router

Category: Setup-dependent integration
Task: Demonstrates streaming hosted file_search through the iterative_request_router

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# File Search Callout Filter — Streaming (stream: true)
# Requires `--features openai-responses` because these filters are opt-in.
#
# Demonstrates streaming hosted file_search through the iterative_request_router.
# The IRR sends the request to the model, executes any returned file_search_call
# through the vector store, and streams the enriched context back through a second
# model inference before finalizing the SSE lifecycle with citation annotations.
#
# This config uses the streaming step order mandated by §5 of the #313 spec:
# `openai_stream_events` is REQUIRED ahead of `openai_file_search_callout` and
# `openai_responses_proxy`. It always composes the current IRR execution into
# one logical Responses stream. The proxy derives
# the streaming transport automatically from the effective outbound `stream:
# true`; there is no operator opt-in. Typed streaming commits response.completed
# to the client as it arrives, so a loop-terminal error detected after that event
# must reach the client through the logical-stream finalizer, which defers each
# round's terminal event and can replace it with an error frame. Without the
# stream-events finalizer that error is silently dropped after the success is
# already on the wire.
#
# openai_agentic_loop is the sole loop owner: it parses each streamed
# model round, records file-search assignments, and publishes the single
# continuation signal (action=loop|done). openai_file_search_callout is a
# request-phase dispatcher that executes those assignments at request-body
# EOS on IRR re-entry (#1046). openai_stream_events composes every round
# into one logical Responses stream.
#
# Configuration:
#   vector_store_url:         Vector store API base URL. Structural validation at
#                             startup rejects a non-http(s) scheme, userinfo, or a
#                             query/fragment. Private/loopback targets (literal or
#                             resolved) are gated at connect time and permitted
#                             only when insecure_options.allow_private_upstreams
#                             is set.
#   outbound_chain:           Filter chain every vector-store sub-request runs
#                             through (cross-cutting concerns such as credential
#                             injection). Must be inline ({ name, filters }); a
#                             named filter_chains reference cannot resolve inside
#                             the iterative_request_router step this filter runs
#                             in. The destination comes from vector_store_url, so
#                             the chain never selects an upstream.
#   timeout_ms:               Whole-call timeout in milliseconds
#   max_response_bytes:       Maximum response bytes retained per callout
#   max_total_response_bytes: Maximum successful bytes across one fan-out
#   on_failure:               closed (fail closed) or open (fail open)
#   forward_headers:          Headers to forward from the client request
#                             (e.g. Authorization) to the vector store API
#
# Shared sub-request connector settings under `runtime` are restart-required.
# When `runtime.subrequest_circuit_breaker` is set, client-aware filters such
# as this one inherit peer-level failure isolation from the shared client:
#
# runtime:
#   subrequest_circuit_breaker:
#     consecutive_failures: 5
#     recovery_window_secs: 30

listeners:
  - name: ai-gateway
    address: "127.0.0.1:8080"
    filter_chains: [file-search-streaming-pipeline]

filter_chains:
  - name: file-search-streaming-pipeline
    filters:
      - filter: openai_responses_request
      - filter: openai_tool_parse
      - filter: iterative_request_router
        initial_step: inference
        max_iterations: 8
        timeout_ms: 120000
        step_timeout_ms: 60000
        max_stream_response_bytes: 67108864
        # Retains one 64 MiB request plus one 64 MiB response and metadata.
        max_state_bytes: 136314880
        steps:
          - name: inference
            filters:
              - filter: openai_stream_events
              # Request-phase dispatcher: at request-body EOS on each IRR
              # re-entry it executes the file_search_call items the loop owner
              # assigned in the prior streamed round, reconciling each in place
              # inside ResponsesState.accumulated_output. It never parses the
              # response body and never decides whether another round runs.
              - filter: openai_file_search_callout
                vector_store_url: http://127.0.0.1:8001
                # Every vector-store sub-request runs through this chain. It must
                # be inline ({ name, filters }): this filter runs inside an
                # iterative_request_router step, whose pipeline is built with an
                # empty named-chain map, so a top-level filter_chains reference
                # cannot resolve here. The destination comes from
                # vector_store_url, so the chain never selects an upstream.
                outbound_chain:
                  name: vector-store-outbound
                  filters:
                    - filter: headers
                      request_set:
                        - name: X-Vector-Store-Client
                          value: praxis-ai-gateway
                timeout_ms: 5000
                max_response_bytes: 10485760
                max_total_response_bytes: 67108864
                # The filter and the enclosing iterative router may use
                # different values; the smaller limit wins at runtime.
                max_state_bytes: 136314880
                on_failure: closed
                forward_headers:
                  - authorization
              # Sole loop owner: parses each streamed model round, records
              # file-search assignments for the dispatcher, and publishes the
              # single continuation signal (action=loop|done). max_infer_iters
              # must stay below the IRR max_iterations safety cap (7 + 1 = 8).
              - filter: openai_agentic_loop
                max_infer_iters: 7
              - filter: openai_responses_proxy
              - filter: headers
                request_set:
                  - name: Content-Type
                    value: application/json
              - filter: router
                routes:
                  - path: "/v1/responses"
                    cluster: "inference"
              - filter: load_balancer
                clusters:
                  - name: "inference"
                    endpoints:
                      - "127.0.0.1:3001"
            on_result:
              - filter: openai_agentic_loop
                key: action
                value: loop
                next: inference
              - default: true
                done: true

insecure_options:
  allow_private_endpoints: true # example proxies to local backends
  # Central SSRF control for outbound callouts: permits the vector_store_url's
  # resolved address to be private/loopback at connect time. Required only when
  # the vector store runs on a private network (e.g. local development).
  allow_private_upstreams: true