Full Flow Agentic

Runs the complete Responses API pipeline through an agentic iterative_request_router that executes hosted file_search, web_search, and MCP tool calls in a model-tool-model loop, persisting both buffered and streaming (stream: true) responses

Category: Setup-dependent integration
Task: Runs the complete Responses API pipeline through an agentic iterative_request_router that executes hosted file_search, web_search, and MCP tool calls in a model-tool-model loop, persisting both buffered and streaming (stream: true) responses

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

Companion resources from the same snapshot:

# Responses API Full Flow — Unified Agentic Gateway
# Requires `--features openai-file-resolve-filter,openai-conversations,openai-mcp-tools,store-sqlite` because these filters are opt-in.
#
# Runs the complete Responses API pipeline through an agentic
# iterative_request_router that executes hosted file_search, web_search,
# and MCP tool calls in a model-tool-model loop, persisting both buffered
# and streaming (`stream: true`) responses. This is the single config an
# operator would deploy for a Responses API gateway that must serve the
# full agentic loop, the Conversations API, and the dedicated
# prompts/embeddings/files/vector_stores services behind one listener. The
# request trace is continued through every nested web-search, OGX, and MCP
# callout with a fresh span ID per hop.
#
# The request is classified once, then a top-level router freezes one logical
# cluster plus its application_provider/application_protocol metadata. That
# binding drives two mutually exclusive paths:
#
#   - application_provider=openai: a terminal branch forwards the original
#     Responses request directly. Managed validation, response state,
#     persistence/rehydration, hosted-tool resolution, and the IRR never run.
#   - every managed provider: the bound-body barrier runs provider-aware body
#     filters once, then the IRR executes the agentic loop.
#
# ResponsesState created on the managed path persists across all IRR
# iterations via extension swapping:
#
#   0. state_owner maps deployment-specific trusted ingress headers into one
#      normalized tenant/issuer/subject context for owner-aware filters.
#      callout_credentials captures per-user Brave and OGX provider keys from
#      trusted x-user-brave-key and x-user-ogx-key ingress headers into the
#      `brave_search` and `ogx_files` slots, then strips both source headers
#      before inference or IRR processing. The trusted boundary MUST
#      unconditionally delete-then-set those source headers on every request;
#      append-on-success is insufficient because this filter cannot authenticate
#      header provenance. The IRR carries the resulting typed request extension
#      across iterations, while each callout explicitly selects its slot for its
#      exact-authority provider destination.
#      project_state_owner_headers then projects that context as x-tenant-id and
#      x-user-id for OGX and explicitly allowlisted callouts. The raw ingress
#      names are stripped, and the projection is repeated inside the IRR step
#      because it owns a separate destination request lifecycle.
#      These adapters do not themselves change persistence predicates: this
#      context-only configuration must be combined with owner-scoped store
#      enforcement before it is exposed as a multi-owner state service.
#
#   0. openai_operation publishes the typed registry match consumed by
#      openai_conversations. It must remain earlier in this filter chain.
#
#   1. openai_conversations handles the Conversations API only for managed
#      providers. The router binds /v1/conversations first: an OpenAI binding
#      takes the direct terminal branch, while a managed binding runs the
#      Praxis CRUD implementation. On the managed Responses response path the
#      filter also appends input+output items back to the conversation after a
#      completed buffered or streamed response. For streams it writes from the
#      canonical terminal state before releasing response.completed, including
#      when the Response itself has store:false.
#
#   2. openai_responses_format classifies the request and promotes
#      format, model, stream, and mode to internal routing headers.
#      on_invalid is left at the default (passthrough) so that
#      Conversations API traffic is not rejected before the conversations
#      filter can handle it in the body phase.
#
#   3. The binding router selects the logical provider/protocol. The filters
#      below that declare bound-body access then run once against the buffered
#      request. Their `unless application_provider=openai` conditions make the
#      OpenAI route true passthrough. openai_responses_validate checks JSON,
#      initializes ResponsesState, and rejects gateway-unsupported fields such
#      as background=true only on managed routes.
#
#   4. On managed routes, openai_tool_parse parses the tools array and
#      tool_choice from Responses API requests and promotes summary facts
#      (has_tools, has_web_search, tool_choice, function_count, etc.) to
#      metadata and filter results. It does not mutate the request body and is
#      skipped with the other gateway-owned body work on direct OpenAI routes.
#
#   5. openai_response_store persists responses to SQLite and registers
#      the store backend so downstream filters (rehydrate, compact) can
#      read from it. It runs once at the post-binding body barrier before
#      the iterative_request_router, so on the response path — which runs in reverse config
#      order — it persists AFTER the IRR step has composed the terminal
#      event. This is what makes streaming persistence work: the store
#      reads the accumulated response object that openai_stream_events
#      builds inside the IRR.
#
#   6. openai_responses_rehydrate loads conversation context from
#      previous_response_id or conversation field by fetching stored
#      responses/conversations and prepending their message history to the
#      current request. previous_response_id takes precedence when both
#      are present.
#
#   7. openai_file_resolve resolves file_id references in the current
#      input and rehydrated history through a configured Files API. Runs
#      after rehydrate so stored history is available, and before
#      responses_proxy so the rebuilt upstream body contains file_data or
#      image_url rather than stale file_id references.
#
#   8. openai_doc_extract converts input_file content parts to input_text
#      for backends that do not natively support input_file (e.g. vLLM,
#      llm-d). Text-safe content (text/*, application/json,
#      application/xml) is decoded from base64, validated as UTF-8, and
#      forwarded as plain text. Unsupported formats are left unchanged
#      (on_unsupported: continue). This filter is optional — omit it when
#      the inference backend natively supports input_file.
#
#   9. openai_mcp_tool_resolve resolves MCP tool entries from the tools
#      array by calling tools/list on each upstream MCP server. Runs after
#      rehydrate so previous_tools is available for cross-request caching.
#      Writes mcp_tool_map to ResponsesState for downstream tool dispatch.
#      Projected identity headers are forwarded only to operator-configured
#      connector_id targets, never arbitrary client-selected server_url values.
#
#   10. Two terminal branches consume the frozen binding before the IRR.
#       One forwards provider-owned OpenAI Responses traffic unchanged; the
#       other preserves the direct service/WebSocket routes. Both use
#       `cluster_source: bound_upstream`, so no second routing decision can
#       disagree with the provider selected for the request.
#
#   11. The iterative_request_router owns the agentic loop for classified
#      openai_responses create requests. For a native Responses upstream,
#      openai_responses_proxy rebuilds the outbound body from ResponsesState
#      and automatically selects the typed streaming transport for an
#      effective `stream: true` request, so each backend SSE chunk reaches
#      openai_stream_events incrementally; for a Chat Completions upstream,
#      responses_to_chat_completions does the equivalent build and translates
#      Chat responses back to Responses. Both now run AFTER the step load
#      balancer, each gated on the selected upstream's protocol (see the
#      "Protocol-adaptive translation" note below).
#      The step load balancer consumes the request-level binding directly;
#      there is no router inside the IRR. openai_agentic_loop is the sole loop owner, sole
#      response parser, and sole transition authority: it extracts
#      completed function_call, web_search_call, file_search_call, and MCP
#      items, records the assignment each request-phase dispatcher must
#      execute next, and publishes the single loop/done decision.
#
# IRR inference filter chain (canonical #1046 order with
# protocol-adaptive translation):
#   Request phase (forward):
#     openai_stream_events → openai_web_search → openai_mcp_dispatch
#       → openai_file_search_callout → openai_agentic_loop
#       → headers → load_balancer(cluster_source=bound_upstream)
#       → openai_responses_proxy         (when selected upstream = Responses)
#       → responses_to_chat_completions  (when selected upstream = Chat)
#       → path_rewrite                   (when selected upstream = Chat)
#   Response phase (reverse):
#     path_rewrite → responses_to_chat_completions → openai_responses_proxy
#       → load_balancer → headers → openai_agentic_loop
#       → (dispatchers have no response phase) → openai_stream_events
#   The protocol-adaptive body builders run AFTER the load balancer because
#   they carry selected_upstream conditions, which are rejected
#   at build time unless an unconditional load balancer is guaranteed to run
#   before them. On the response path they therefore run FIRST, converting a
#   Chat Completions response into Responses before openai_agentic_loop and
#   openai_stream_events observe it. Exactly one body builder fires per
#   request, keyed to the selected cluster's declared application_protocol;
#   the match fails closed when the cluster declares none.
#   The three dispatchers (openai_web_search, openai_mcp_dispatch,
#   openai_file_search_callout) have no response phase: each executes the
#   loop owner's assigned calls at request-body EOS on the next iteration,
#   reconciling results in place inside ResponsesState.accumulated_output.
#   Each is inert unless the model emits the matching hosted tool call.
#
# Transition rules (evaluated after each step's response-body):
#   openai_agentic_loop.action = "loop" → next: inference
#   default                             → done (exit to client)
#   The loop owner is the sole transition authority (issue #1046); the
#   dispatch filters never drive the IRR transition.
#
# Config knobs:
#   max_infer_iters: application-level iteration cap on openai_agentic_loop
#   max_iterations:  infrastructure-level safety cap on the IRR. It must be
#                    at least max_infer_iters + 1 (initial inference plus
#                    the allowed tool-backed inference continuations). Here
#                    7 + 1 = 8.
#
# Streaming persistence:
#   Removing openai_stream_events (or moving openai_response_store after
#   the IRR) silently breaks streaming persistence: the store persists
#   ResponsesState.response_object, and only openai_stream_events
#   populates it for `stream: true` responses. openai_stream_events is
#   also REQUIRED for loop correctness — openai_responses_proxy selects
#   typed streaming automatically for `stream: true`, which commits
#   `response.completed` to the client as it arrives, so a loop-terminal
#   error detected by openai_agentic_loop can only reach the client
#   through this filter's logical-stream finalizer. When it is absent from
#   a streaming step, openai_agentic_loop fails closed with a 500 before
#   any backend request. Keep the store before the IRR and openai_stream_events
#   INSIDE the IRR step so both buffered and streaming responses are
#   persisted and retrievable via GET /v1/responses/{id} (served before the IRR
#   by openai_response_store).
#
# Protocol-adaptive translation:
#   This config serves ONE Responses gateway pipeline regardless of whether
#   the inference backend speaks the Responses API natively or only Chat
#   Completions. The IRR step's load balancer publishes the selected
#   cluster's typed application metadata (http.application_protocol /
#   http.application_provider), and three filters carry selected_upstream
#   conditions that read it:
#     - openai_responses_proxy         when application_protocol=openai_responses
#     - responses_to_chat_completions  when application_protocol=openai_chat_completions
#     - path_rewrite                   when application_protocol=openai_chat_completions
#   The load balancer declares both supported protocol variants because core
#   validates selected_upstream matcher values against the clusters that can be
#   selected. The top-level binding router maps `vllm-chat` to
#   chat-completions-backend and other managed models to inference-backend;
#   deployments can replace that illustrative model policy. selected_upstream fails
#   closed: declare application_protocol on every selectable cluster, or neither
#   branch fires and no outbound body is built. The gated filters must sit AFTER
#   the unconditional load balancer, so build-time validation can prove a load
#   balancer always runs before them.
#
# Security: StreamBuffer body callouts run before this listener's
# header-phase filters. Deploy this example behind an outer authentication
# and authorization boundary before enabling the required
# allow_pre_security_callout acknowledgement below.
#
# A request is stateful when any of these hold:
#   - previous_response_id is set
#   - tools array is non-empty
#   - store is true (the OpenAI spec default when omitted)
#   - background is true
#   - conversation is set
#   - prompt.id is set
#
# Example requests:
#
#   # Stateful (store defaults to true) — validated, looped, persisted
#   curl -X POST http://localhost:8080/v1/responses \
#     -H "Content-Type: application/json" \
#     -d '{"model":"gpt-4.1","input":"Hello, world!"}'
#
#   # Streaming with persistence (store defaults to true)
#   curl -N -X POST http://localhost:8080/v1/responses \
#     -H "Content-Type: application/json" \
#     -d '{"model":"gpt-4.1","input":"Hello","stream":true}'
#
#   # Stateless (store=false, no stateful markers) — validated and routed
#   curl -X POST http://localhost:8080/v1/responses \
#     -H "Content-Type: application/json" \
#     -d '{"model":"gpt-4.1","input":"Hello","store":false}'
#
#   # Multi-turn with previous_response_id
#   curl -X POST http://localhost:8080/v1/responses \
#     -H "Content-Type: application/json" \
#     -d '{"model":"gpt-4.1","input":"What next?","previous_response_id":"resp_abc"}'
#
#   # File search — model issues file_search_call, executed via vector store
#   curl -X POST http://localhost:8080/v1/responses \
#     -H "Content-Type: application/json" \
#     -H "Authorization: Bearer token" \
#     -d '{
#       "model": "gpt-4.1",
#       "input": "What does the README say about deployment?",
#       "tools": [{"type": "file_search", "vector_store_ids": ["vs_abc"]}]
#     }'
#
#   # MCP tool loop — model issues an MCP call, executed via the server URL
#   curl -X POST http://localhost:8080/v1/responses \
#     -H "Content-Type: application/json" \
#     -d '{
#       "model": "gpt-4.1",
#       "input": "What is the weather in SF?",
#       "tools": [{"type": "mcp", "server_label": "weather",
#                  "server_url": "http://mcp-weather:8001/mcp",
#                  "allowed_tools": ["get_weather"],
#                  "require_approval": "never"}]
#     }'
#
#   # Retrieve a stored response (buffered or streamed)
#   curl http://localhost:8080/v1/responses/<response-id>
#
#   # Responses WebSocket (requires a WebSocket-capable backend) —
#   # takes the bypass branch, never the IRR
#   websocat -H="Authorization: Bearer $OPENAI_API_KEY" \
#     ws://localhost:8080/v1/responses
#
# Web search requires the trusted boundary to provide x-user-brave-key. The
# WEB_SEARCH_API_KEY field remains required by the provider configuration, so
# set it before starting the proxy; when `user_credential: brave_search` is
# configured, a missing per-user slot fails closed instead of falling back to
# that shared key. openai_web_search stays inert until the model emits a
# web_search_call.
#
# Listener and cluster transport timeouts are intentionally omitted.
# Configured downstream read and upstream idle/read/write timeouts remain
# active after a 101 upgrade and can terminate long-lived WebSockets.
#
# Direct OpenAI upstream variant:
#   The example routes model `gpt-5` to openai-responses-backend and tags that
#   binding with application_provider=openai. Configure that cluster's endpoint
#   as api.openai.com:443 with tls.sni/Host api.openai.com, and either preserve
#   the caller's Authorization header or inject an operator-owned credential in
#   the direct-openai branch. Do not put the OpenAI cluster in the IRR: the
#   provider tag intentionally terminates through direct-openai-dispatch before
#   validation, local persistence/rehydration, or agentic iteration.
#
# Build:
#   cargo build -p praxis-ai-proxy --features openai-file-resolve-filter,openai-conversations,openai-mcp-tools,store-sqlite

listeners:
  - name: ai-gateway
    address: "127.0.0.1:8080"
    filter_chains: [full-flow-agentic-pipeline]

filter_chains:
  - name: full-flow-agentic-pipeline
    filters:
      - filter: trace_context
        # Establish typed correlation before IRR snapshots request extensions.
        # Child callouts retain this request/trace identity while Praxis core
        # creates a fresh span ID for each outbound hop.

      - filter: state_owner
        # These source names are deployment-specific. The authentication
        # boundary must overwrite them and prevent clients from bypassing it.
        mode: trusted_headers
        tenant:
          header: x-auth-tenant
        issuer:
          static: urn:example:gateway
        subject:
          header: x-auth-user

      - filter: callout_credentials
        # The trusted authentication boundary must delete any client-supplied
        # instance and then set exactly one authenticated value on every
        # request. This filter establishes typed request context and strips the
        # ingress header; it does not authenticate the header's provenance.
        credentials:
          - slot: brave_search
            source_header: x-user-brave-key
          - slot: ogx_files
            source_header: x-user-ogx-key
          - slot: mcp_gateway
            source_header: x-user-mcp-key
        assertions:
          # The same establishing filter keeps assertions in a separate typed
          # slot map. The trusted boundary MUST delete any client-supplied
          # x-mcp-authorized value and then set the verified, short-lived
          # assertion. Praxis forwards it unchanged only to configured
          # connector_id destinations; the MCP Gateway decides authorization.
          - slot: mcp_gateway
            source_header: x-mcp-authorized

      # Emit the normalized owner in the contract expected by OGX and by the
      # explicitly configured downstream callouts below. Raw x-auth-* inputs
      # are stripped and never used directly as destination assertions.
      - filter: project_state_owner_headers
        tenant_header: x-tenant-id
        subject_header: x-user-id

      # Publishes the typed operation consumed by openai_conversations.
      - filter: openai_operation

      - filter: openai_responses_format
        on_invalid: continue
        headers:
          format: x-praxis-ai-format
          model: x-praxis-ai-model
          stream: x-praxis-ai-stream
          mode: x-praxis-responses-mode

      # Bind one logical destination before any provider-owned body work. The
      # classifier overwrites the reserved x-praxis-ai-model header from the
      # parsed body, so clients cannot spoof a different promoted value. This
      # example uses `gpt-5` for the direct OpenAI path, `vllm-chat` for a
      # managed Chat Completions backend, and any other Responses model for a
      # managed native-Responses backend. Deployments may replace this model
      # policy with another trusted routing fact.
      - filter: router
        routes:
          - path_prefix: "/v1/prompts"
            cluster: "prompts-api"
          - path_prefix: "/v1/embeddings"
            cluster: "embeddings-api"
          - path_prefix: "/v1/files"
            cluster: "files-api"
          # Conversations requests carry no model, so this route is the
          # deployment's explicit ownership policy. The managed default below
          # selects the Praxis implementation. Point it at
          # openai-responses-backend to pass the API through to OpenAI instead.
          - path_prefix: "/v1/conversations"
            cluster: "inference-backend"
          - path: "/v1/responses"
            headers:
              upgrade: "websocket"
            cluster: "responses-websocket-backend"
          - path: "/v1/responses"
            headers:
              x-praxis-ai-format: "openai_responses"
              x-praxis-ai-model: "gpt-5"
            cluster: "openai-responses-backend"
          - path: "/v1/responses"
            headers:
              x-praxis-ai-format: "openai_responses"
              x-praxis-ai-model: "vllm-chat"
            cluster: "chat-completions-backend"
          - path_prefix: "/v1/responses"
            headers:
              x-praxis-ai-format: "openai_responses"
            cluster: "inference-backend"
          - path_prefix: "/v1/vector_stores"
            cluster: "vector-stores-backend"

      # Core runs the bound-body barrier immediately after the router and
      # before evaluating this branch. Each gateway-owned body filter below
      # therefore still needs its `unless application_provider: openai` gate.
      # This terminal branch then dispatches the original request directly and
      # prevents the remaining header chain, including IRR, from running. Both
      # layers are necessary. Endpoint selection consumes the frozen logical
      # binding rather than making a second routing decision.
      #
      # `headers` is used here only as a branch anchor; it does not modify any
      # headers. Its condition makes the entry run only when the router bound
      # the request to a cluster whose application_provider is `openai`.
      - filter: headers
        conditions:
          - when:
              bound_upstream:
                application_provider: openai
        # Branches are evaluated only when their host filter runs. This branch
        # has no `on_result`, so it is unconditional once the OpenAI condition
        # above matched.
        branch_chains:
          - name: direct-openai
            # `terminal` means: run this branch, dispatch its selected upstream,
            # and stop the parent chain. In particular, do not continue to the
            # iterative_request_router near the end of this pipeline.
            rejoin: terminal
            chains:
              - name: direct-openai-dispatch
                filters:
                  - filter: load_balancer
                    # Reuse the cluster already chosen by the router. This must
                    # not perform another routing decision based on a rewritten
                    # request later in the pipeline.
                    cluster_source: bound_upstream
                    clusters:
                      # This name must match the router-selected cluster. The
                      # load balancer resolves that logical binding to a concrete
                      # endpoint and forwards the original request there.
                      - name: "openai-responses-backend"
                        http:
                          application_protocol: openai_responses
                          application_provider: openai
                        endpoints:
                          - "127.0.0.1:3001"

      # Non-agentic service and WebSocket routes retain the historical bypass,
      # now driven by the same binding router rather than a router nested in a
      # branch (binding pipelines permit exactly one top-level router).
      - filter: headers
        conditions:
          - when:
              bound_upstream:
                application_provider: direct
        branch_chains:
          - name: direct-services
            rejoin: terminal
            chains:
              - name: direct-services-dispatch
                filters:
                  - filter: load_balancer
                    cluster_source: bound_upstream
                    clusters:
                      - name: "prompts-api"
                        http:
                          application_provider: direct
                        endpoints:
                          - "127.0.0.1:9998"
                      - name: "embeddings-api"
                        http:
                          application_provider: direct
                        endpoints:
                          - "127.0.0.1:9997"
                      - name: "files-api"
                        http:
                          application_provider: direct
                        endpoints:
                          - "127.0.0.1:9999"
                      - name: "responses-websocket-backend"
                        http:
                          application_provider: direct
                        endpoints:
                          - "127.0.0.1:3001"
                      - name: "vector-stores-backend"
                        http:
                          application_provider: direct
                        endpoints:
                          - "127.0.0.1:3002"

      # The filters below participate in the once-per-request bound-body phase.
      # OpenAI owns these capabilities on its direct path, so the condition
      # suppresses their request, response, and response-body lifecycle hooks.
      - filter: openai_responses_validate
        conditions:
          - unless:
              bound_upstream:
                application_provider: openai

      - filter: openai_tool_parse
        conditions:
          - unless:
              bound_upstream:
                application_provider: openai

      - filter: openai_conversations
        backend: sqlite
        database_url: "sqlite://responses.db?mode=rwc"
        conversations_table: openai_conversations
        items_table: openai_conversation_items
        conditions:
          - unless:
              bound_upstream:
                application_provider: openai

      - filter: openai_response_store
        backend: sqlite
        # In-memory:
        #   database_url: "sqlite::memory:"
        # File-backed:
        database_url: "sqlite://responses.db?mode=rwc"
        responses_table: openai_responses
        conversations_table: openai_conversations
        conditions:
          - unless:
              bound_upstream:
                application_provider: openai

      - filter: openai_responses_rehydrate
        conditions:
          - unless:
              bound_upstream:
                application_provider: openai

      - filter: openai_file_resolve
        files_api_url: "http://127.0.0.1:9999"
        # Require the caller-scoped OGX credential captured above. It is
        # injected only on configured file_id callouts; file_url downloads
        # remain on the separate credential-free resolver.
        user_credential: ogx_files
        allow_pre_security_callout: true
        # file_id callouts run through this outbound filter chain via the
        # filtered subrequest executor, which enforces destination authority,
        # DNS/SSRF (from the pipeline's allow_private_upstreams), TLS/SNI, and
        # Host centrally — this holds even with an empty or omitted chain. The
        # chain only re-projects the trusted tenant/subject identity from the
        # StateOwner extension; client file_url downloads never traverse it.
        outbound_chain:
          name: files-api-outbound
          filters:
            - filter: project_state_owner_headers
              tenant_header: x-tenant-id
              subject_header: x-user-id
        on_missing: reject
        timeout_ms: 10000
        conditions:
          - unless:
              bound_upstream:
                application_provider: openai

      - filter: openai_doc_extract
        allow_pre_security_callout: true
        on_unsupported: continue
        conditions:
          - unless:
              bound_upstream:
                application_provider: openai

      - filter: openai_mcp_tool_resolve
        user_credential: mcp_gateway
        authorization_assertion: mcp_gateway
        forward_headers:
          - x-tenant-id
          - x-user-id
        # Ambient trusted headers are forwarded only for connector-backed MCP
        # entries; arbitrary client-supplied server_url targets never receive
        # them. Clients reference this destination with connector_id.
        connectors:
          - id: trusted-mcp
            server_url: http://127.0.0.1:8001/mcp
        conditions:
          - unless:
              bound_upstream:
                application_provider: openai

      # The router needs a /v1/responses prefix route so managed GET/DELETE
      # resource endpoints can reach the local store. That prefix also binds
      # unknown subpaths. After recognized store operations have terminated,
      # reject any remaining managed subpath here so it cannot enter the IRR.
      # Provider-owned OpenAI traffic already left through direct-openai.
      - filter: static_response
        status: 404
        body: '{"error":{"message":"unsupported managed Responses operation","type":"invalid_request_error","param":null,"code":"invalid_request_error"}}'
        headers:
          - name: Content-Type
            value: application/json
        conditions:
          - when:
              path_prefix: "/v1/responses"
          - unless:
              path: "/v1/responses"

      # Gateway-owned Responses requests flow here after both terminal direct
      # branches have been considered. OpenAI-bound requests skip the IRR
      # structurally because direct-openai rejoins at terminal above. Do not add
      # a bound_upstream condition to IRR itself: IRR owns an ordinary pre-read
      # body hook, which cannot be gated on metadata published later by the
      # binding router. The IRR runs the agentic model-tool-model loop: each round
      # streams through openai_stream_events, the three request-phase
      # dispatchers execute the loop owner's assigned hosted-tool calls, and
      # openai_agentic_loop parses each response and publishes the single
      # loop/done transition. The bound-body openai_response_store persists the
      # composed object on the response path (reverse order). The inference
      # step consumes the frozen logical binding directly; it must not route
      # again or attempt to replace the request-level provider decision.
      - filter: iterative_request_router
        initial_step: inference
        # Safety cap: at least max_infer_iters + 1 (7 + 1 = 8).
        max_iterations: 8
        # The IRR's 30s default is too short for production model streams.
        # Allow six minutes total for time-to-first-byte and logical
        # streaming across the loop; each step inherits a five-minute
        # deadline (the vLLM integration tests already run at 300s).
        timeout_ms: 360000
        step_timeout_ms: 300000
        max_response_bytes: 67108864
        max_stream_response_bytes: 67108864
        max_state_bytes: 136314880
        steps:
          - name: inference
            filters:
              # Re-project inside the destination-owned subrequest. Core moves
              # the normalized StateOwner extension into each IRR step; this
              # strips any still-uncommitted raw ingress identity headers and
              # recreates only the OGX/callout contract.
              - filter: project_state_owner_headers
                tenant_header: x-tenant-id
                subject_header: x-user-id

              # Composes every inference SSE stream into one logical
              # Responses lifecycle: preserves one response identity across
              # rounds and withholds per-round terminal events until the IRR
              # transition is known. A single inference round is just a
              # one-round logical stream, so this filter always composes and
              # must run inside this IRR step. REQUIRED for both streaming
              # persistence and loop-terminal error delivery (see the
              # "Streaming persistence" note above); openai_agentic_loop
              # fails closed with a 500 on an effective stream:true request
              # when it is absent.
              #
              # Its limits are set at or above responses_to_chat_completions'
              # converter bounds below so no frame the client receives is ever
              # rejected here and dropped from the store (the accumulator must
              # never be tighter than the converter): max_events (200000) >=
              # converter max_stream_events (100000); max_buffer_bytes (64 MiB)
              # > converter max_emitted_sse_frame_bytes (48 MiB); timeout_secs
              # (600) >= converter stream_timeout_secs (300). These apply on the
              # native path too and stay within the IRR's byte/time caps above.
              - filter: openai_stream_events
                max_buffer_bytes: 67108864
                max_events: 200000
                timeout_secs: 600

              # Request-phase dispatcher: runs first on re-entry to execute
              # web_search_call assignments prepared by the loop owner. Inert
              # until the model emits a web_search_call.
              - filter: openai_web_search
                provider: brave
                api_key: ${WEB_SEARCH_API_KEY}
                # Require the per-user key captured by the outer
                # callout_credentials filter. The secret is injected only at
                # the exact resolved provider authority.
                user_credential: brave_search
                max_calls_per_round: 32
              # Request-phase dispatcher: runs second on re-entry to execute
              # MCP calls prepared and classified by the loop owner. Inert
              # until the model emits an MCP tool call.
              - filter: openai_mcp_dispatch
                user_credential: mcp_gateway
                authorization_assertion: mcp_gateway
                forward_headers:
                  - x-tenant-id
                  - x-user-id
                max_calls_per_round: 32
                max_parallel_calls: 8
                max_result_bytes: 1048576
                max_total_result_bytes: 8388608
              # Request-phase dispatcher: at request-body EOS on each IRR
              # re-entry it executes the file_search_call items the loop owner
              # assigned in the prior response, reconciling each in place
              # inside ResponsesState.accumulated_output. It never parses the
              # response body and never decides whether another round runs,
              # so it is inert unless the model emits a hosted file_search_call.
              - filter: openai_file_search_callout
                vector_store_url: http://127.0.0.1:3002
                user_credential: ogx_files
                # Every vector-store sub-request runs through the filtered
                # subrequest executor, which enforces destination authority,
                # DNS/SSRF, TLS/SNI, and Host centrally (destination from
                # vector_store_url) — this holds even with an empty or omitted
                # chain. The outbound chain only re-projects the trusted
                # tenant/subject identity from the StateOwner extension.
                outbound_chain:
                  name: vector-store-outbound
                  filters:
                    - filter: project_state_owner_headers
                      tenant_header: x-tenant-id
                      subject_header: x-user-id
                timeout_ms: 5000
                max_response_bytes: 10485760
                max_total_response_bytes: 67108864
                max_state_bytes: 136314880
                on_failure: closed
              # Sole loop owner: parses each model response, records the
              # hosted-tool assignments for the dispatchers, and publishes the
              # single continuation signal (action=loop|done). max_infer_iters
              # must stay below the IRR max_iterations safety cap (7 + 1 = 8).
              - filter: openai_agentic_loop
                max_infer_iters: 7
              # Protocol adapters run after the load balancer publishes the
              # selected cluster's typed application metadata.
              - filter: headers
                request_set:
                  - name: Content-Type
                    value: application/json
              # Select an endpoint from the request-level binding and publish
              # exchange-local application metadata. Every selected_upstream
              # condition below reads this fresh selection, so the matching
              # protocol adapter runs independently on every IRR round.
              - filter: load_balancer
                cluster_source: bound_upstream
                clusters:
                  - name: "inference-backend"
                    # Gateway-owned native Responses backend.
                    http:
                      application_protocol: openai_responses
                      application_provider: vllm
                    endpoints:
                      - "127.0.0.1:3001"
                  - name: "chat-completions-backend"
                    # Gateway-owned Chat Completions backend.
                    http:
                      application_protocol: openai_chat_completions
                      application_provider: vllm
                    endpoints:
                      - "127.0.0.1:3001"

              # ---- Protocol-adaptive body builders (run AFTER selection) ----
              # Exactly one branch fires, chosen by the selected upstream's
              # declared application_protocol. selected_upstream fails closed:
              # always declare http.application_protocol on the cluster above.

              # Native Responses upstream: enforce provider-aware prompt
              # support, rebuild the outbound body from ResponsesState, and
              # select the typed streaming transport for `stream: true`.
              - filter: openai_responses_proxy
                conditions:
                  - when:
                      selected_upstream:
                        application_protocol: openai_responses

              # Chat Completions upstream: translate the canonical Responses
              # request into Chat Completions on the request path, and
              # translate the Chat Completions response (finite JSON or streamed
              # SSE) back into a Responses resource / Responses SSE on the
              # response path. Because response filters run in reverse, this
              # converts the raw Chat response into Responses BEFORE
              # openai_agentic_loop and openai_stream_events observe it (both
              # require Responses). It fully owns outbound body build for Chat
              # backends. It also rejects OpenAI prompt templates, which have no
              # Chat representation. Stream bounds stay at/under
              # openai_stream_events'
              # limits so no frame the client receives is dropped from the store
              # (cf. responses-to-chat-completions.yaml).
              - filter: responses_to_chat_completions
                max_rewritten_body_bytes: 33554432
                max_emitted_sse_frame_bytes: 50331648
                max_stream_events: 100000
                stream_timeout_secs: 300
                conditions:
                  - when:
                      selected_upstream:
                        application_protocol: openai_chat_completions

              # The translator owns body/response conversion but deliberately
              # leaves the request URI unchanged. A Chat Completions backend
              # still expects POST /v1/chat/completions, so rewrite that transport
              # detail separately. The identical protocol gate plus create
              # verb/path guards ensure only translated requests are rewritten.
              - filter: path_rewrite
                replace:
                  pattern: "^/v1/responses/?$"
                  replacement: "/v1/chat/completions"
                conditions:
                  - when:
                      selected_upstream:
                        application_protocol: openai_chat_completions
                      path_prefix: "/v1/responses"
                      methods: [POST]
            on_result:
              # openai_agentic_loop is the sole loop authority (issue #1046):
              # it parses each model response, records assignments for the
              # request-phase dispatchers (web_search, mcp_dispatch,
              # file_search_callout), and publishes the single continuation
              # signal. The IRR transitions only on the owner's action, never
              # on a dispatcher's.
              - filter: openai_agentic_loop
                key: action
                value: loop
                next: inference
              - default: true
                done: true

insecure_options:
  allow_private_endpoints: true # example proxies to local backends
  # Central SSRF control: permits resolved callout addresses (vector_store_url,
  # file_id Files API) to be private/loopback at connect time (needed for
  # local/private vector stores and Files API backends).
  allow_private_upstreams: true