Mcp Streaming

Demonstrates MCP tool calls over the filtered-subrequest transport with SSE streaming support

Category: Setup-dependent integration
Task: Demonstrates MCP tool calls over the filtered-subrequest transport with SSE streaming support

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# MCP Streaming Callout Transport
# Requires `--features openai-mcp-tools,store-sqlite` because these filters are opt-in.
#
# Demonstrates MCP tool calls over the filtered-subrequest transport with SSE
# streaming support. The openai_mcp_dispatch filter executes MCP calls against
# server URLs supplied in the client request, routing them through a filtered
# outbound chain that can observe and stamp headers without selecting an
# upstream (the MCP transport itself stages the destination).
#
# SSE streaming support:
#   The MCP callout transport can consume both buffered JSON and SSE-streamed
#   tool results from MCP servers. When the MCP server returns a tools/call
#   result as `text/event-stream`, the transport parses the SSE stream, applies
#   byte and time budgets, and materializes the result for dispatch. Oversized
#   SSE streams are classified as 413 per the configured max_result_bytes.
#
# Outbound chain requirements:
#   The openai_mcp_streaming_selector filter is required in any MCP outbound
#   chain to enable SSE transport negotiation. For an INLINE outbound_chain
#   (name + filters), the selector is auto-injected as the first entry; you do
#   NOT add it manually. For a NAMED outbound_chain (a top-level filter_chains
#   reference), you MUST list `- filter: openai_mcp_streaming_selector` as the
#   first, unconditional filter or the proxy refuses to start.
#
# StreamBuffer validation:
#   A response-body-buffering (StreamBuffer) filter in the MCP outbound_chain
#   is rejected at boot, since SSE streams must be consumed incrementally and
#   cannot be buffered.
#
# This example uses an INLINE outbound_chain for openai_mcp_dispatch, so the
# selector is auto-injected. Only the request-stamping headers filter is
# explicitly listed.
#
# Example request:
#
#   curl -X POST http://localhost:8080/v1/responses \
#     -H "Content-Type: application/json" \
#     -d '{
#       "model": "gpt-4.1",
#       "input": "What is the weather in SF?",
#       "tools": [{
#         "type": "mcp",
#         "server_label": "weather",
#         "server_url": "http://mcp-weather:8001/mcp",
#         "allowed_tools": ["get_weather"],
#         "require_approval": "never"
#       }]
#     }'
#
# Build:
#   cargo build -p praxis-ai-proxy --features openai-mcp-tools,store-sqlite

listeners:
  - name: ai-gateway
    address: "127.0.0.1:8080"
    filter_chains: [mcp-streaming-pipeline]

filter_chains:
  - name: mcp-streaming-pipeline
    filters:
      - filter: openai_responses_request
        on_invalid: reject
        headers:
          format: x-praxis-ai-format
          model: x-praxis-ai-model
          stream: x-praxis-ai-stream

      - filter: openai_tool_parse

      - filter: state_owner
        mode: single_tenant
        tenant_id: default

      - filter: openai_response_store
        backend: sqlite
        database_url: "sqlite://responses.db?mode=rwc"
        responses_table: openai_responses
        conversations_table: openai_conversations

      - filter: openai_mcp_tool_resolve
        timeout_ms: 5000

      - filter: iterative_request_router
        initial_step: inference
        max_iterations: 11
        max_stream_response_bytes: 67108864
        timeout_ms: 360000
        steps:
          - name: inference
            filters:
              # Composes every inference SSE stream into one logical Responses
              # lifecycle: preserves one response identity across rounds and
              # withholds per-round terminal events until the IRR transition is
              # known. Stays before openai_agentic_loop so response-phase reverse
              # order lets the loop parse the raw SSE.
              #
              # REQUIRED because `openai_responses_proxy` below selects typed
              # streaming automatically for an effective `"stream": true`
              # request.
              - filter: openai_stream_events

              # Request-phase dispatcher: executes MCP calls prepared by the
              # loop owner. The outbound_chain is INLINE (name + filters), so
              # openai_mcp_streaming_selector is auto-injected as the first
              # entry; we only list the request-stamping headers filter here.
              #
              # This filter runs inside an iterative_request_router step, whose
              # pipeline is built with an empty named-chain map, so only inline
              # outbound_chain definitions are supported here (top-level
              # filter_chains references cannot resolve).
              #
              # Do NOT put static credentials in this chain: MCP destinations
              # are client-selected (server_url), so a blanket credential header
              # would leak to attacker-chosen hosts.
              - filter: openai_mcp_dispatch
                outbound_chain:
                  name: mcp-dispatch-outbound
                  filters:
                    - filter: headers
                      request_set:
                        - name: X-MCP-Client
                          value: praxis-ai-gateway
                max_calls_per_round: 32
                max_parallel_calls: 8
                max_result_bytes: 1048576
                max_total_result_bytes: 8388608

              - filter: openai_agentic_loop
                max_infer_iters: 10

              # Selects typed streaming automatically when the effective
              # outbound body carries `"stream": true`.
              - filter: openai_responses_proxy

              - filter: router
                routes:
                  - path_prefix: "/"
                    cluster: "inference-backend"

              - filter: load_balancer
                clusters:
                  - name: "inference-backend"
                    read_timeout_ms: 300000
                    endpoints:
                      - "127.0.0.1:3001"
            on_result:
              # openai_agentic_loop is the sole loop authority: it parses each
              # model response, records assignments for the request-phase
              # dispatchers, and publishes the single continuation signal.
              - filter: openai_agentic_loop
                key: action
                value: loop
                next: inference
              - default: true
                done: true

insecure_options:
  allow_private_endpoints: true
  # Required for loopback MCP server URLs supplied in the client request.
  allow_private_upstreams: true