Irr Terminal Streaming

Demonstrates a single-step iterative_request_router pipeline that exposes a native OpenAI Responses SSE body incrementally. openai_responses_proxy always advertises the streaming capability and selects Praxis’s typed streaming transport automatically for an effective “stream”: true request; there is no operator opt-in

Category: Setup-dependent integration
Task: Demonstrates a single-step iterative_request_router pipeline that exposes a native OpenAI Responses SSE body incrementally. openai_responses_proxy always advertises the streaming capability and selects Praxis’s typed streaming transport automatically for an effective “stream”: true request; there is no operator opt-in

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# Terminal Responses Streaming through IRR
# Requires `--features openai-responses` because these filters are opt-in.
#
# Demonstrates a single-step iterative_request_router pipeline that exposes a
# native OpenAI Responses SSE body incrementally. `openai_responses_proxy`
# always advertises the streaming capability and selects Praxis's typed
# streaming transport automatically for an effective `"stream": true` request;
# there is no operator opt-in. Streaming is not a Praxis IRR YAML transport
# mode.
#
# `openai_responses_format` publishes descriptive client intent. The final
# outbound serializer, `openai_responses_proxy`, independently keeps the
# provider-visible `"stream": true` field aligned with Praxis's typed
# `SubRequestResponseMode::Streaming` selection.
#
# Every response filter in a streaming-capable step must use `BodyMode::Stream`.
# IRR may also resume the same downstream stream after response-dependent
# transitions; the agentic-loop example demonstrates that multi-round form.
#
# The typed response mode selected by `openai_responses_proxy` arms the
# step-local `openai_stream_events` filter. It receives each SSE chunk before it
# is sent to the client and preserves parser state through stream completion.
#
# Example:
#
#   curl -N http://localhost:8080/v1/responses \
#     -H "Content-Type: application/json" \
#     -d '{"model":"gpt-4.1","input":"Say hello","stream":true}'

listeners:
  - name: ai-gateway
    address: "127.0.0.1:8080"
    filter_chains: [responses-terminal-stream]

filter_chains:
  - name: responses-terminal-stream
    filters:
      - filter: openai_responses_request
        on_invalid: reject

      - filter: iterative_request_router
        initial_step: inference
        max_iterations: 1
        # The IRR's 30s default is too short for production model streams.
        # Allow six minutes total for time-to-first-byte and response streaming.
        timeout_ms: 360000
        steps:
          - name: inference
            filters:
              - filter: openai_responses_proxy

              - filter: headers
                request_set:
                  - name: Content-Type
                    value: application/json

              - filter: router
                routes:
                  - path: "/v1/responses"
                    cluster: "inference-backend"

              - filter: load_balancer
                clusters:
                  - name: "inference-backend"
                    endpoints:
                      - "127.0.0.1:3001"

              # After load_balancer so IRR body hooks see the selected peer.
              # timeout_secs starts at the first SSE chunk; each chunk recaps
              # leftover budget onto the live body.
              - filter: openai_stream_events
            on_result:
              - default: true
                done: true

insecure_options:
  allow_private_endpoints: true # example proxies to local backends