Irr Terminal Streaming
Demonstrates a single-step iterative_request_router pipeline that exposes a native OpenAI Responses SSE body incrementally. openai_responses_proxy always advertises the streaming capability and selects Praxis’s typed streaming transport automatically for an effective “stream”: true request; there is no operator opt-in
Category: Setup-dependent integration
Task: Demonstrates a single-step iterative_request_router pipeline that exposes a native OpenAI Responses SSE body incrementally. openai_responses_proxy always advertises the streaming capability and selects Praxis’s typed streaming transport automatically for an effective “stream”: true request; there is no operator opt-in
Prerequisites: The external service, credentials, or certificates referenced by this configuration.
This configuration comes from the selected release. The example has not been run here; external services are not bundled.
Download the source file.
# Terminal Responses Streaming through IRR
# Requires `--features openai-responses` because these filters are opt-in.
#
# Demonstrates a single-step iterative_request_router pipeline that exposes a
# native OpenAI Responses SSE body incrementally. `openai_responses_proxy`
# always advertises the streaming capability and selects Praxis's typed
# streaming transport automatically for an effective `"stream": true` request;
# there is no operator opt-in. Streaming is not a Praxis IRR YAML transport
# mode.
#
# `openai_responses_format` publishes descriptive client intent. The final
# outbound serializer, `openai_responses_proxy`, independently keeps the
# provider-visible `"stream": true` field aligned with Praxis's typed
# `SubRequestResponseMode::Streaming` selection.
#
# Every response filter in a streaming-capable step must use `BodyMode::Stream`.
# IRR may also resume the same downstream stream after response-dependent
# transitions; the agentic-loop example demonstrates that multi-round form.
#
# The typed response mode selected by `openai_responses_proxy` arms the
# step-local `openai_stream_events` filter. It receives each SSE chunk before it
# is sent to the client and preserves parser state through stream completion.
#
# Example:
#
# curl -N http://localhost:8080/v1/responses \
# -H "Content-Type: application/json" \
# -d '{"model":"gpt-4.1","input":"Say hello","stream":true}'
listeners:
- name: ai-gateway
address: "127.0.0.1:8080"
filter_chains: [responses-terminal-stream]
filter_chains:
- name: responses-terminal-stream
filters:
- filter: openai_responses_request
on_invalid: reject
- filter: iterative_request_router
initial_step: inference
max_iterations: 1
# The IRR's 30s default is too short for production model streams.
# Allow six minutes total for time-to-first-byte and response streaming.
timeout_ms: 360000
steps:
- name: inference
filters:
- filter: openai_responses_proxy
- filter: headers
request_set:
- name: Content-Type
value: application/json
- filter: router
routes:
- path: "/v1/responses"
cluster: "inference-backend"
- filter: load_balancer
clusters:
- name: "inference-backend"
endpoints:
- "127.0.0.1:3001"
# After load_balancer so IRR body hooks see the selected peer.
# timeout_secs starts at the first SSE chunk; each chunk recaps
# leftover budget onto the live body.
- filter: openai_stream_events
on_result:
- default: true
done: true
insecure_options:
allow_private_endpoints: true # example proxies to local backends