Web Search Scoped Credentials

A scoped-credentials variant of web-search-chat-completions.yaml

Category: Setup-dependent integration
Task: A scoped-credentials variant of web-search-chat-completions.yaml

Prerequisites: The external service, credentials, or certificates referenced by this configuration.

Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

Companion resources from the same snapshot:

# Responses Web Search with Per-User Callout Credentials + Owner Attribution
#
# A scoped-credentials variant of web-search-chat-completions.yaml. Accepts
# OpenAI Responses requests with hosted web search against a Chat Completions
# backend, and demonstrates PR1 of issue #880:
#
#   * state_owner maps trusted ingress identity headers (x-auth-tenant /
#     x-auth-user) into a normalized owner the web-search callout projects
#     into its nested provider subrequest.
#   * callout_credentials captures a PER-USER provider key from the trusted
#     ingress header x-user-brave-key into the `brave_search` slot and strips
#     the header so it never reaches an upstream. openai_web_search stages that
#     per-user secret as the Brave `x-subscription-token` on the provider
#     callout instead of the shared api_key. A request missing the header fails
#     closed with a 401 before any provider callout runs.
#
# Both filters run in the OUTER chain (before the iterative_request_router):
# core threads their extensions into every IRR step, and the ingress header is
# stripped before the IRR clones request headers.
#
#   curl -N http://localhost:8080/v1/responses \
#     -H "Content-Type: application/json" \
#     -H "x-auth-tenant: acme" -H "x-auth-user: alice" \
#     -H "x-user-brave-key: $ALICE_BRAVE_KEY" \
#     -d '{"model":"gpt-4.1","input":"What is the weather in SF?","tools":[{"type":"web_search"}]}'

listeners:
  - name: responses-web-search-chat-gateway
    address: "127.0.0.1:8080"
    filter_chains: [responses-web-search-chat]

filter_chains:
  - name: responses-web-search-chat
    filters:
      - filter: state_owner
        # Deployment-specific trusted ingress identity. The authentication
        # boundary must overwrite these and prevent clients from spoofing them.
        mode: trusted_headers
        tenant:
          header: x-auth-tenant
        issuer:
          static: urn:example:gateway
        subject:
          header: x-auth-user

      - filter: callout_credentials
        # Capture the per-user Brave key from a trusted ingress header into the
        # `brave_search` slot and strip the header so it never reaches an
        # upstream. openai_web_search stages this per-user secret instead of the
        # shared api_key below.
        #
        # SECURITY: each source_header is trusted-boundary-owned. The
        # authentication boundary MUST unconditionally delete then set every
        # source_header on every request, so a client cannot spoof another
        # user's credential by supplying the header itself.
        credentials:
          - slot: brave_search
            source_header: x-user-brave-key

      - filter: openai_responses_format
        on_invalid: reject
        headers:
          format: x-praxis-ai-format
          model: x-praxis-ai-model
          stream: x-praxis-ai-stream

      - filter: openai_responses_validate

      - filter: openai_tool_parse

      - filter: iterative_request_router
        initial_step: inference
        max_iterations: 11
        # Terminal streaming rounds resume real inference after the web-search
        # dispatch, so allow generous per-request and per-step deadlines for
        # slower backends (mirrors full-flow-agentic.yaml). The default step
        # deadline is too tight for a full translated streaming generation.
        timeout_ms: 120000
        step_timeout_ms: 60000
        max_stream_response_bytes: 67108864
        steps:
          - name: inference
            filters:
              # Parses every translated inference SSE stream, preserves one
              # logical response identity across rounds, and withholds per-round
              # terminal events until the IRR transition is known. Runs last in
              # the response phase so it observes the translated Responses SSE. A
              # no-op on finite (non-SSE) responses.
              - filter: openai_stream_events

              - filter: openai_web_search
                provider: brave
                api_key: ${WEB_SEARCH_API_KEY}
                # Require the per-user Brave key captured into the `brave_search`
                # slot by callout_credentials. A request without it fails closed
                # with a 401 (missing_callout_context) before any provider callout.
                user_credential: brave_search
                # Each provider callout executes through this outbound chain; the
                # filtered-subrequest executor enforces destination authority,
                # DNS/SSRF, TLS/SNI, and Host centrally.
                outbound_chain:
                  name: web-search-outbound
                  filters:
                    - filter: request_id

              - filter: openai_agentic_loop
                max_infer_iters: 10

              - filter: responses_to_chat_completions
                max_rewritten_body_bytes: 67108864

              - filter: path_rewrite
                replace:
                  pattern: "^/v1/responses/?$"
                  replacement: "/v1/chat/completions"
                conditions:
                  - when:
                      path_prefix: "/v1/responses"
                      methods: [POST]

              - filter: headers
                request_set:
                  - name: Content-Type
                    value: application/json

              - filter: router
                routes:
                  - path: "/v1/chat/completions"
                    cluster: "chat-completions-backend"

              - filter: load_balancer
                clusters:
                  - name: "chat-completions-backend"
                    endpoints:
                      - "127.0.0.1:3001"
            on_result:
              # openai_agentic_loop is the sole loop authority (issue #1046): it
              # parses each model response, records the assignment consumed by
              # the request-phase openai_web_search dispatcher, and publishes the
              # single continuation signal. The IRR transitions only on the
              # owner's action, never on a dispatcher's.
              - filter: openai_agentic_loop
                key: action
                value: loop
                next: inference
              - default: true
                done: true

insecure_options:
  allow_private_endpoints: true # example proxies to local backends