Intelligent Route Management Skip

Demonstrates the intelligent_route management-path skip list: management and discovery endpoints (model listing, subscriptions, API-key management, health) bypass model resolution entirely, while inference paths are still resolved by model and fail closed on an unknown model

Category: Practical
Task: Demonstrates the intelligent_route management-path skip list: management and discovery endpoints (model listing, subscriptions, API-key management, health) bypass model resolution entirely, while inference paths are still resolved by model and fail closed on an unknown model

Prerequisites: The product runtime and any backend services referenced by this configuration.

Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# Intelligent Route: Management-Path Skip List
#
# Demonstrates the `intelligent_route` management-path skip list: management
# and discovery endpoints (model listing, subscriptions, API-key management,
# health) bypass model resolution entirely, while inference paths are still
# resolved by model and fail closed on an unknown model.
#
# Skipping matters because a management request may carry a JSON body `model`
# field that would otherwise be resolved and fail closed (404). `skip_paths`
# bypass resolution regardless of body content; the optional `management_cluster`
# routes them to a management backend so `load_balancer` can terminate without a
# downstream `router`. Omit `management_cluster` when a fronting gateway already
# serves these paths; keep it for the standalone proxy shown here.
#
listeners:
  - name: proxy
    address: "0.0.0.0:8080"
    filter_chains:
      - main

filter_chains:
  - name: main
    filters:
      # Promote the JSON body's `model` field into X-Model and strip any
      # client-supplied X-Model so the routed model cannot be spoofed.
      - filter: model_to_header
        header: X-Model

      # Resolve inference clusters by model. Management/discovery paths listed
      # in skip_paths bypass this resolution and are routed straight to
      # management_cluster, so load_balancer can serve them directly.
      - filter: intelligent_route
        local_site: site-a
        model_header: X-Model
        management_cluster: management-api
        skip_paths:
          - /v1/models
          - /v1/subscriptions
          - /v1/api-keys
          - /health
        candidates:
          - kind: inference_model
            name: granite-3.3-8b
            site: site-a
            cluster: inference-local

      # Forward to whichever cluster intelligent_route selected.
      - filter: load_balancer
        clusters:
          - name: inference-local
            endpoints:
              - "127.0.0.1:8001"
          - name: management-api
            endpoints:
              - "127.0.0.1:8003"

admin:
  address: "127.0.0.1:9901"
shutdown_timeout_secs: 5

insecure_options:
  allow_private_endpoints: true # example proxies to local backends