Intelligent Route Overlay

Routes requests using a routing overlay file (routing-overlay.json) instead of inline YAML candidates. The overlay is rendered by the operator into a Kubernetes ConfigMap and projected as a volume mount

Category: Practical
Task: Routes requests using a routing overlay file (routing-overlay.json) instead of inline YAML candidates. The overlay is rendered by the operator into a Kubernetes ConfigMap and projected as a volume mount

Prerequisites: The product runtime and any backend services referenced by this configuration.

Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# Intelligent Route: Overlay File Mode
#
# Routes requests using a routing overlay file (`routing-overlay.json`)
# instead of inline YAML candidates.  The overlay is rendered by
# the operator into a Kubernetes ConfigMap and projected as a volume mount.
#
# When the overlay file changes on disk, the filter detects the change via
# inotify, debounces filesystem events, validates the new overlay, and
# atomically swaps the routing snapshot.  In-flight requests continue using
# their existing snapshot.  No pod restart, rollout restart, or Kubernetes
# API call is required.
#
# If the new overlay is invalid or unreadable, the previous snapshot is
# retained and an error is logged.
#
listeners:
  - name: proxy
    address: "0.0.0.0:8080"
    filter_chains:
      - main

filter_chains:
  - name: main
    filters:
      # Promote the JSON body's `model` field into X-Model and strip any
      # client-supplied X-Model so the routed model cannot be spoofed.
      - filter: model_to_header
        header: X-Model

      # Route from the routing overlay file.
      - filter: intelligent_route
        overlay_file: /etc/praxis/routing/routing-overlay.json
        model_header: X-Model
        reload:
          enabled: true
          debounce_ms: 500

      # Forward to whichever cluster intelligent_route selected.
      - filter: load_balancer
        clusters:
          - name: granite-local
            endpoints:
              - "127.0.0.1:8001"
          - name: llama-remote
            endpoints:
              - "127.0.0.1:8002"

admin:
  address: "127.0.0.1:9901"
shutdown_timeout_secs: 5

insecure_options:
  allow_private_endpoints: true # example proxies to local backends