Intelligent Route Inference

Use model identity to select the matching inference backend.

Category: Practical
Task: Route inference requests by model

Prerequisites: Praxis AI runtime; A model-to-header mapping; Reachable inference backends; Provider credentials when required

Expected outcome: The configured model identity selects an inference backend.

Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# Intelligent Route: Inference Model Routing
#
# Routes requests to different upstream clusters based on the inference
# model name carried in a configured request header.  The header value is
# set by the `model_to_header` filter, which promotes the JSON body's
# `model` field and strips any client-supplied copy of the header so
# routing cannot be spoofed from the wire.
#
# `intelligent_route` selects a candidate by static scoring and writes the
# cluster name into `ctx.cluster`.  The downstream `load_balancer` reads
# that cluster.  No `router` filter is required between them.
#
# Scoring: fresh > stale, local-site > remote on equal freshness.
# First configured candidate wins on equal score.
#
# This filter is registered by praxis-ai-proxy, not Praxis core, because
# it encodes AI-specific routing semantics.
#
listeners:
  - name: proxy
    address: "0.0.0.0:8080"
    filter_chains:
      - main

filter_chains:
  - name: main
    filters:
      # Promote the JSON body's `model` field into X-Model and strip any
      # client-supplied X-Model so the routed model cannot be spoofed.
      - filter: model_to_header
        header: X-Model

      # Route to the cluster that serves the requested model.
      # Selects by highest score: fresh > stale, local > remote.
      - filter: intelligent_route
        local_site: site-a
        model_header: X-Model
        candidates:
          - kind: inference_model
            name: granite-3.3-8b
            site: site-a
            cluster: granite-local
            fresh: true

          - kind: inference_model
            name: llama-3.2-8b
            site: site-b
            cluster: llama-remote
            fresh: true

      # Forward to whichever cluster intelligent_route selected.
      - filter: load_balancer
        clusters:
          - name: granite-local
            endpoints:
              - "127.0.0.1:8001"
          - name: llama-remote
            endpoints:
              - "127.0.0.1:8002"

admin:
  address: "127.0.0.1:9901"
shutdown_timeout_secs: 5

insecure_options:
  allow_private_endpoints: true # example proxies to local backends