Ai Inference Body Based Routing

Routes LLM API requests to different backends based on the model field in the JSON request body

Category: Practical
Task: Routes LLM API requests to different backends based on the model field in the JSON request body

Prerequisites: The product runtime and any backend services referenced by this configuration.

Run it: Use ghcr.io/praxis-proxy/ai:0.5.0 and follow the container quickstart to mount and start the configuration.

This configuration comes from the selected release. The example has not been run here; external services are not bundled.

Download the source file.

# AI Inference Body-Based Routing
#
# Routes LLM API requests to different backends based on the
# `model` field in the JSON request body. Uses the json_body_field
# filter to extract the model name into an X-Model header, then
# the router matches on that header to select the correct cluster.
#
# This demonstrates StreamBuffer mode: the request body is
# inspected before upstream selection, enabling routing decisions
# based on body content.
#
# Example request:
#
#   curl -X POST http://localhost:8080/v1/chat/completions \
#     -H "Content-Type: application/json" \
#     -d '{"model": "mistral-7b-instruct", "messages": [...]}'
#
# This request is routed to the mistral cluster. A request with
# "model": "granite-3.1-8b" goes to the granite cluster. Unknown
# models fall through to the default cluster.

listeners:
  - name: inference-gateway
    address: "0.0.0.0:8080" # dev value; binds all interfaces
    filter_chains:
      - extract-model
      - routing

filter_chains:
  # Extract the model field from the JSON body and promote it
  # to the X-Model request header.
  - name: extract-model
    filters:
      - filter: json_body_field
        field: model
        header: X-Model

  # Route based on the promoted X-Model header.
  - name: routing
    filters:
      - filter: router
        routes:
          - path_prefix: "/"
            headers:
              x-model: "mistral-7b-instruct"
            cluster: mistral
          - path_prefix: "/"
            headers:
              x-model: "granite-3.1-8b"
            cluster: granite
          - path_prefix: "/"
            cluster: default
      - filter: load_balancer
        clusters:
          - name: mistral
            endpoints:
              - "10.0.1.1:8080"
              - "10.0.1.2:8080"
          - name: granite
            endpoints:
              - "10.0.2.1:8080"
              - "10.0.2.2:8080"
          - name: default
            endpoints:
              - "10.0.3.1:8080"

insecure_options:
  allow_private_endpoints: true # example proxies to local backends