AI Inference
On this page
Body-aware classification, routing, and enrichment for AI inference traffic, built on the filter pipeline and StreamBuffer body access pattern.
Overview
AI inference filters classify request bodies to determine the API format (OpenAI Responses, Anthropic Messages, Chat Completions), extract routing signals (model, stream mode, store flag), and promote them to headers, metadata, and filter results for downstream routing via branch chains.
Request Body
|
v
Classifier (pure function)
|
v
Format Filter (promotes facts to headers/metadata/results)
|
v
Validate Filter (JSON parsing, metadata enrichment, ID generation)
|
v
Branch Chains / Router (routing decisions)
|
v
Upstream
Classification Pipeline
Format Detection
The classifier (classifier/mod.rs) is a pure
function with no I/O. It parses the request body
JSON once and returns a ClassifiedRequest struct
with extracted facts.
Detection precedence:
inputfield present or path-based responses endpoint: Responses APImessages+max_tokens+ Anthropic signals (systemor typed content blocks): Anthropic Messages APImessagesalone: Chat Completions API- Valid JSON without recognized fields: UnknownJson
- Invalid JSON: InvalidJson
- Non-JSON content type: NonJson
Path-based classification handles sub-resource
endpoints (GET /v1/responses/{id},
POST /v1/responses/{id}/cancel, etc.) that lack
a request body.
Metadata Propagation
The format filter promotes classified facts using three channels:
- Filter metadata: durable key-value pairs
(e.g.
openai_responses_format.model) that persist across Pingora phases. Used for cross-filter communication. - Extra request headers: added to the upstream
request (e.g.
X-Praxis-AI-Format). Used for header-based routing in the router filter. - Filter results: written to
FilterResultSetfor branch chain condition evaluation.
All promoted values are validated against a 256-byte length limit and checked for control characters before propagation.
Inside iterative_request_router steps, metadata is
reset for credential isolation but extensions persist.
Filters that read classifier metadata must have the
classifier re-run inside their step. See
examples/configs/inference/fallback-with-translation.yaml
for the full pattern.
Stateful vs Stateless
Responses API requests are classified as “stateful” when any of these hold:
previous_response_idis presenttoolsis presentstoreis not explicitlyfalsehas_conversationis truehas_prompt_idis true
Requests with background: true are rejected before mode
classification because Praxis does not implement the
asynchronous Responses lifecycle.
Stateful mode influences routing decisions (e.g. directing to clusters with response store access).
StreamBuffer Body Access
StreamBuffer is the key enabler for AI inference filters. It accumulates request body chunks and defers upstream forwarding until the filter releases or end-of-stream:
- Buffer the JSON envelope (model name, parameters, prompt prefix).
- Extract routing signals from the buffered bytes.
- Select the upstream based on body content.
- Release the buffered prefix and stream the remainder.
This peek-then-stream pattern avoids the latency of external processor architectures while providing body visibility where it matters.
Filters declare BodyAccess::ReadOnly +
BodyMode::StreamBuffer { max_bytes } to opt in.
Only PromptEnrichFilter uses ReadWrite (it
modifies the messages array).
Filters
model_to_header
Extracts the model field from JSON request bodies
and promotes it to a configurable header (default
X-Model). Enables header-based routing to
provider-specific clusters.
openai_responses_format
Classifies AI API request bodies and promotes format, model, stream, store, background, and mode to headers, metadata, and filter results.
openai_responses_validate
Parses Responses API request JSON, enriches filter metadata,
and generates cryptographically random response and
conversation IDs with resp_ and conv_ prefixes.
Provider-owned parameter combinations pass through unchanged.
anthropic_messages_format
Classifies Anthropic Messages API requests and promotes format metadata.
prompt_enrich
Injects system or user messages into
OpenAI-compatible chat completion request bodies.
Static configured messages are prepended or appended
to the messages array. Uses BodyAccess::ReadWrite.
credential_injection
Per-cluster API key injection with client credential stripping. Supports inline values and environment variable sources.
openai_response_store
Persists non-streaming Responses API responses. See Response Store for details.
Key Files
apis/src/classifier/mod.rs: pure format classifierapis/src/openai/responses/mod.rs:ResponsesFormatFilterfilters/src/inference/model_to_header.rs:ModelToHeaderFilterfilters/src/prompt_enrich/: prompt enrichment filterapis/src/anthropic/: Anthropic Messages format filter