a2a
Extracts A2A protocol metadata from JSON-RPC request bodies and promotes method, family, task ID, streaming detection, and version to request headers, filter results, and durable metadata for routing.
Praxis AI registers AI-specific filters into the Praxis filter pipeline. The generated filter reference is the authoritative inventory of filter names, descriptions, and configuration documentation.
For pipeline execution, filter traits, body access, conditional execution, filter chains, and core filters, see the Praxis core filter documentation.
AI filters are organized across two crates:
apis/src/ Provider API integrations
anthropic/ Anthropic Messages API
openai/ OpenAI Responses and Conversations APIs
classifier/ Shared request classification
store/ Response and conversation persistence
filters/src/ Cross-provider behavior
agentic/ MCP and A2A
guardrails/ AI content guardrails
inference/ Model routing
prompt_enrich/ Prompt injection
token_usage/ Token counting and headers
Praxis AI also inherits all base proxy filters from Praxis core through
FilterRegistry::with_builtins(). This includes routing, load balancing,
headers, credential injection, JSON-RPC parsing, CORS, compression, IP ACLs,
and other general proxy behavior. Those filters are intentionally documented
in the Praxis core filter reference, not duplicated here.
In-tree AI filters are registered by
praxis_ai_filters::register_ai_filters. Downstream
consumers that only need AI filters (for example an
Envoy ExtProc) can depend on praxis-ai-filters without
the proxy crate:
use praxis_filter::FilterRegistry;
let mut registry = FilterRegistry::with_builtins();
praxis_ai_filters::register_ai_filters(&mut registry);
// Or: let registry = praxis_ai_filters::build_ai_registry();
praxis-ai-proxy builds the full registry in three
ownership layers:
let mut registry = FilterRegistry::with_builtins();
praxis_ai_filters::register_ai_filters(&mut registry);
register_external_filters(&mut registry); // proxy-only
This keeps core filters, in-tree AI filters, and auto-discovered extensions at clear ownership boundaries. External filter auto-discovery stays proxy-only.
Pipelines that use OpenAI store or rehydrate filters must also install the response-store extension:
pipeline.add_pipeline_extension(
Box::new(praxis_ai_apis::store::ResponseStoreRegistry::new()),
);
The AI proxy does this in server/src/pipelines.rs.
Other hosts (such as ExtProc) must do the same.
Extracts A2A protocol metadata from JSON-RPC request bodies and promotes method, family, task ID, streaming detection, and version to request headers, filter results, and durable metadata for routing.
Calls an external AI guardrail provider to evaluate request and response bodies. The provider determines whether content should be passed, blocked, or redacted.
Classifies Anthropic Messages API requests and promotes routing facts to headers, metadata, and filter results.
Normalizes Anthropic Messages protocol headers for native backends.
Transforms Anthropic Messages API requests to Chat Completions-compatible request bodies and transforms compatible responses back.
Transforms streaming SSE responses between the Chat Completions and Anthropic Messages formats, processing each chunk as it arrives.
Validates Anthropic Messages request bodies for proxy-owned JSON envelope requirements.
Executes server-owned WebSearch tool calls in an Anthropic Messages loop.
Signs outbound requests to AWS services using SigV4.
Injects an Azure AD (Entra ID) bearer token into outbound requests.
Establishing filter that captures typed per-user callout secrets from ingress headers.
Replaces caller credentials with the upstream credential selected by intelligent_route or provider_route.
Extend Praxis AI with custom filters using the Praxis core filter interfaces and registration model.
Integrates with an external metering service for pre-request balance checks and post-response token usage reporting.
AI filters provided by Praxis AI. For base proxy filters (router, load balancer, headers, CORS, etc.), see the Praxis core filter reference.
Injects a GCP OAuth2 access token into outbound requests.
Calls an external HTTP service during request processing and feeds its response into branch-chain evaluation.
Captures request headers matching a configured prefix into filter_metadata and removes them from the upstream request.
Selects an upstream cluster from a site/capability descriptor by matching either an inference model name or MCP tool name.
Rewrites publisher-ID body model values to the short model name for LLMISvc / KServe routing; the routing header is left unchanged.
Extracts MCP JSON-RPC method, tool, resource, prompt, version, and session metadata into filter-visible request fields.
Promotes the JSON “model” field from the request body to a request header.
Agentic loop controller for the Responses API pipeline.
Transforms requests targeting Azure OpenAI deployments into standard Chat Completions-compatible form and normalizes responses back.
Transforms OpenAI Chat Completions requests into Vertex AI Gemini generateContent format and translates responses back.
Lowers rich client-owned tool declarations to private function tools for a function-only Responses backend and restores the typed items on the way back.
Serve OpenAI Conversations API endpoints locally through the configured response store.
Converts input_file content parts to input_text for backends that do not support input_file natively (e.g. vLLM, llm-d).
Resolves file_id and file_url references in Responses API input by fetching content from a Files API or remote URL via ApiClient and inlining the base64-encoded content in the provider-native field.
Dispatches the loop owner’s pending file-search assignments against a vector store API compatible backend.
Executes MCP tool calls against upstream MCP servers within the Responses API agentic loop.
Resolves MCP tool entries from the Responses API tools array into concrete tool definitions by calling tools/list on each upstream MCP server.
Classifies supported OpenAI operations from the request head.
Persists Responses API responses to the configured response store backend.
Summarizes conversation history when the token count exceeds a configured threshold.
Classifies AI API request bodies and promotes routing facts to headers, metadata, and filter results without mutating the body.
Rewrites the model field in Responses and Chat Completions request bodies.
Rebuilds the request body from ResponsesState when present.
Validates previous_response_id by fetching the stored response, confirming its status is “completed”, and populating ResponsesState with the full conversation history (stored turns + current input).
Processes the Responses create request body once and initializes state.
Validates and enriches Responses API requests.
Composes the current IRR execution into one logical Responses stream.
Parses tool definitions and tool_choice from Responses API request bodies and promotes routing facts to metadata and filter results without mutating the body.
Web search filter for model-driven web_search_call dispatch.
Projects a normalized [StateOwner] into destination-specific HTTP headers.
Injects statically configured messages into the messages array of OpenAI-compatible chat completion request bodies.
Exact provider-local mapping from an authenticated intelligent routing selection to a private backend cluster.
Translates canonical Responses create requests for a Chat Completions backend.
Establishes a normalized [StateOwner] from trusted identity sources.
Measures time-to-first-token for streaming AI responses.
Extracts token usage from AI inference responses and writes unified counts to [filter_metadata].
Token-denominated rate limiter: reserves an estimated cost at admission, reconciles against actual usage after the response completes.
Injects Praxis-Token-Input, Praxis-Token-Output, and Praxis-Token-Total headers into downstream responses when token usage data is present in [filter_metadata].