Filter Reference

AI filters provided by Praxis AI. For base proxy filters (router, load balancer, headers, CORS, etc.), see the Praxis core filter reference.
On this page

AI filters provided by Praxis AI. For base proxy filters (router, load balancer, headers, CORS, etc.), see the Praxis core filter reference.

Provider APIs (praxis-ai-apis)

Anthropic

FilterDescription
anthropic_messages_formatClassifies Anthropic Messages API requests and promotes routing facts to headers, metadata, and filter results.
anthropic_messages_protocolNormalizes Anthropic Messages protocol headers for native backends.
anthropic_messages_to_chat_completionsTransforms Anthropic Messages API requests to Chat Completions-compatible request bodies and transforms compatible responses back. The name refers to the Chat Completions wire shape, not the OpenAI Responses API; any Chat Completions-compatible backend is a valid target, not only OpenAI.
anthropic_messages_to_chat_completions_streamTransforms streaming SSE responses between the Chat Completions and Anthropic Messages formats, processing each chunk as it arrives.
anthropic_validateValidates Anthropic Messages request bodies for proxy-owned JSON envelope requirements.
anthropic_web_searchExecutes server-owned WebSearch tool calls in an Anthropic Messages loop.

Azure

FilterDescription
openai_chat_completions_to_azureai_chat_completionsTransforms requests targeting Azure OpenAI deployments into standard Chat Completions-compatible form and normalizes responses back.

General

FilterDescription
callout_credentialsEstablishing filter that captures typed per-user callout secrets from ingress headers.
project_state_owner_headersProjects a normalized [StateOwner] into destination-specific HTTP headers.
state_ownerEstablishes a normalized [StateOwner] from trusted identity sources.

OpenAI

FilterDescription
openai_agentic_loopAgentic loop controller for the Responses API pipeline.
openai_client_tool_compatLowers rich client-owned tool declarations to private function tools for a function-only Responses backend and restores the typed items on the way back.
openai_conversationsHandles all /v1/conversations endpoints locally.
openai_doc_extractConverts input_file content parts to input_text for backends that do not support input_file natively (e.g. vLLM, llm-d).
openai_file_resolveResolves file_id and file_url references in Responses API input by fetching content from a Files API or remote URL via ApiClient and inlining the base64-encoded content in the provider-native field.
openai_file_search_calloutDispatches the loop owner’s pending file-search assignments against a vector store API compatible backend.
openai_mcp_dispatchExecutes MCP tool calls against upstream MCP servers within the Responses API agentic loop.
openai_mcp_tool_resolveResolves MCP tool entries from the Responses API tools array into concrete tool definitions by calling tools/list on each upstream MCP server.
openai_operationClassifies supported OpenAI operations from the request head.
openai_response_storePersists Responses API responses to the configured response store backend.
openai_responses_compactSummarizes conversation history when the token count exceeds a configured threshold.
openai_responses_formatClassifies AI API request bodies and promotes routing facts to headers, metadata, and filter results without mutating the body.
openai_responses_model_rewriteRewrites the model field in Responses and Chat Completions request bodies.
openai_responses_proxyRebuilds the request body from ResponsesState when present.
openai_responses_rehydrateValidates previous_response_id by fetching the stored response, confirming its status is "completed", and populating ResponsesState with the full conversation history (stored turns + current input).
openai_responses_requestProcesses a Responses request body once and initializes state.
openai_responses_validateValidates and enriches Responses API requests.
openai_stream_eventsComposes the current IRR execution into one logical Responses stream.
openai_tool_parseParses tool definitions and tool_choice from Responses API request bodies and promotes routing facts to metadata and filter results without mutating the body.
openai_web_searchWeb search filter for model-driven web_search_call dispatch.
responses_to_chat_completionsTranslates canonical Responses create requests for a Chat Completions backend.

Vertex

FilterDescription
openai_chat_completions_to_vertexai_geminiTransforms OpenAI Chat Completions requests into Vertex AI Gemini generateContent format and translates responses back.

Cross-Provider Filters (praxis-ai-filters)

Agentic

FilterDescription
a2aExtracts A2A protocol metadata from JSON-RPC request bodies and promotes method, family, task ID, streaming detection, and version to request headers, filter results, and durable metadata for routing.
mcpExtracts MCP protocol metadata from JSON-RPC request bodies and promotes method, tool/resource/prompt name, JSON-RPC kind, protocol version, and session presence to request headers/filter results; stores session ID in durable metadata.

AWS

FilterDescription
aws_sigv4_signSigns outbound requests to AWS services using SigV4.

Azure

FilterDescription
azure_adInjects an Azure AD (Entra ID) bearer token into outbound requests.

Callout

FilterDescription
http_calloutCalls an external HTTP service during request processing and feeds its response into branch-chain evaluation.

GCP

FilterDescription
gcp_adcInjects a GCP OAuth2 access token into outbound requests.

Guardrails

FilterDescription
ai_guardrailsCalls an external AI guardrail provider to evaluate request and response bodies. The provider determines whether content should be passed, blocked, or redacted.

Identity Guard

FilterDescription
identity_header_guardCaptures request headers matching a configured prefix into filter_metadata and removes them from the upstream request.

Inference

FilterDescription
llmisvc_model_provider_resolverRewrites publisher-ID body model values to the short model name for LLMISvc / KServe routing; the routing header is left unchanged.
model_to_headerPromotes the JSON "model" field from the request body to a request header.

Metering

FilterDescription
external_meteringIntegrates with an external metering service for pre-request balance checks and post-response token usage reporting.

Prompt Enrich

FilterDescription
prompt_enrichInjects statically configured messages into the messages array of OpenAI-compatible chat completion request bodies.

Routing

FilterDescription
credential_injectReplaces caller credentials with the upstream credential selected by intelligent_route or provider_route.
intelligent_routeSelects an upstream cluster from a site/capability descriptor by matching either an inference model name or MCP tool name.
provider_routeExact provider-local mapping from an authenticated intelligent routing selection to a private backend cluster.

Time To First Token

FilterDescription
time_to_first_tokenMeasures time-to-first-token for streaming AI responses.

Token Rate Limit

FilterDescription
token_rate_limitToken-denominated rate limiter: reserves an estimated cost at admission, reconciles against actual usage after the response completes. Evaluates an ordered list of rules, each with its own optional match condition, algorithm choice, and budget.

Token Usage

FilterDescription
stream_usage_injectInjects stream_options.include_usage = true into streaming OpenAI chat-completions requests so the upstream response contains token usage.
token_ceilingRejects provider-bound requests whose estimated serialized request body or requested output exceeds configured limits. Missing output-limit fields are rejected when max_output_tokens is configured; this strict behavior prevents providers from applying an unbounded model default. At least one of max_input_tokens or max_output_tokens is required, and max_body_bytes defaults to 2 MiB.
token_countExtracts token usage from AI inference responses and writes unified counts to [filter_metadata].
token_usage_headersInjects Praxis-Token-Input, Praxis-Token-Output, and Praxis-Token-Total headers into downstream responses when token usage data is present in [filter_metadata]. Also injects Praxis-Token-Status when usage capture failed (e.g. overflow), so an unavailable count is never silently indistinguishable from a genuine zero.