A2A Agent Card Routing
Routes agent card discovery requests to dedicated backends
Configurations in Praxis AI.
Routes agent card discovery requests to dedicated backends
Routes A2A requests by body-derived method, family, context ID, task ID, and streaming detection
Captures task and context ownership from SendMessage JSON responses and SendStreamingMessage / SubscribeToTask SSE responses, then routes follow-up requests back to the backend cluster that created the task or owns the context
Routes LLM API requests to different backends based on the model field in the JSON request body
8 configurations in Anthropic for Praxis AI.
Signs outbound requests to an AWS service (Bedrock, in this example) using Signature Version 4. Credentials are static, sourced from environment variables — see the module docs on Sigv4SignFilter for the planned OIDC/default-credential-chain follow-up
1 configurations in Azure for Praxis AI.
Acquires an Entra ID bearer token via the client-credentials grant and injects “Authorization: Bearer " on every proxied request to Azure OpenAI
Injects per-cluster API credentials into upstream requests and strips client-provided credentials to prevent forwarding
Pre-request balance check and post-response token usage reporting against an external metering service
Acquires an OAuth2 access token from the GCE/GKE metadata server (source: adc or metadata) and injects “Authorization: Bearer " on every proxied request to Vertex AI
Captures identity headers matching a prefix into filter metadata and strips them before forwarding upstream
1 configurations in Inference for Praxis AI.
Demonstrates every candidate capability and selection input handled by intelligent_route today
Use model identity to select the matching inference backend.
Demonstrates the intelligent_route management-path skip list: management and discovery endpoints (model listing, subscriptions, API-key management, health) bypass model resolution entirely, while inference paths are still resolved by model and fail closed on an unknown model
Routes MCP tools/call requests to the cluster that owns the requested tool, using the mcp.name metadata set by the mcp filter
Routes requests using a routing overlay file (routing-overlay.json) instead of inline YAML candidates. The overlay is rendered by the operator into a Kubernetes ConfigMap and projected as a volume mount
Routes JSON-RPC 2.0 requests to different backends based on the “method” field in the JSON request body
Screens every request body through Lakera Guard for content moderation before forwarding to the upstream
A real llm-d EPP or test processor returns the trusted x-gateway-destination-endpoint header
Rewrites publisher-ID body model values to the short model name for LLMISvc / KServe routing; the routing header (default X-Model) is left unchanged so routing can still use the publisher ID
Routes MCP requests by body-derived method and tool name
Configurable stateless MCP broker using the final MCP 2026-07-28 stateless profile
Routes LLM API requests to different backends based on the “model” field in the JSON request body
Evaluates incoming requests against a NeMo Guardrails service
Evaluates upstream responses against a NeMo Guardrails service
1 configurations in Openai for Praxis AI.
1 configurations in Payload Processing for Praxis AI.
Demonstrates the production boundary used when an external authenticator injects separate tenant and subject headers. state_owner consumes those assertions into an immutable internal owner and strips the inbound copies. project_state_owner_headers then recreates destination-specific headers from that normalized context
Injects system messages into OpenAI-compatible chat completion requests before forwarding to the upstream provider
This listener requires downstream mTLS. peer_identity_trust authenticates and authorizes the edge gateway before AI-owned x-ai-routing- fields can influence provider-local routing
Measures the elapsed time from request receipt to the first non-empty SSE body chunk and records a praxis_ai_ttft_seconds Prometheus histogram labeled by model
Extracts token usage from AI inference responses (streaming and non-streaming) and makes counts available to downstream filters via filter metadata as token.input, token.output, and token.total
Apply token-aware rate limiting to inference requests and reconcile reported usage.
Extends token-rate-limit.yaml with per-rule algorithm choice (ai#789 / praxis#551): each rule in rules: independently picks sliding_window or token_bucket, matched by a static header value. team-alpha gets an exact trailing-window budget; team-beta gets a continuously-refilling bucket
Extends token-rate-limit.yaml with graduated enforcement tiers (proposal S1, ai#881)
Inject Praxis-Token-Input, Praxis-Token-Output, and Praxis-Token-Total headers into downstream responses when token counts are available in filter metadata
1 configurations in Vertex for Praxis AI.