All configurations

33 configurations in All configurations for Praxis AI.

Configurations in Praxis AI.

Subcategories

Configurations

  • A2A Agent Card Routing — Routes agent card discovery requests to dedicated backends
  • A2A Classifier Routing — Routes A2A requests by body-derived method, family, context ID, task ID, and streaming detection
  • A2A Task Routing — Captures task and context ownership from SendMessage JSON responses and SendStreamingMessage / SubscribeToTask SSE responses, then routes follow-up requests back to the backend cluster that created the task or owns the context
  • Ai Inference Body Based Routing — Routes LLM API requests to different backends based on the model field in the JSON request body
  • Aws Sigv4 — Signs outbound requests to an AWS service (Bedrock, in this example) using Signature Version 4. Credentials are static, sourced from environment variables — see the module docs on Sigv4SignFilter for the planned OIDC/default-credential-chain follow-up
  • Azure Ad — Acquires an Entra ID bearer token via the client-credentials grant and injects “Authorization: Bearer " on every proxied request to Azure OpenAI
  • Credential Injection — Injects per-cluster API credentials into upstream requests and strips client-provided credentials to prevent forwarding
  • External Metering — Pre-request balance check and post-response token usage reporting against an external metering service
  • Gcp Adc — Acquires an OAuth2 access token from the GCE/GKE metadata server (source: adc or metadata) and injects “Authorization: Bearer " on every proxied request to Vertex AI
  • Identity Header Guard — Captures identity headers matching a prefix into filter metadata and strips them before forwarding upstream
  • Intelligent Route All Capabilities — Demonstrates every candidate capability and selection input handled by intelligent_route today
  • Intelligent Route Inference — Use model identity to select the matching inference backend.
  • Intelligent Route Management Skip — Demonstrates the intelligent_route management-path skip list: management and discovery endpoints (model listing, subscriptions, API-key management, health) bypass model resolution entirely, while inference paths are still resolved by model and fail closed on an unknown model
  • Intelligent Route Mcp — Routes MCP tools/call requests to the cluster that owns the requested tool, using the mcp.name metadata set by the mcp filter
  • Intelligent Route Overlay — Routes requests using a routing overlay file (routing-overlay.json) instead of inline YAML candidates. The overlay is rendered by the operator into a Kubernetes ConfigMap and projected as a volume mount
  • Json Rpc Routing — Routes JSON-RPC 2.0 requests to different backends based on the “method” field in the JSON request body
  • Lakera Guard — Screens every request body through Lakera Guard for content moderation before forwarding to the upstream
  • Llmd Ext Proc Routing — A real llm-d EPP or test processor returns the trusted x-gateway-destination-endpoint header
  • Llmisvc Model Provider Resolver — Rewrites publisher-ID body model values to the short model name for LLMISvc / KServe routing; the routing header (default X-Model) is left unchanged so routing can still use the publisher ID
  • Mcp Classifier Routing — Routes MCP requests by body-derived method and tool name
  • Mcp Stateless Broker — Configurable stateless MCP broker using the final MCP 2026-07-28 stateless profile
  • Model To Header Routing — Routes LLM API requests to different backends based on the “model” field in the JSON request body
  • Nemo Guardrails Response — Evaluates upstream responses against a NeMo Guardrails service
  • Nemo Guardrails — Evaluates incoming requests against a NeMo Guardrails service
  • Project State Owner Headers — Demonstrates the production boundary used when an external authenticator injects separate tenant and subject headers. state_owner consumes those assertions into an immutable internal owner and strips the inbound copies. project_state_owner_headers then recreates destination-specific headers from that normalized context
  • Prompt Enrichment — Injects system messages into OpenAI-compatible chat completion requests before forwarding to the upstream provider
  • Provider Route — This listener requires downstream mTLS. peer_identity_trust authenticates and authorizes the edge gateway before AI-owned x-ai-routing- fields can influence provider-local routing
  • Time To First Token — Measures the elapsed time from request receipt to the first non-empty SSE body chunk and records a praxis_ai_ttft_seconds Prometheus histogram labeled by model
  • Token Counting — Extracts token usage from AI inference responses (streaming and non-streaming) and makes counts available to downstream filters via filter metadata as token.input, token.output, and token.total
  • Token Rate Limit Mixed Algorithms — Extends token-rate-limit.yaml with per-rule algorithm choice (ai#789 / praxis#551): each rule in rules: independently picks sliding_window or token_bucket, matched by a static header value. team-alpha gets an exact trailing-window budget; team-beta gets a continuously-refilling bucket
  • Token Rate Limit Soft Tiers — Extends token-rate-limit.yaml with graduated enforcement tiers (proposal S1, ai#881)
  • Token Rate Limit — Apply token-aware rate limiting to inference requests and reconcile reported usage.
  • Token Usage Headers — Inject Praxis-Token-Input, Praxis-Token-Output, and Praxis-Token-Total headers into downstream responses when token counts are available in filter metadata

A2A Agent Card Routing

Routes agent card discovery requests to dedicated backends

A2A Classifier Routing

Routes A2A requests by body-derived method, family, context ID, task ID, and streaming detection

A2A Task Routing

Captures task and context ownership from SendMessage JSON responses and SendStreamingMessage / SubscribeToTask SSE responses, then routes follow-up requests back to the backend cluster that created the task or owns the context

Ai Inference Body Based Routing

Routes LLM API requests to different backends based on the model field in the JSON request body

Anthropic

8 configurations in Anthropic for Praxis AI.

Aws Sigv4

Signs outbound requests to an AWS service (Bedrock, in this example) using Signature Version 4. Credentials are static, sourced from environment variables — see the module docs on Sigv4SignFilter for the planned OIDC/default-credential-chain follow-up

Azure

1 configurations in Azure for Praxis AI.

Azure Ad

Acquires an Entra ID bearer token via the client-credentials grant and injects “Authorization: Bearer " on every proxied request to Azure OpenAI

Credential Injection

Injects per-cluster API credentials into upstream requests and strips client-provided credentials to prevent forwarding

External Metering

Pre-request balance check and post-response token usage reporting against an external metering service

Gcp Adc

Acquires an OAuth2 access token from the GCE/GKE metadata server (source: adc or metadata) and injects “Authorization: Bearer " on every proxied request to Vertex AI

Identity Header Guard

Captures identity headers matching a prefix into filter metadata and strips them before forwarding upstream

Inference

1 configurations in Inference for Praxis AI.

Intelligent Route All Capabilities

Demonstrates every candidate capability and selection input handled by intelligent_route today

Intelligent Route Inference

Use model identity to select the matching inference backend.

Intelligent Route Management Skip

Demonstrates the intelligent_route management-path skip list: management and discovery endpoints (model listing, subscriptions, API-key management, health) bypass model resolution entirely, while inference paths are still resolved by model and fail closed on an unknown model

Intelligent Route Mcp

Routes MCP tools/call requests to the cluster that owns the requested tool, using the mcp.name metadata set by the mcp filter

Intelligent Route Overlay

Routes requests using a routing overlay file (routing-overlay.json) instead of inline YAML candidates. The overlay is rendered by the operator into a Kubernetes ConfigMap and projected as a volume mount

Json Rpc Routing

Routes JSON-RPC 2.0 requests to different backends based on the “method” field in the JSON request body

Lakera Guard

Screens every request body through Lakera Guard for content moderation before forwarding to the upstream

Llmd Ext Proc Routing

A real llm-d EPP or test processor returns the trusted x-gateway-destination-endpoint header

Llmisvc Model Provider Resolver

Rewrites publisher-ID body model values to the short model name for LLMISvc / KServe routing; the routing header (default X-Model) is left unchanged so routing can still use the publisher ID

Mcp Classifier Routing

Routes MCP requests by body-derived method and tool name

Mcp Stateless Broker

Configurable stateless MCP broker using the final MCP 2026-07-28 stateless profile

Model To Header Routing

Routes LLM API requests to different backends based on the “model” field in the JSON request body

Nemo Guardrails

Evaluates incoming requests against a NeMo Guardrails service

Nemo Guardrails Response

Evaluates upstream responses against a NeMo Guardrails service

Openai

1 configurations in Openai for Praxis AI.

Payload Processing

1 configurations in Payload Processing for Praxis AI.

Project State Owner Headers

Demonstrates the production boundary used when an external authenticator injects separate tenant and subject headers. state_owner consumes those assertions into an immutable internal owner and strips the inbound copies. project_state_owner_headers then recreates destination-specific headers from that normalized context

Prompt Enrichment

Injects system messages into OpenAI-compatible chat completion requests before forwarding to the upstream provider

Provider Route

This listener requires downstream mTLS. peer_identity_trust authenticates and authorizes the edge gateway before AI-owned x-ai-routing- fields can influence provider-local routing

Time To First Token

Measures the elapsed time from request receipt to the first non-empty SSE body chunk and records a praxis_ai_ttft_seconds Prometheus histogram labeled by model

Token Counting

Extracts token usage from AI inference responses (streaming and non-streaming) and makes counts available to downstream filters via filter metadata as token.input, token.output, and token.total

Token Rate Limit

Apply token-aware rate limiting to inference requests and reconcile reported usage.

Token Rate Limit Mixed Algorithms

Extends token-rate-limit.yaml with per-rule algorithm choice (ai#789 / praxis#551): each rule in rules: independently picks sliding_window or token_bucket, matched by a static header value. team-alpha gets an exact trailing-window budget; team-beta gets a continuously-refilling bucket

Token Rate Limit Soft Tiers

Extends token-rate-limit.yaml with graduated enforcement tiers (proposal S1, ai#881)

Token Usage Headers

Inject Praxis-Token-Input, Praxis-Token-Output, and Praxis-Token-Total headers into downstream responses when token counts are available in filter metadata

Vertex

1 configurations in Vertex for Praxis AI.