openai_responses_compact

Summarizes conversation history when the token count exceeds a configured threshold.
On this page

Summarizes conversation history when the token count exceeds a configured threshold.

Configuration Notes

compact_threshold in context_management must be an integer of at least 1000. Invalid or missing compact_threshold values produce an invalid_request_error.

Compaction applies in two scenarios:

  • Rehydrated history - stored history loaded via previous_response_id or conversation. Only the stored history is summarized; the current turn is preserved.

  • Explicit compact - POST /v1/responses/compact with a required model and an inline input conversation and/or a previous_response_id. Loads any stored history, appends the inline input, summarizes the combined conversation, and returns a response.compaction object (with output and usage) per the OpenAI contract.

Direct input requests (full conversation in input with no stored history) skip reactive compaction because state.input == state.messages - there is no separable “current turn” to preserve after summarization. Requests without rehydrated history are released without compaction.

Praxis runs StreamBuffer body hooks before header-phase request filters. This filter therefore requires allow_pre_security_callout: true and should only be used behind an outer authentication and authorization boundary.

Configuration

FieldTypeRequiredDescription
allow_pre_security_calloutboolnoAllow summarization callouts from the StreamBuffer pre-read phase, before header-phase security filters execute. This must be explicitly enabled only when an outer trust boundary authenticates and authorizes requests before they reach this listener.
inference_urlstringyesURL of the inference backend for summarization calls. E.g., "http://localhost:11434/v1/chat/completions"
allow_private_inference_urlboolnoAllow the inference target to resolve to non-public addresses.
default_modelstringnoDefault model for summarization when not overridden in the request’s context_management.
tiktoken_encodingstringnoTiktoken encoding name for local token estimation of the conversation text.
summary_prefixstringnoPrefix prepended to the summary when translating compaction items to backend messages. Defaults to "[Previous conversation summary]\n\n".
timeout_msintegernoCallout timeout in milliseconds.
on_failureclosed | opennoFailure mode for the inference callout.
status_on_errorintegernoHTTP error status code (400..=599) to return when rejecting on error.

Examples

Example 1

filter: openai_responses_compact
allow_pre_security_callout: true
inference_url: "http://localhost:11434/v1/chat/completions"
allow_private_inference_url: true
default_model: llama3.2:1b

Example 2

filter: openai_responses_compact
allow_pre_security_callout: true
inference_url: "http://localhost:11434/v1/chat/completions"
allow_private_inference_url: true
default_model: gpt-4o-mini
tiktoken_encoding: cl100k_base
summary_prefix: "[Previous conversation summary]\n\n"
timeout_ms: 30000
on_failure: closed
status_on_error: 502