openai_doc_extract

Converts input_file content parts to input_text for backends that do not support input_file natively (e.g. vLLM, llm-d).
On this page

Converts input_file content parts to input_text for backends that do not support input_file natively (e.g. vLLM, llm-d).

Configuration

FieldTypeRequiredDescription
allow_pre_security_calloutboolnoAcknowledge that StreamBuffer body processing runs before header-phase security filters. Must be true.
max_rewritten_body_bytesintegernoMaximum size in bytes of the request body this filter produces after converting input_file parts to input_text (default 64 MiB). Raw request body size is governed by the pipeline’s body_limits, not this field.
max_content_bytesintegernoMaximum decoded content bytes per file (default 10 MiB).
max_file_referencesintegernoMaximum input_file parts processed per request (default 32).
max_total_text_bytesintegernoMaximum total extracted text bytes across all files in one request (default 64 MiB).
on_unsupportedcontinue | rejectnoBehavior when a file’s MIME type is not text-safe.

Examples

Example 1

filter: openai_doc_extract
allow_pre_security_callout: true

Example 2

filter: openai_doc_extract
allow_pre_security_callout: true
on_unsupported: continue
max_rewritten_body_bytes: 67108864
max_content_bytes: 10485760
max_file_references: 32
max_total_text_bytes: 67108864