Anthropic Messages Replay Test Plan
On this page
Status: draft
Owner: franciscojavierarceo
What?
Define the Anthropic Messages API shapes Praxis should cover in replay, passthrough, and translation tests. The goal is not to reimplement Anthropic’s full request validator. Praxis should validate only the envelope and fields it owns locally, then preserve the request and response bodies as much as possible when passing them through.
This test plan extends the broader Messages API filter proposal in 00484_anthropic-messages-api-filters.md with concrete fixture and test coverage.
Sources
The test plan is based on:
- Anthropic’s Messages API documentation
- Anthropic’s tool-use documentation: overview, defining tools, handling tool calls, and the tool reference
- Anthropic’s extended thinking and stop reason documentation
- A local aggregate scan of Claude Code JSONL sessions under
~/.claude/projects
The local scan was used only for shape discovery. It should not be copied into fixtures verbatim unless the content is reviewed and sanitized first.
Local Session Inventory
The local Claude Code sessions are dominated by client-tool traffic and are a good source for realistic replay shapes:
- 114 JSONL session files and 10,300 records
- Message roles: 4,295 assistant records and 3,082 user records
- Content block types:
tool_use: 2,615tool_result: 2,615text: 1,491- string content: 434
thinking: 222image: 3
- Image source types:
base64only - Stop reasons:
tool_use: 2,948end_turn: 325stop_sequence: 17
- Common local tool names include
Bash,Read,Edit,Agent,TaskUpdate,Write,TaskCreate,Skill,WebFetch, andWebSearch. - Tool result content appears as plain strings,
is_errorstrings, lists of text blocks, and one list containing an image block.
The scan did not find server tool response shapes such as server_tool_use.
Those cases should be curated from public documentation or generated synthetic
fixtures.
Validation Policy
Praxis should keep Anthropic Messages validation intentionally narrow:
- Validate that request bodies are valid JSON objects before downstream filters depend on them.
- Classify and promote only proxy-owned routing facts such as protocol format, model, streaming mode, and whether tools are present.
- Preserve unknown request and response fields.
- Leave Anthropic-owned semantics to the backend, including model capability checks, role ordering, parameter ranges, tool schema validity, and unsupported feature combinations.
- When transforming to OpenAI Chat Completions, test only the fields Praxis actually maps or intentionally drops with an observable warning.
- When importing replay fixtures, preserve the original
source_recordsstructure. Generated replay requests and responses may be added next to the source records, but source records must not be modified to make the test pass.
For response passthrough, the expected behavior is pure passthrough minus Praxis-owned source, protocol, and transport metadata. Content blocks, generated JSON, tool calls, thinking blocks, usage fields, stop reasons, and provider metadata should not be rewritten on native Anthropic paths.
Test Matrix
Request Envelope and Classification
Cover these cases with unit or integration tests:
- Minimal Messages request with
model,max_tokens, and one user message. - String user content and array-based
textcontent. - Requests recognized by
/v1/messagespath oranthropic-versionheader even when body structure overlaps with OpenAI Chat Completions. stream: trueclassification and routing metadata.toolspresence setting the internal tools metadata.- Top-level
systemas a string and as text blocks. - Mid-conversation
systemmessages preserved when present. - Last-assistant prefill content preserved when present.
- Malformed JSON or non-object JSON rejected by the validate filter.
Content Blocks
Cover all content block shapes Praxis may see or transform:
- User and assistant
textblocks. - Image blocks with
source.typevalues:base64urlfile
- Image media types:
image/jpegimage/pngimage/gifimage/webp
thinkingandredacted_thinkingblocks in source records and responses.- Tool result nested content blocks:
textimagedocumentsearch_result
The current local sessions only show base64 images, so URL, file, document, and search-result fixtures should be curated.
Tool Definitions
Cover tool definitions that Anthropic accepts and Praxis may transform:
- User-defined tool with
name,description, andinput_schema. - Tool
namevalues using the documented alphanumeric, underscore, and hyphen shape. - Optional
input_examplespreserved on passthrough paths. - Optional properties preserved or intentionally handled:
cache_controlstrictdefer_loadingallowed_callers
tool_choicevariants:autoanytoolnone
- Parallel-tool controls such as
disable_parallel_tool_useand their OpenAI translation equivalent when supported. - Client-side tool names seen in local sessions, especially
Bash,Read,Edit, andWebFetch. - Anthropic-schema client tools from the public reference:
memorybashtext_editorcomputer
- Server tools from the public reference:
web_searchweb_fetchcode_executionadvisortool_searchmcp_toolset
Server-tool fixtures should be synthetic unless we capture a real local session that includes them.
Tool Call Lifecycle
Cover complete tool loops, not only isolated content blocks:
- Assistant message with a single
tool_useblock andstop_reason: "tool_use". - Assistant text followed by
tool_use. - Assistant message with multiple parallel
tool_useblocks. - User follow-up containing
tool_resultblocks immediately after the matching assistant tool request. tool_result.contentas:- string
- list of text blocks
- list containing an image block
- empty or omitted content
tool_result.is_error: true.- Mixed server and client tool blocks, such as
server_tool_useplus client-sidetool_use. - Pending server-tool continuations where the follow-up user message contains
only
tool_resultblocks. - Programmatic tool-calling metadata such as
callerandcontainerwhen a fixture is available.
Ordering requirements should be documented, but Praxis should enforce them only where a transformation would otherwise produce an invalid backend request.
Response and Stop Reasons
Cover each documented stop reason with either a passthrough fixture or a translation unit test:
end_turnmax_tokensstop_sequencetool_usepause_turnrefusalmodel_context_window_exceeded
The refusal case should include stop_details. Tool-use responses should
include usage fields and at least one content block with type: "tool_use".
Streaming
Streaming replay is deferred until the replay schema can represent SSE
responses. Unit and integration tests for anthropic_messages_to_chat_completions_stream should still
cover:
- Text
content_block_deltaevents. - Tool call input deltas.
- Usage deltas.
- Stop events and final message events.
- Partial UTF-8 across chunks.
- Mixed text and tool-use content in one response.
- Finish-reason mapping for tool use and max tokens.
Fixture Plan
Use sanitized local Claude Code sessions for:
- A basic text-only Messages exchange.
- A base64 image request.
- A client-tool cycle with
Bash,Read, orEdit. - A client-tool cycle with
is_error: true. - A
tool_resultwith list-based text content. - The existing list-based image tool result case.
- A response containing
thinkingfollowed by final text or tool use.
Use curated fixtures for:
- URL and file image sources.
- GIF and WebP media types if not captured locally.
- Document and search-result tool result blocks.
- Server tools and
server_tool_use. pause_turn,refusal, andmodel_context_window_exceeded.- SSE replay once the fixture schema supports it.
Every fixture should include enough assertions to prove the behavior that made it worth adding. A fixture that only returns HTTP 200 is insufficient unless the shape itself is the regression target.
Priority
P0
- Add a Claude Code tool-use replay fixture that preserves source records and verifies the generated request/response remains replayable.
- Add Anthropic-to-OpenAI translation coverage for
tool_useandtool_result, including backend-visible OpenAI tool calls and tool-role messages. - Add a thinking fixture that proves source records preserve thinking content even when translation drops unsupported thinking blocks for OpenAI backends.
- Add a tool-result error fixture.
- Add a list-content tool-result fixture.
P1
- Add URL and file image source coverage.
- Add stop-reason fixtures for
max_tokens,stop_sequence, andrefusal. - Add tool-choice translation tests for
auto,any,tool, andnone. - Add optional tool-definition property passthrough tests.
P2
- Add curated server-tool fixtures.
- Add
pause_turnandmodel_context_window_exceededfixtures. - Extend replay fixtures to represent streaming SSE responses.
- Add programmatic tool-calling metadata coverage if we capture or synthesize a representative request.
Open Questions
- Should replay fixtures grow a first-class SSE response format, or should SSE stay covered only by stream parser tests for now?
- Should
anthropic_validateremain envelope-only, or should Praxis validate a small set of transformation-owned invariants before converting to OpenAI Chat Completions? - How much should committed fixtures reflect Claude Code-specific tool names versus public Anthropic tool names?
- What sanitization workflow should we use for large local tool results so the original source record structure is preserved without leaking private local content?