sub2api/backend/internal/pkg/apicompat/chatcompletions_responses_request_invariants_test.go
visa2 f10bca8155 refactor(apicompat): redesign the Codex Responses ↔ Chat Completions bridge
Codex CLI speaks the OpenAI Responses protocol (streaming, store:false), while
many upstreams (e.g. DeepSeek in thinking mode) only expose Chat Completions.
The bridge that translates between the two had grown field by field and leaned
on Go's serialization defaults, which both the Responses client (Codex) and the
Chat upstream reject in ways the official OpenAI endpoints tolerate.

Problems this fixes (all observed running Codex CLI against a DeepSeek upstream):
  - Streaming reasoning was never shown in the Codex TUI (the answer appeared
    with no visible thinking): reasoning deltas were emitted before the reasoning
    item was opened, so the strict client discarded them.
  - A tool-using turn could wedge the session into a "no response" state: the
    function_call stream was never closed (no function_call_arguments.done /
    output_item.done), so Codex never saw the tool call complete.
  - Parallel tool calls were rejected upstream (400/502): each function_call
    became its own assistant message, producing consecutive assistant messages
    with mismatched tool replies.
  - A tool turn was rejected with "reasoning_content in the thinking mode must be
    passed back": the reasoning that produced the tool call was dropped instead
    of being returned on the assistant message.
  - Items with no Chat equivalent (web_search_call, ...) and Codex's
    command-approval notice landed between an assistant tool_calls message and
    its tool reply, triggering "An assistant message with 'tool_calls' must be
    followed by tool messages responding to each 'tool_call_id'".
  - Interrupt/reconnect left an unanswered or dangling tool_call in the history,
    triggering the same 400.

The shared root cause is reliance on serialization defaults — omitempty dropping
protocol-required zero values, and unrecognized item types falling through a
generic path — rather than deliberately reproducing the target protocol. The
bridge is reworked into two explicit layers.

Request direction (Responses input -> Chat messages): a parse -> build ->
normalize pipeline.
  - reasoning_content is carried back on the assistant message that produced a
    tool call (DeepSeek thinking mode requires it to continue the same thought)
  - consecutive function_call items (parallel tool calls) are merged into a
    single assistant message's tool_calls array
  - item types with no Chat equivalent are skipped instead of leaking through a
    generic path
  - normalizeChatMessages is the single invariant gate: it guarantees every
    assistant tool_calls message is immediately followed by one tool reply per
    tool_call_id — reordering any intervening message (such as a command-approval
    notice) to after the replies, dropping unanswered tool_calls and orphan tool
    replies, and preserving bare passthrough tool messages.

Response direction (Chat SSE -> Responses SSE): ResponsesStreamEvent.MarshalJSON
constructs each streamed event explicitly so protocol-required fields are always
present (output_index/content_index/summary_index at 0, message content:[],
reasoning summary:[], function_call call_id/name/arguments, output_text part
text/annotations/logprobs). This is a single source of truth that removes any
post-hoc JSON patching. Reasoning is emitted as its own output item, opened
before its deltas, and tool calls are fully closed
(function_call_arguments.done + output_item.done with complete arguments).

Tests cover request-direction message invariants against golden Codex request
shapes (parallel calls, unknown items, intervening messages, partial/dangling
calls), per-event wire completeness, and streaming lifecycle ordering.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-31 16:14:58 +08:00

188 lines
7.5 KiB
Go

package apicompat
import (
"encoding/json"
"testing"
"github.com/stretchr/testify/require"
)
// assertChatInvariants enforces the DeepSeek / OpenAI Chat Completions message
// invariants that, when violated, surface as upstream 400s. Used to validate the
// request-direction converter against golden codex request shapes.
func assertChatInvariants(t *testing.T, messages []ChatMessage) {
t.Helper()
for i, m := range messages {
// Every assistant tool_calls message must be immediately followed by one
// tool message per tool_call_id, in order.
if len(m.ToolCalls) > 0 {
for j, tc := range m.ToolCalls {
k := i + 1 + j
require.Lessf(t, k, len(messages), "tool_call %s has no following tool message", tc.ID)
require.Equalf(t, "tool", messages[k].Role, "tool_call %s not followed by a tool message", tc.ID)
require.Equalf(t, tc.ID, messages[k].ToolCallID, "tool reply order mismatch for %s", tc.ID)
}
}
// No two consecutive assistant messages.
if i > 0 && m.Role == "assistant" && messages[i-1].Role == "assistant" {
t.Fatalf("consecutive assistant messages at %d", i)
}
// No orphan tool replies.
if m.Role == "tool" {
require.NotEmptyf(t, m.ToolCallID, "tool message without tool_call_id at %d", i)
}
}
}
func convertGolden(t *testing.T, input string) []ChatMessage {
t.Helper()
msgs, err := responsesInputToChatMessages("You are a helpful assistant.", json.RawMessage(input))
require.NoError(t, err)
return msgs
}
// Golden sample: a single tool-call turn (codex runs one shell/curl command),
// the shape that produced the original "no response" / 400.
func TestGolden_SingleToolCall(t *testing.T) {
msgs := convertGolden(t, `[
{"type":"message","role":"user","content":[{"type":"input_text","text":"latest sha?"}]},
{"type":"reasoning","summary":[{"type":"summary_text","text":"need to run curl"}]},
{"type":"function_call","call_id":"call_a","name":"exec_command","arguments":"{\"cmd\":\"curl x\"}"},
{"type":"function_call_output","call_id":"call_a","output":"deadbeef"}
]`)
assertChatInvariants(t, msgs)
// reasoning_content must ride on the assistant tool-call message.
var asst *ChatMessage
for i := range msgs {
if len(msgs[i].ToolCalls) > 0 {
asst = &msgs[i]
}
}
require.NotNil(t, asst)
require.Equal(t, "need to run curl", asst.ReasoningContent)
}
// Golden sample: parallel tool calls (codex runs git log + git tag at once).
func TestGolden_ParallelToolCalls(t *testing.T) {
msgs := convertGolden(t, `[
{"type":"message","role":"user","content":[{"type":"input_text","text":"features?"}]},
{"type":"reasoning","summary":[{"type":"summary_text","text":"inspect repo"}]},
{"type":"function_call","call_id":"c0","name":"exec_command","arguments":"{\"cmd\":\"git log\"}"},
{"type":"function_call","call_id":"c1","name":"exec_command","arguments":"{\"cmd\":\"git tag\"}"},
{"type":"function_call_output","call_id":"c0","output":"log"},
{"type":"function_call_output","call_id":"c1","output":"tags"}
]`)
assertChatInvariants(t, msgs)
// Both parallel calls share ONE assistant message.
var toolMsgs int
for _, m := range msgs {
if len(m.ToolCalls) == 2 {
require.Equal(t, "c0", m.ToolCalls[0].ID)
require.Equal(t, "c1", m.ToolCalls[1].ID)
}
if m.Role == "tool" {
toolMsgs++
}
}
require.Equal(t, 2, toolMsgs)
}
// Golden sample: an unknown item type (web_search_call from a 联网查询) sitting
// between a function_call and its output must not break tool↔reply adjacency.
func TestGolden_UnknownItemBetweenToolCallAndOutput(t *testing.T) {
msgs := convertGolden(t, `[
{"type":"message","role":"user","content":[{"type":"input_text","text":"search"}]},
{"type":"reasoning","summary":[{"type":"summary_text","text":"let me search"}]},
{"type":"function_call","call_id":"c0","name":"exec_command","arguments":"{}"},
{"type":"web_search_call","id":"ws_1","status":"completed","action":{"type":"search","query":"x"}},
{"type":"function_call_output","call_id":"c0","output":"result"}
]`)
assertChatInvariants(t, msgs)
}
// Sequential tool calls (a tool reply between two calls) must stay in distinct
// assistant messages.
func TestRequest_SequentialToolCallsStaySeparate(t *testing.T) {
msgs := convertGolden(t, `[
{"type":"function_call","call_id":"c1","name":"exec","arguments":"{}"},
{"type":"function_call_output","call_id":"c1","output":"r1"},
{"type":"function_call","call_id":"c2","name":"exec","arguments":"{}"},
{"type":"function_call_output","call_id":"c2","output":"r2"}
]`)
assertChatInvariants(t, msgs)
assistants := 0
for _, m := range msgs {
if len(m.ToolCalls) == 1 {
assistants++
}
}
require.Equal(t, 2, assistants)
}
// Golden sample: codex injects a message (e.g. an "Approved command prefix
// saved" notice) between a function_call and its output. The intervening message
// must be moved after the tool reply so the assistant tool_calls is immediately
// followed by its reply.
func TestGolden_MessageBetweenToolCallAndOutput(t *testing.T) {
msgs := convertGolden(t, `[
{"type":"message","role":"user","content":[{"type":"input_text","text":"do it"}]},
{"type":"reasoning","summary":[{"type":"summary_text","text":"run cmd"}]},
{"type":"function_call","call_id":"A","name":"exec","arguments":"{}"},
{"type":"message","role":"developer","content":[{"type":"input_text","text":"Approved command prefix saved"}]},
{"type":"function_call_output","call_id":"A","output":"ok"}
]`)
assertChatInvariants(t, msgs)
// The assistant tool_calls message is immediately followed by its tool reply.
for i, m := range msgs {
if len(m.ToolCalls) > 0 {
require.Equal(t, "tool", msgs[i+1].Role)
require.Equal(t, "A", msgs[i+1].ToolCallID)
}
}
}
// Golden sample: a parallel tool call where one sibling's output is missing
// (codex interrupted/reconnected mid-execution). The unanswered tool_call must
// be dropped so the remaining assistant tool_calls are all answered.
func TestGolden_PartialParallelDropsUnansweredCall(t *testing.T) {
msgs := convertGolden(t, `[
{"type":"message","role":"user","content":[{"type":"input_text","text":"q"}]},
{"type":"reasoning","summary":[{"type":"summary_text","text":"r"}]},
{"type":"function_call","call_id":"A","name":"exec","arguments":"{}"},
{"type":"function_call","call_id":"B","name":"exec","arguments":"{}"},
{"type":"function_call_output","call_id":"A","output":"oa"}
]`)
assertChatInvariants(t, msgs)
for _, m := range msgs {
for _, tc := range m.ToolCalls {
require.NotEqual(t, "B", tc.ID, "unanswered tool_call B should have been dropped")
}
}
}
// Golden sample: a dangling tool_call at the end of the history (no output yet).
// The assistant message holding only that call must be dropped entirely.
func TestGolden_DanglingToolCallDropped(t *testing.T) {
msgs := convertGolden(t, `[
{"type":"message","role":"user","content":[{"type":"input_text","text":"q"}]},
{"type":"reasoning","summary":[{"type":"summary_text","text":"r"}]},
{"type":"function_call","call_id":"A","name":"exec","arguments":"{}"}
]`)
assertChatInvariants(t, msgs)
for _, m := range msgs {
require.Empty(t, m.ToolCalls, "dangling unanswered tool_call should have been dropped")
}
}
// normalizeChatMessages drops an orphan tool reply whose tool_call was never
// announced.
func TestNormalize_DropsOrphanToolReply(t *testing.T) {
msgs := convertGolden(t, `[
{"type":"message","role":"user","content":[{"type":"input_text","text":"q"}]},
{"type":"function_call_output","call_id":"ghost","output":"orphan"}
]`)
for _, m := range msgs {
require.NotEqualf(t, "tool", m.Role, "orphan tool reply should have been dropped")
}
}