Skip to content

[codex] Fix BYOK Agent tools, subagents and native todos - #485

Open
lawyer112 wants to merge 4 commits into
leookun:mainfrom
lawyer112:fix/deepseek-agent-tool-subagents
Open

lawyer112 wants to merge 4 commits into
leookun:mainfrom
lawyer112:fix/deepseek-agent-tool-subagents

Conversation

@lawyer112

@lawyer112 lawyer112 commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

  • Accept the current Cursor CLI's initial subagent message, which omits message_id, by using its stable run_id as that message's identity. Root requests without a message ID still fail, and explicit IDs remain unchanged.
  • Handle the two team metadata RPCs locally for the built-in BYOK identity; requests from real Cursor accounts continue to the upstream service.
  • Preserve the previous TodoWrite list in the completed Cursor call and send the full merged list in its success result. Cursor can then display additions and completion progress correctly. Reuse the existing checkpoint and todo merge logic; leave canonical model arguments and tool results unchanged.

Reproduction and verification

  • Before the initial change, DeepSeek V4.1 Flash could call TodoWrite, but Read and Glob returned unauthenticated, while Task/Explore failed with Cursor user message action has no message_id. The missing ID and stable child run ID were verified in the CLI's actual request trace.
  • Native IDE testing then found that successful TodoWrite calls always displayed Checked to-do list: the completed call's args and result contained the same updated list. Partial merge results also omitted unchanged todos. Three new wire regression tests cover additions, status-only merge patches with retained content/dependencies, and clearing the list.
  • cargo test -p cursor-server --lib: 164 passed.
  • cargo test -p cursor-server --test tool_round --test conversation_delivery --test connect_wire --test cursor_transport_lifecycle --quiet: 30 passed with the initial protocol fix.
  • cargo fmt --all -- --check, desktop npm run check, and release desktop/server builds passed. The desktop release build was repeated after the TodoWrite change.
  • Native Windows Cursor CLI using the configured deepseek/deepseek-v4.1-flash through the installed desktop build of this branch: TodoWrite, Glob, Read, Shell, and Task/Explore completed on 2026-09-30. All 9 provider calls returned HTTP 200. The root agent and Explore subagent independently read the same local file marker; Shell returned a distinct echo marker. The test used --print --force --trust, and success was checked in the run and tool records, not solely from the model's final answer.
  • The same CLI checks also passed with the configured Claude Opus 5 (Anthropic, 8 provider calls), GPT-6 Astra (OpenAI Responses, 9 calls), and Grok 4.6 (OpenAI Responses, 7 calls). All calls returned HTTP 200, and each Explore subagent used the selected parent model.
  • The final installed desktop build was tested in the native Windows Cursor 3.21.18 IDE with DeepSeek V4.1 Flash. Creating three pending todos displayed Added 3 to-dos; a status-only merge preserved all three contents and displayed completed/in-progress/pending rows in the expandable native to-do summary. Completing the remaining items displayed 3 of 3 To-dos Completed. A real Explore task showed its Completed card, opened an independent Sent by parent chat, and returned the fixture marker via an actual child Read call. Its final parent/child run made 6 provider calls, all HTTP 200 with the selected model. Run/tool records, native stored todo state, and the outgoing Cursor wire trace were checked alongside the UI.
  • A deployment follow-up found that the earlier bare Cargo release command left the management frontend in development mode, although the Agent service worked. Rebuilt using npm run tauri:build -- --no-bundle --ci -- --locked, then verified the installed management dashboard and the normal Windows desktop shortcuts. Its embedded HTML, JavaScript, and CSS all returned HTTP 200 with no Vite server running. Added the standalone desktop build command and frontend verification requirements to the README.

Known limits

  • The full server test suite currently fails at server/tests/prefix_stability.rs::every_captured_mode_owns_and_renders_its_runtime_template (expected <user_query> wrapper). The failing test and prompt assets are unchanged by this PR.
  • Live OpenAI Chat validation with the two configured Kimi routes is blocked by an upstream 403 stating that the current subscription has no Kimi Code access; neither route reached a tool call.
  • The two team metadata routes overlap with [codex] Fix BYOK agent routing and add Windows/macOS CLI launchers #447. This PR is independently based on main and focuses on Agent tool and subagent behavior.
  • Native IDE validation covers the above tools, to-do summary, and Explore task UI. Other Cursor Agent features and all configured model routes have not been exhaustively tested.

DeepSeek reasoning and tool-result continuation

  • Preserve canonical DeepSeek thinking when continuing a conversation through Chat Completions after it previously used Responses.
  • Include explicit empty reasoning_content for DeepSeek assistant tool calls that emitted no thinking; prefer native Chat replay state when present.
  • Document protocol selection and the required model reselection in Cursor after changing the endpoint.

Validation:

  • cargo test -p cursor-server provider::openai_chat::tests --locked: 8 passed, including four new regression cases.
  • npm run tauri:build -- --no-bundle --ci -- --locked: production desktop build succeeded with embedded frontend assets.
  • The recorded failing Responses request reproduced HTTP 400. The patched native Chat adapter accepted the same 254 projected history messages and 36 tools at reasoning_effort=max, with 279644 input tokens, and produced a tool call. Tools suggested by the recorded private conversation were not executed.
  • An isolated live test executed a read of its single authorized marker file, submitted the actual tool result, and verified the marker in the final answer across two provider calls with maximum reasoning.
  • The installed production build served its frontend assets and completed a fresh upstream request with HTTP 200.

Provider bodies, conversation contents, credentials, and local diagnostics are excluded from this PR.

@lawyer112 lawyer112 changed the title [codex] Fix BYOK CLI subagent startup and tool metadata [codex] Fix BYOK Agent tools, subagents and native todos Sep 30, 2026
Xiashangning pushed a commit to Xiashangning/cursor-byok that referenced this pull request Oct 1, 2026
From upstream PR 485 (leookun#485), two of its three parts:

- Subagent message_id fallback: Cursor CLI omits the initial subagent
  message ID. When subagent_type_name and run_id are present, fall back
  to the stable client run ID instead of rejecting the request.
  Adapted to develop's action() signature (suppressed parameter kept).

- TodoWrite projection: keep the pre-update todo list in the completed
  call's args and the merged list in the result, so Cursor renders the
  per-item change instead of a bare "Checked to-do list". Uses the
  resolved state already produced by validated_todo_write in the
  dispatcher instead of upstream's project_todo_completion recomputation
  in conversation output.

Dropped the PR's team metadata part (GetTeamRepos/GetTeamAdminSettings
routes, compatibility::team_configuration, is_local_path entries):
develop already implements it since 0778d01.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant