Agent protocol

An agent provider implements these JSON-RPC methods. Every method takes an object with agentId (the host adds it) plus the fields below. A JavaScript plugin registers an agent with api.agents.register (ui-plugins.md, Agents): the host calls the handler of the same name on it (list_options or listOptions), and emit(event) sends the agent/event notifications below. The official providers (Codex, Claude Code, OpenCode and ACP) are such plugins; see "Providers in JavaScript" below. No provider is a native program (process-plugins.md).

Method Params Result
agent/initialize AgentInfo
agent/list_options { workspace } { options: ConfigOption[] }
agent/list_commands { sessionId } { commands: SlashCommand[] }
agent/list_sessions { workspace } { sessions: [{ id, title?, createdAt?, updatedAt? }] }
agent/read_session { workspace, sessionId } { items: TranscriptItem[] }
agent/create_session { workspace, options, tools?, instructions? } { sessionId }
agent/resume_session { sessionId, workspace, options, tools?, instructions? } {}
agent/close_session { sessionId } {}
agent/prompt { sessionId, input: { blocks, itemId?, delivery?: { inputId, attemptId, intent: send | steer | queue } } } { runId, receipt?: { evidence: local_write | native_admission | consumption, nativeInputId? } }; itemId identifies the user row for rollback, while delivery identifies the exact delivery attempt
agent/cancel { sessionId, reason?: stop | steer } {}; simulated-steer cancellation must not manufacture native completion on timeout
agent/cancel_task { sessionId, taskId } {} — stops one subagent; only with capabilities.cancelTask
agent/set_option { sessionId, optionId, value } { options: ConfigOption[] } (full snapshot)
agent/respond_to_approval { approvalId, optionId } {}
agent/respond_to_question { questionId, answer: { values, cancelled } } {}
agent/authenticate { method } {}
agent/logout {}
agent/list_skills { workspace } { skills: [{ name, description, source, manualOnly }] } — only with capabilities.skills; a prompt then carries { type: "skill", name, input } blocks
agent/rollback { sessionId, itemId } {} — forget that user message and everything after it; only with capabilities.rollback
agent/fork_session { sessionId, workspace, options, tools?, instructions? } { sessionId } (the copy); only with capabilities.fork
agent/compact { sessionId } {} — the result arrives as a compacted event; only with capabilities.compact
agent/usage_limits { limits: UsageLimits | null }
agent/update {} — upgrade the agent's own program with the installer that owns it (AgentInfo.maintenance)

A method an agent does not implement fails with "does not answer"; the host treats that like the capability being off. An agent with no slash commands answers list_commands with an empty list.

tools are the plugin tools this session may call (MCP Tool shapes) and instructions is text the agent must follow on top of its own (plugin rules, and the code mode guide). Both come again on every create, resume and fork, because not every agent keeps them. Declare the tools in the agent's own way and pass each call to the host with the request host/tools.call { agentId, sessionId?, workspace?, name, input, callId? }; it answers an MCP CallToolResult and never fails the turn. See tools.md. An agent that takes extra tools only as an MCP server it connects to (OpenCode, ACP agents) gets one from the host: host/mcp.serve { agentId, sessionId?, workspace?, tools, url? } answers { url }, a streamable HTTP MCP server on the loopback address whose tools/call runs like host/tools.call (name workspace without sessionId when one server serves every session of a folder: the host then finds the calling chat by the running turn). Serving again with that url replaces what it serves; host/mcp.close { url } stops it, and it stops when the plugin unloads.

input.blocks are ContentBlocks: { type: "text", text }, { type: "image", mimeType, data } (base64), { type: "file_ref", path }.

Model choices may supply group (the provider's display name) and reasoningLevels (the reasoning choices valid for that model). If any model supplies levels, the catalog is per-model: an empty list means no reasoning control for that model. Keep the reasoning ConfigOption in the catalog even when the initially selected model has no levels, so draft chats can switch models before creating a session. The host resolves its choices against the selected model before applying stored effort or creating a session; incompatible effort becomes null (provider default). Agents with a single reasoning list can leave all model-level lists empty.

Any choice may supply aliases: earlier values that now mean that choice. When an agent stops offering a value in favor of another (Claude Code maps the CLI's opus to claude-opus-5-5[1m]), it lists the old value as an alias of the new choice, and the host maps a chat's stored value and remembered defaults to it instead of falling back to the catalog default.

Any choice may supply icon (inline SVG; currentColor follows the text), which the composer draws in the picker and on the pill while the choice is chosen, and transient: true for a choice that holds for its chat only: the host does not remember it for new chats, which open with the choice remembered before it. Claude Code's Ultracode effort level is both: it runs many agents per turn, and the CLI never keeps it either.

AgentInfo.promptKeywords lists words the agent reacts to in a prompt, { word, label, description }, such as Claude Code's ultracode, which turns that one turn into a dynamic workflow. The agent acts on the word itself; while the prompt contains it as a word of its own, the composer shows the label and description over the prompt box.

The chat model picker uses provider groups, searches names, IDs and groups, and stores favorites by agent and complete model ID in plugin storage.

Events

While a run is active the plugin sends the notification agent/event with an AgentEvent:

{ sessionId, runId?, taskId?, event: <kind>, ...fields }

taskId names the subagent the event belongs to (see Subagents below).

event Fields
text_delta itemId, text, `mode: append
reasoning_delta itemId, text, mode
tool_call_started ToolCall fields
tool_call_updated id, and any of status, title, kind, input, outputDelta, content, locations
approval id, title, toolCall?, `options: [{ id, name, kind: allow_once
question id, message?, url? (external HTTP(S) action), responseMode?: tool (default) or message, `fields: [{ id, label, description?, kind: text
plan `entries: [{ content, status: pending
usage usedTokens (what the context holds now, not a sum over the turn's requests), contextWindow? (the chat header shows the share left once it is known), costUsd?, inputTokens?, cachedInputTokens?, outputTokens?, reasoningTokens?
session_info title?
config_options options: [ConfigOption] (full snapshot)
commands commands: [SlashCommand]
notice `level: info
task TaskInfo (see Subagents)
run_started none; the envelope's runId names the new run
run_finished `outcome: { status: completed
input_rejected inputId, attemptId, reason; positive native non-admission, emitted before rejecting the prompt call; never for an uncertain transport write
input_consumed inputId, nativeInputId?; correlated native model pickup, not a locally buffered line or API admission alone
usage_limits windows?, resetCredits?, spendLimit?, recovery? (UsageRecovery, below)
usage_blocked recovery: UsageRecovery; run-scoped proven quota interruption, which may still be followed by native retries/grace

A run must end with exactly one run_finished.

A run starts when the host calls prompt. When the agent begins a turn that nobody prompted (for example, it answers a background task that ended), it sends run_started with a new runId first. The host then shows the chat as running, cancel reaches the turn, and the transcript is not replaced while it streams. A prompted run never sends run_started.

Steering, delivery evidence and usage recovery

capabilities.steer means a native continuation-boundary submission. Keep the active runId for such a steer; do not start a second turn while the first start is in flight. capabilities.simulatedSteer means the host may cancel and reserve a same-session replacement after actual completion and host settlement. These are different capabilities. Provider-specific mapping belongs in the provider, never shared host/kernel code. Explicit after-run queues stay host-held; a provider must not also enqueue them in a second native inbox.

An omitted receipt is weak evidence. A successful write is not proof of native admission or consumption. Include native correlation ids where supported, emit input_consumed only when the native protocol establishes pickup, and fence delayed completions/receipts against their run/input identities. Providers lacking that evidence leave it unknown, visible and non-replayable. Never advertise amendment or withdrawal of already-admitted input without a native operation.

UsageRecovery is { availability: allowed | blocked | unknown | unsupported, observedAt, source, reason, identity?, scope?, resetsAt? }. Identity must identify the effective native account/configuration, not just an agent name or utilization. Missing permission/identity/reset evidence stays unknown. Select the latest exhausted relevant reset; if any relevant window lacks a deadline, omit it. Warnings, purchased-overage permission, output-token caps, authentication errors, generic 429s, and a full usage meter alone are not quota interruption evidence.

Emit usage_blocked with the actual run id, without stopping tools or retries. Only a failed native run followed by host settlement can arm recovery. Fresh availability reads must validate the same identity/scope and preserve blocked evidence instead of replacing it with sparse/null updates. A provider without a native quota reporting contract exposes the limitation, not guessed reset times.

Subagents

An agent with capabilities.subagents reports every subagent it starts with a task event carrying the full TaskInfo snapshot, sent again whenever something about it changes:

{ id, title, status: queued|running|waiting|idle|completed|failed|cancelled,
  toolCallId?,     // the tool call that started it; the UI nests it there
  parentTaskId?,   // the subagent that started this one, at any depth
  name?, model?, effort?, prompt?,
  background?,     // it can outlive the turn that started it
  activity?,       // what it is doing now, such as the last tool
  summary?,        // latest progress while running, the result once done
  startedAt?, endedAt?, usage?: Usage, toolUses?,
  phase?,          // the workflow phase an agent of a workflow runs in
  workflow?: { phases: [{ title }], result?, logs? } }
  • Every event the subagent produces (text, reasoning, tools, approvals, questions, usage) carries its taskId in the envelope. The host keeps those items inside the task, at any depth.
  • approval and question events with a taskId stay pending on the chat (the chat is what waits); the host sets their taskId so the UI can say which subagent asks.
  • usage with a taskId is the subagent's; it never replaces the chat's context meter. plan, run_started and run_finished with a taskId are ignored: a subagent's plan and turns are its own.
  • idle is a subagent that finished its work but can be given more; waiting is one blocked on the user.
  • read_session returns each subagent as a task item ({ role: "task", task: TaskInfo, items }) at the place it started, with its own transcript inside, read from the agent's own record of that subagent.
  • cancel stops the run and every subagent in it. cancel_task stops one subagent and leaves the run going; the task then ends as cancelled.

Plugins show subagents too: a tool's ctx.subagent (tools.md) and a prompt a plugin runs through an agent with models.use (api.models.prompt) are task items of the chat under the tool call, with ids the host gives (sub-..., and sub-.../<id> for a subagent of such a prompt's session). The host answers cancel_task for them itself, and passes a nested one on as cancel_task of the hidden session. To a provider such a prompt is an ordinary session: create_session (with no tools), one prompt, its events, then close_session; its approvals and questions reach the user through the chat it shows in, and come back as respond_to_approval and respond_to_question like any other.

Workflows

A task with workflow set runs a script of agents in phases (Claude Code's dynamic workflows). It is one task, published under the tool call that started it; name is the script's own name, prompt its source, background true. Every agent it starts is a task of its own with the workflow as parentTaskId and its phase in phase, so a reader that knows nothing of workflows still sees a tree of subagents.

  • workflow.phases lists the phases in the script's order, including the ones no agent has reached; an agent's phase names one of them.
  • queued is an agent that was started but waits for an earlier phase or a free slot. It counts as work in hand, like running and waiting.
  • workflow.result is what the script returned (any JSON value) and workflow.logs what it logged; both arrive when the run ends. summary is the result as one line of text.
  • Agents of a workflow stop with it: cancel_task on the workflow stops the run; on one of its agents it fails.
  • The agents' own transcripts are events tagged with their task ids, like any subagent's. They can arrive later than the agent's status: Claude Code writes them to files the plugin follows.

ToolCall:

{ id, name, kind: read|edit|delete|move|search|execute|think|fetch|task|other,
  title, status: pending|running|completed|failed|cancelled, input,
  content: [ { type: "text", text }
           | { type: "diff", path, oldText?, newText?, diff? }
           | { type: "terminal", command, cwd?, output, exitCode? } ],
  locations: [{ path, line? }] }

TranscriptItem (history): { id, createdAt?, role: "user", blocks }, { role: "assistant", text }, { role: "reasoning", text }, { role: "tool", call }, { role: "plan", plan }, { role: "notice", text }.

Background questions

A question with responseMode: "message" does not pause the run. The host keeps it available across turn completion until answered or dismissed. Its fields use the same choices and free-text controls as ordinary questions. Answers become a normal user message (field answers in display order, separated by blank lines), steering an active run when the provider advertises capabilities.steer, or starting a new turn when idle. Dismissal only resolves the card. These questions do not call agent/respond_to_question; the host emits question_resolved after successful delivery. Pending cards are session state, not restored after app restart.

Providers in JavaScript

A provider is a sandboxed plugin like any other: it starts its agent's program through the broker, reads the program's output as streams, and emits agent events. plugins/codex/ is the reference for a program that speaks JSON-RPC; plugins/claude/ for one with a line protocol of its own (Claude Code's stream-json), plugin tools served as an MCP server over that protocol, and the agent's own records read from disk; plugins/opencode/ for a program that is an HTTP server (opencode serve, reached with api.net and a localhost:* grant, its events read as server-sent events) and plugin tools given through the host's loopback MCP server (host/mcp.serve); plugins/acp/ for a provider whose agents are only known at run time (the ACP registry, downloaded and cached in plugin storage, plus the user's custom.json in the plugin's data folder), each started through process.any (a path, npx or uvx), and for the client side of a protocol with requests in both directions (fs/*, terminal/*, permissions and elicitations the agent asks the client for).

  • Manifest. "main": "main.js", contributes.agents for the agents it always serves, and exactly the permissions it uses, each with its reason: agents.provide; process naming its CLI (and the installers it asks about updates, such as npm and brew); env for the variables it reads; net for a host it opens or fetches (localhost:* for a server it starts on a port of its choosing); fs.read for the folders where the agent keeps its own records (Claude Code's ~/.claude/**), and workspace when it reads project files.
  • Registration. api.agents.register(definition) in activate, with one handler per method above; it returns { emit }. The host lists a plugin's agents once, when it has loaded. A provider that learns of more agents later (accounts in its storage, which answers only once the plugin is connected, or a registry it downloads) registers them; the runtime tells the host (host/agents.changed), which lists them again. api.agents.unregister(id) takes one back the same way, and the host stops it (the ACP plugin drops a registry agent the registry no longer lists). A definition may carry icon (inline SVG) and description: the host shows them for an agent that has not been started yet (one the user has not turned on), until its initialize answers. When an agent learns after initialize that its answer is out of date (it needs a sign-in, which most ACP agents say only at session/new), the plugin calls host/agents.changed with { agentIds: [id] } and the host asks it again.
  • Handlers run in the plugin's runtime: they may await host calls and the program, but must not busy-loop (the 10 s limit, ui-plugins.md). prompt answers with the run id at once; the run goes on in events and ends with exactly one run_finished.
  • Plugin tools. Declare tools to the agent in its own way and pass each call with api.host.tools.call({ agentId, sessionId, name, input, callId }) (callHostTool in the SDK turns a failure into an error result). A plugin cannot listen on a port, so an agent that connects to MCP servers gets the host's: api.host.call("host/mcp.serve", { agentId, workspace, tools }) answers the url to give the agent (OpenCode: POST /mcp with a remote server; it names each tool convergence_<tool>; an ACP agent: an http entry of mcpServers in session/new, session/load and session/resume, only when its initialize answer says mcpCapabilities.http, else the chat gets a notice that the tools are not available). The address is on this machine's loopback, so an agent server on another computer cannot use it.
  • Instructions. instructions from create, resume and fork go to the agent as extra system text (Codex developerInstructions, OpenCode's system on each prompt). ACP has no system prompt, so the ACP plugin sends them as a marked text block ahead of a session's first prompt and leaves that block out of a replayed history.
  • Sign-in. A browser opens only because the user asked: api.openUrl works for a provider (an http or https link) while the user signs in to one of its agents from settings, that is while the host's agent/authenticate call is out; it resolves { opened: false, reason } for a link it kept closed, which the sign-in should report with the link. A sign-in link an agent asks for at any other time (an ACP URL elicitation, a URL an agent prints) goes to the chat as a notice the user follows.
  • Installed version. api.process.which(program) (on the wire host/process.which) answers { path, realPath } for a program the plugin may run; the real path proves which installer owns it (installerOf in the SDK).
  • Files. api.fs.read, list and stat answer within the fs.read grants. The broker checks the path a call resolves to, links followed, so a file reached through a link out of the granted folders is refused; read the agent's records best effort (a refusal or a missing file is "nothing there") rather than failing the call. A file an agent appends to can be followed by its size (stat) and read from the byte offset already seen (read(path, { encoding: "binary" })).
  • State. What a provider must remember across restarts (such as the host item each prompt became) goes in plugin storage (api.host.storage), keyed by session. A file the user edits goes in the plugin's data folder (api.paths.data, the data scope of its fs grants), such as the ACP plugin's custom.json. What a native provider (before the port) kept in <data dir>/agents/<folder>/ is moved there as legacy/ the first time the JavaScript plugin that replaced it loads; the Claude, Codex and OpenCode plugins read it once into their storage.
  • Ids. crypto.randomUUID() makes the version 4 UUIDs a CLI may require for a session id.

The provider SDK (plugins/sdk/)

The SDK ships with the official plugins. An official plugin imports it by its path beside the plugin (import { spawnJsonRpc } from "../sdk/process.js"): the host sends the plugin's files under <folder>/ and the SDK's under sdk/, so the import resolves in the runtime as it does for node --test on disk, and saving the SDK reloads the plugins that use it. Only an SDK the installation shipped (or an unchanged copy of it) is mounted into an official plugin.

Module What it has
jsonrpc.js JsonRpcConnection (requests with timeouts, cancellation by AbortSignal with an optional cancel notification, notifications, server requests answered through a Responder with ok, err or discard, a closed connection failing what is still out), lines and pump for newline-delimited JSON over a stream
process.js spawnJsonRpc(api, program, args, { cwd, env, name, defaultTimeout, onRequest, onNotification, onExit }): a child's stdio as a JSON-RPC peer, writes kept in order, the stderr tail kept for onExit, shutdown() (TERM, then KILL); spawnLines(api, program, args, { cwd, env, name, onLine, onExit }): the same for a child with a line protocol of its own (send(message) writes one JSON line); run(api, program, args, { env, timeout }) for a program's whole output
sse.js parseSse(chunks) for any body stream, sse(api, url, init) over api.net.fetch, sseJson
agent.js events.* builders for every event kind, agentEvent, outcome, newId, uuid (crypto.randomUUID), permissionMode (the four shared modes), callHostTool
mcp.js handleMcp(api, { agentId, sessionId, tools }, message): an MCP tool server (initialize, ping, tools/list, tools/call through host/tools.call) for agents that take extra tools only as MCP servers and relay the messages; MCP_SERVER (convergence, so the agent names a tool mcp__convergence__<tool>)
diff.js unifiedDiff(path, oldText, newText): a unified diff with three lines of context, for agents that report an edit as replaced text
maintenance.js which (api.process.which, nulls when refused), installerOf, npmPrefix, latestNpm, latestBrew, isNewer
testing.js fakes for node --test: fakeApi (a spawned program is a scripted FakePeer, host calls and process.which a function, fs as given, paths.data, net.fetch answered by onFetch with fakeResponses, openUrl recorded and answered by onOpenUrl, agents.register and unregister recorded), diskFs (an api.fs over the real disk, with the paths a grant would refuse), memoryStorage (plugin storage for onHost), FakeStream (also a streamed response body, such as server-sent events), settle

Tests: node --test plugins/sdk/*.test.mjs plugins/codex/*.test.mjs plugins/claude/*.test.mjs plugins/opencode/*.test.mjs plugins/acp/*.test.mjs. The live tests are cargo test -p convergence-host --test e2e -- --ignored: they build the plugin host and run real prompts through the plugins, among them a Codex turn that runs a shell command and calls a plugin tool, a Claude Code turn on Haiku that calls a plugin tool through its MCP control channel, and an OpenCode turn on its free big-pickle model that calls a plugin tool through the host's loopback MCP server (opencode_calls_a_plugin_tool; run it by name: it puts ~/.opencode/bin/opencode, or CONVERGENCE_E2E_OPENCODE, first on the login PATH), and an ACP turn (Kilo on a free model of its gateway, or CONVERGENCE_E2E_ACP_AGENT and CONVERGENCE_E2E_ACP_MODEL) that calls a plugin tool through the same loopback server (acp_calls_a_plugin_tool; acp_initializes_and_lists_its_options spends no usage).

Source: plugins/docs/agent-protocol.md in the Divergence repository, built with this site.