Official

codex

Codex agent provider: runs the Codex CLI's app-server for each account.

The app opens the listing; nothing installs until an agent in your Plugins workspace has read the files and you enable the plugin. In a terminal: cvg install convergence/codex@0.2.0

Permissions in 0.2.0

  • Provide agents agents.provideMediumAdds agents to the app.Provide the Codex agent and pass its tool calls to plugin tools
  • Run named programs processMediumStarts the listed programs.Run the Codex CLI, and ask or tell the installer that owns it (npm or Homebrew) about a newer versionPrograms: codexnpmbrew
  • Environment variables envMediumReads the listed environment variables.Find each account's Codex home, and pass extra app-server arguments set in CODEX_ARGSVariables: HOMECODEX_HOMECODEX_ARGS
  • Read files fs.readMediumReads files in the listed places.Read, once, the accounts the previous Codex provider keptPlaces: its own data folder

Files

NOTES.md26.1 KB
# Codex provider plugin — decisions and protocol notes

Backend: `codex app-server` (stdio, newline JSON-RPC 2.0), CLI 0.160.0.
The original mapping and streaming behaviour were checked on 0.155.1.
Steering, input correlation and recovery fields were rechecked with installed
0.160.0 `codex app-server generate-ts --experimental` on 2026-10-05. Read-only
native account/usage calls confirmed the identity and availability fields;
an ephemeral native shell turn confirmed expected-turn mismatch rejects with
`-32600` ("expected active turn id … but found …"). No model prompt was sent;
the new scheduling races are tested with scripted peers, not a live model turn.

## Layout

The plugin is sandboxed TypeScript (`"main": "main.ts"`), with scoped Effect handlers. It reaches the
CLI only through the broker (`api.process`), and speaks JSON-RPC to it
with the provider SDK (`../sdk/`, see `plugins/docs/agent-protocol.md`).
It replaced a Rust program, deleted with the other native providers; what
that program kept is read once from `legacy/` (`legacy.ts`).

| File           | Role                                                                                         |
| -------------- | -------------------------------------------------------------------------------------------- |
| `main.ts`      | `activate`: one agent per configured account                                                 |
| `instances.ts` | the accounts (plugin storage key `instances`)                                                |
| `legacy.ts`    | the Rust plugin's `instances.json`, read once into storage                                   |
| `map.ts`       | pure item-to-protocol mapping, and the thread helpers of the v2 schema (all unit-tested)     |
| `params.ts`    | options, `thread/start` / `resume` / `fork` / `turn/start` params, the prompt as `UserInput` |
| `agent.ts`     | connection, session and run state, every agent method                                        |
| `icon.ts`      | `icon.svg` as a string (the runtime strips TypeScript and imports JSON)                      |
| `*.test.ts`    | `node --test`; `agent.test.ts` drives the agent against a scripted app-server                |

## Connection

One app-server process per account, started lazily on first use and
restarted transparently if it died (`CodexAgent.client`; concurrent first
calls share one start). The broker finds `codex` on the login PATH
(`process: ["codex"]`); a path in `CODEX_PATH` is not honoured any more,
since a path needs `process.any`. Handshake is `initialize` with
`clientInfo { name: "convergence" }` and
`capabilities { experimentalApi: true, requestAttestation: false }`, then
the `initialized` notification.

`experimentalApi` is on because `item/tool/requestUserInput` (the question
surface) and the `plan` item type are gated behind it. `requestAttestation`
is off, so `attestation/generate` never arrives.

## Sessions

A Divergence session id **is** a Codex thread id.

- `create_session` → `thread/start { cwd, model?, approvalPolicy?, sandbox? }`
- `resume_session` → `thread/resume { threadId, cwd, excludeTurns: true, … }`
  (`excludeTurns` keeps the resume cheap; history comes from `read_session`)
- `list_sessions` → `thread/list { cwd: <workspace>, limit: 100 }`.
  **`cwd` is mandatory in practice**: without it the app-server returns the
  threads of every workspace.
- `read_session` → `thread/read { includeTurns: true }`, mapped turn by
  turn into `TranscriptItem`s. The derived title (thread `name`, else
  `preview`, else the first user message, trimmed to one line) is pushed as
  a `SessionInfo` event because `read_session` itself only returns items.
- `close_session` → `thread/unsubscribe`, best effort.

## Runs

Legacy `prompt` returns a run id immediately and `turn/start` is awaited in the
background (no timeout, since a turn has no useful timeout). A prompt carrying
`delivery` awaits native admission and returns an input receipt as well.
The run ends on `turn/completed` — `completed`/`interrupted`/`failed` map to
`Completed`/`Cancelled`/`Failed{turn.error.message}`. `finishRun` takes the
run out of the session state, so a second signal for the same run is a
no-op and **exactly one `RunFinished` is emitted**. The paths that end a run
are: `turn/completed`, a failed `turn/start` request, `cancel`, and the app-server
exiting (the pump fails every live run when its stdout closes). `error`, even
with `willRetry: false`, is a notice, not a terminal turn signal.

Notifications route by `threadId`, with known retired native turn IDs fenced.
A completion for a different live turn cannot finish it. A differing ID alone
does not suppress child activity or a real `turn/started`. Autonomous turns in
known sessions emit `run_started` with a fresh run ID; child turns never finish
the parent run. Late start replies cannot retarget a successor.

**Steering:** a `prompt` sent while a run is active becomes
`turn/steer { expectedTurnId }` and returns the _existing_ run id, so the
two prompts share one run and one `RunFinished`. `capabilities.steer` is
true for this reason. While `turn/start` is in flight, steers wait for its
native identity rather than issuing another start. Native validation/precondition
errors (`-32600`, `-32601`, `-32602`) emit correlated `input_rejected`; a lost
write/reply, internal error or malformed response is uncertain, never auto-retried.
The host owns explicit queues; passing `delivery.intent: "queue"` is rejected.

`delivery.attemptId` is sent as `clientUserMessageId` on both start and steer.
The native reply is admission, not consumption. A matching `userMessage.clientId`
on `item/started` or `item/completed` emits `input_consumed` once, tagged with the
original run even if it arrives late. Correlation is process-local: a plugin
restart loses that lookup, so unknown input must remain held rather than replayed.
No per-input native edit, withdrawal or amendment capability is advertised.

## Event mapping

| Codex                                                                                   | Divergence                                                                          |
| --------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- |
| `item/agentMessage/delta`                                                               | `TextDelta` (append, `item_id` = `itemId`)                                          |
| `item/reasoning/textDelta`, `…/summaryTextDelta`                                        | `ReasoningDelta` (append)                                                           |
| `item/completed` `agentMessage` / `reasoning`                                           | `TextDelta`/`ReasoningDelta` **Replace**, only when no delta streamed for that item |
| `item/started` / `item/completed` tool items                                            | `ToolCallStarted` once, then `ToolCallUpdated`                                      |
| `item/commandExecution/outputDelta`, `item/fileChange/outputDelta`                      | `ToolCallUpdated { output_delta }`                                                  |
| `item/fileChange/patchUpdated`                                                          | `ToolCallUpdated { content: Diff[] }`                                               |
| `turn/plan/updated`                                                                     | `Plan`                                                                              |
| `thread/tokenUsage/updated`                                                             | `Usage`                                                                             |
| `thread/name/updated`                                                                   | `SessionInfo { title }`                                                             |
| `error`, `warning`                                                                      | `Notice` (error / warning)                                                          |
| `deprecationNotice`, `configWarning`, `guardianWarning`, `windows/worldWritableWarning` | `Notice` (warning)                                                                  |
| `model/rerouted`                                                                        | `Notice` (info), naming both models                                                 |
| `item/reasoning/summaryPartAdded`                                                       | a blank line, so summary parts do not run together                                  |
| `item/mcpToolCall/progress`                                                             | `ToolCallUpdated { output_delta }`                                                  |
| `contextCompaction` item                                                                | `Compacted`                                                                         |
| `collabAgentToolCall` spawn, `subAgentActivity` started, `thread/started` (spawned)     | `Task`, plus the child thread's own events tagged with `task_id`                    |
| `account/rateLimits/updated`                                                            | `UsageLimits`                                                                       |
| `serverRequest/resolved`                                                                | `ApprovalResolved` / `QuestionResolved`                                             |

Deltas are forwarded byte for byte; whitespace-only chunks are kept (the
live probe streamed `"P"`, `"ONG"`, `" "` as separate deltas). One
exception: reasoning whose whole text is `codex exec`'s "Reading
additional input from stdin..." is not reasoning and is left out, live
(the first text of an item is held while it could still be only that) and
in history.

Tool kinds come from Codex's structured `commandActions`, never from the
command string: all-`read` → `Read`; all read/`listFiles`/`search` →
`Search`; anything else, or no actions at all → `Execute`. `fileChange` →
`Edit` with one `ToolContent::Diff` per change carrying the unified `diff`
only (Codex never sends whole-file snapshots). `mcpToolCall` and
`dynamicToolCall` → `Other`, `webSearch` → `Fetch`, `imageView` → `Read`.

## Approvals and questions

`item/commandExecution/requestApproval` and
`item/fileChange/requestApproval` become `Approval` events. Both decision
enums share four values, so the option ids are the Codex decision strings
themselves: `accept` → AllowOnce, `acceptForSession` → AllowAlways,
`decline` → RejectOnce, `cancel` → RejectAlways. The JSON-RPC response is
withheld until `respond_to_approval` fires a oneshot; `cancel` resolves
every pending approval of the session with `"cancel"` so the app-server
unblocks, and a dropped sender also answers `"cancel"` rather than hanging.

`item/tool/requestUserInput` becomes a `Question` — one `QuestionField` per
Codex question, `Select` when it carries options and `Text` otherwise,
`allow_other` from `isOther`. The reply is rebuilt as
`{ answers: { <questionId>: { answers: [...] } } }`.

`mcpServer/elicitation/request` has three shapes. With
`_meta.codex_approval_kind: "mcp_tool_call"` (Computer Use and other MCP
tools) it is an `Approval`: `accept` (Allow once), `acceptForSession` when
`_meta.persist` is or holds `"session"`, `acceptAlways` when it holds
`"always"`, and `cancel`. The reply is `{ action: "accept", content: {},
_meta: null | { persist: "session" | "always" } }`, or `cancel`/`decline`
with `content: null`. Mode `url` is a `Question` with `url` and no fields
(HTTP(S) only; anything else is cancelled with an error notice), answered
`accept` or `cancel`. Modes `form`, `openai/form` and `openaiForm` are a
`Question` with one field per `requestedSchema` property (`boolean` →
Boolean, `array` → MultiSelect, `oneOf`/`enum` → Select with the titles,
anything else → Text; `required`, `title`, `description` carried over),
answered with typed content: `integer`/`number` inputs go back as JSON
numbers. Any other mode is refused with method-not-found.

Any other server request (`item/permissions/requestApproval`,
`account/chatgptAuthTokens/refresh`) is answered with a JSON-RPC
method-not-found error so the app-server never waits on us.

## Config options

Fetched live; nothing is hardcoded except the two closed enums noted below.

- `model` (Model) from `model/list`, paginated on `nextCursor`, hidden
  models dropped. Each choice carries **its own** `reasoning_levels` built
  from that model's `supportedReasoningEfforts` rows — no shared list.
- `reasoning` (Reasoning) — the efforts of the currently selected model.
  Changing the model clears the stored reasoning value.
- `permission_mode` (Approval) — the four shared modes, see below.
- `service_tier` (Other) — the speed tiers of the chosen model, from
  `serviceTiers`, preceded by Standard (`default`) as in t3code. Selecting
  Standard explicitly resets a previous Fast selection. Absent for models
  that offer no additional tiers.
- Defaults come from `config/read { cwd }` (`model`,
  `model_reasoning_effort`, `approval_policy`, `sandbox_mode`), which
  returns **snake_case** keys unlike the rest of the v2 API.

Chosen values are sent per turn (`model`, `effort`, `approvalPolicy`,
`sandboxPolicy`, `approvalsReviewer`, `serviceTier`) as well as at
`thread/start`, so changing an option mid-session takes effect on the next
prompt.

## Accounts (instances)

One Codex install can hold several logins, kept apart by `CODEX_HOME`. The
plugin serves one agent per account, described in its own storage under
the key `instances`. The Rust plugin read
`<data dir>/agents/codex/instances.json`; the first time this plugin loads
in its place the host moves that folder into the plugin's data folder
(`legacy/`), and `legacy.ts` adds its entries to the stored ones once
(`fs.read` of the `data` scope):

```json
[{ "id": "personal", "name": "Codex Personal", "home": "~/.codex_personal", "args": [] }]
```

The default account registers in `activate`; the stored ones register once
the storage answers; the runtime tells the host, which lists them. `HOME`, `CODEX_HOME` and `CODEX_ARGS` come from the login
environment (`env` grant).

The default account is always served as `codex`; each entry adds
`codex:<id>`. `AgentInfo.family` is `codex` for all of them and
`continuation_key` is the resolved home, so the UI can offer a compatible
account for an existing chat and refuse an incompatible one. Each account
gets its own app-server process, model catalog and session map.

`CODEX_ARGS` adds app-server arguments for every account, and an instance's
`args` adds them for one. This is how MCP servers reach Codex: the CLI
takes them as `-c mcp_servers.<name>.command=…` overrides, so one mechanism
covers every Codex setting instead of a flag per feature.

## Permissions

The four shared modes (`convergence_protocol::permission_mode`) replace the
old separate approval-policy and sandbox selects:

| Mode              | `approvalPolicy` | `sandbox`            | `approvalsReviewer` |
| ----------------- | ---------------- | -------------------- | ------------------- |
| Supervised        | `untrusted`      | `read-only`          | `user`              |
| Auto-accept edits | `on-request`     | `workspace-write`    | `user`              |
| Auto              | `on-request`     | `workspace-write`    | `auto_review`       |
| Full access       | `never`          | `danger-full-access` | `user`              |

`approvalsReviewer` is always sent, including when it is `user`: leaving it
out on resume keeps the thread's previous reviewer, so `auto_review` would
stay switched on after the user moved back to a stricter mode. A chat that
has not chosen a mode starts from the user's own `config/read` defaults.

## Subagents

Codex runs each subagent in its own thread; the task id is that thread id.
Checked against CLI 0.155.1 and real rollouts (v1 and v2 on disk).

**Registration (live).** A child is registered against the chat from any of:

- a v1 `collabAgentToolCall` with `tool: "spawnAgent"`: every
  `receiverThreadIds` entry; `prompt` (not `input`), `model`,
  `reasoningEffort` and the item id (`tool_call_id`) come from it;
- a v2 `subAgentActivity` with `kind: "started"` (`agentThreadId`,
  `agentPath`). Other kinds address a child that exists; inside a child,
  `interacted` with `agentPath: "/root"` is a message to its parent;
- `thread/started` whose `source.subAgent.thread_spawn.parent_thread_id` is
  a chat here or a known child. Review, compaction and guardian threads
  and strangers' children are ignored.

A child whose parent is itself a child gets the parent's root chat as its
session and `parent_task_id` = the middle child, so grandchildren reach the
host at any depth. Every event of a child thread is sent to the root chat
with `task_id` = the thread id (`emitFor`); the `Task` snapshot itself is
untagged and placed by `parent_task_id`. A collab call or activity naming a
child this process never saw (spawned before a restart) adopts it and
reads its thread record for names.

**Snapshot.** Title = first line of the prompt; v2 sends the instruction
encrypted (the child's `preview` is empty and its first message is an
encrypted `NEW_TASK`), so a v2 title is the task name from the path
(`/root/pipeline_eval_tools` → "pipeline eval tools"); last resort the
nickname. `name` = `agentNickname`, else `agentRole`. `summary` = the
child's latest agent message, replaced by the `agentsStates` message (the
final answer) when the parent's `wait` returns it; never the title.
Status: child turn started → Running, `thread/status/changed` active with
flags → Waiting, child turn completed → Idle (resumable), interrupted →
Cancelled, failed → Failed (error text as summary); `agentsStates`
completed/shutdown → Completed, errored/notFound → Failed, interrupted →
Cancelled (`pendingInit`/`running` snapshots are ignored); v2 activity
completed → Completed. A Completed child is not reopened by a later Idle.
Child `thread/tokenUsage/updated` reports `total.totalTokens` on the task
(cumulative) and never touches the chat's meter; child plans are tagged and
so ignored by the host. `tool_uses` counts the child's tool items and
`activity` is the running tool's title.

**Stop.** `cancel` interrupts every live turn in the chat's subtree first
(`turn/interrupt {threadId: child, turnId}`, 3 s each, in parallel, 10 s
total), then the chat's own turn, and answers `cancel` to every pending
approval of those threads. `cancel_task` (capability on) interrupts that
child and its own children and releases their approvals; the chat's run
keeps going.

**History.** `read_session` reads the thread (`thread/read` without turns,
then `thread/turns/list` pages; single-page `thread/read {includeTurns}`
only as a fallback), then reads every child one generation at a time (at
most 64 threads). `transcriptWithSubagents` puts a `Task` item (id
`task-<thread>`, as the host names a live task) at the spawn or `started`
activity, filled from the parent's items (prompt, model, effort, spawn id,
final `agentsStates` status and message) and the child's record and turns
(nickname, preview, turn times, last agent message as a fallback summary),
with the child's own transcript inside, recursively. `wait`, `sendInput`,
`closeAgent` and the other collab calls are bookkeeping and are left out.
`list_sessions` drops threads with `parentThreadId` or a `thread_spawn`
source.

## Rate limits

`account/rateLimits/read` gives the first snapshot (with
`excludeResetCreditDetails`, since the available count is enough to offer a
reset) and `account/rateLimits/updated` keeps it current. The `primary` and
`secondary` windows are labelled from `windowDurationMins`, so "5 hours"
and "Weekly" come from the data rather than from a guess. Only the window
that is actually full is marked blocked.

0.160.0's `GetAccountRateLimitsResponse` includes nullable `ordinaryUsageAllowed`
and `accountId`. Only the full read's native permission boolean determines
allowed/blocked availability; null is unknown, never inferred from a meter or
elapsed reset. Identity comes from the native snapshot account ID, checked against
`account/read.workspaceRouting.chatgptAccountId` when present (also a native
fallback if the snapshot ID is absent). A mismatch removes recovery authority.
Account changes, login/logout and process exit invalidate cached evidence and
fence in-flight reads. Email, instance name and CODEX_HOME are not account proof.

`account/rateLimits/updated` is sparse and has no identity or permission boolean:
it reports unknown availability and may reuse the last proven account identity.
`usageLimitExceeded` alone emits run-scoped `usage_blocked`, without finishing the
run or interrupting tools. Generic throttling/429, auth errors, warnings and full
meters do not. Reset evidence requires an explicit reached quota/spend-control
marker and every responsible deadline; use the latest across exhausted windows
and buckets. Any missing deadline or workspace credit/billing ceiling omits the
timer. The native error has no responsible window/reset; if no trustworthy usage
snapshot supplies those facts, recovery remains manual. Recovery is account-wide,
not narrowed to whichever bucket currently happens to be exhausted; a bucket
recovering must not change the wait's quota identity. API-key/non-ChatGPT or
older servers may provide no account identity/availability; that stays unknown.

## Rollback, fork and compaction

- `rollback` finds the turn holding the item and calls `thread/revert`
  (`thread/rollback` is deprecated). It rewrites history only; the host
  pairs it with its own file checkpoint, and rewinds the agent first so a
  refusal leaves the files untouched.
- `fork_session` uses `thread/fork` with `excludeTurns`.
- `compact` uses `thread/compact/start`; the `contextCompaction` item then
  arrives as a `Compacted` event.

## Skills

`skills/list` per workspace fills `list_skills`, and the name-to-path map
it builds is what lets a `$name` prompt block become Codex's own
`UserInput { type: "skill", name, path }`. A skill the CLI does not know
falls back to the literal text, so nothing is silently dropped.

## Updating

`AgentInfo.maintenance` reports the installed version and the installer
that owns the binary, proven from the resolved real path
(`api.process.which`): a versioned Homebrew keg or cask,
`<prefix>/lib/node_modules/`, or a native install.
Anything unproven reports the version but stays manual. An npm update pins
`--prefix`, because the `npm` on `PATH` can belong to a different Node than
the one that owns the CLI.

## Deliberate gaps

- **The `granular` approval policy is not offered.** It needs five
  sub-toggles that one `ConfigOption` select cannot express, and the four
  shared permission modes cover what the user actually chooses between.
- **No bundled model manifest.** T3 Code overlays `model-manifest.json` for
  display names and badges. That is the hardcoded-catalog pattern this
  project forbids, so the catalog stays whatever `model/list` reports.
- **`thread/start` takes `sandbox` as a mode string while `turn/start`
  takes a structured `sandboxPolicy`.** `sandboxPolicy()` converts;
  `workspace-write` declares the workspace as the only writable root and
  leaves network access off.
- **`slash_commands` is false.** Codex has no slash-command API; the
  nearest thing is `skills/list`, and skills are invoked through a
  `UserInput { type: "skill" }` block rather than a `/name` prefix. Wiring
  skills into the composer needs a product decision, so it is left out.
- **`plan` items are ignored in the live stream.** The structured plan
  arrives as `turn/plan/updated`, which is what the `Plan` event models; the
  `plan` item only carries free-form markdown. In `read_session` — where
  the notification is not replayed — it is kept as a single-entry `Plan`.
- **`Usage.used_tokens` reports `tokenUsage.last.totalTokens`**, the size of
  the most recent request, because the event documents context-window
  usage. `total` accumulates across the whole thread and would read as
  greater than the context window.
- **Sign-in.** `account/read` tells whether `codex login` happened, and
  the plugin reports `AuthRequired` when it has not. The `chatgpt` auth
  method sends `account/login/start { type: "chatgpt" }` (the Rust plugin
  sent `{}`, which the v2 schema refuses) and opens the `authUrl` it
  answers with. A plugin without views may open it because the runtime
  opens a web link for a provider while the user signs in to one of its
  agents; a link it keeps closed (`{ opened: false, reason }`) fails the
  sign-in with the page's address.
- Not wired: `review/start`, realtime and voice, `feedback/upload`, and the
  host-service families (`fs/*`, `command/exec`, `process/*`). The host owns
  those surfaces in Divergence.
- **No cached status across restarts.** A provider plugin is only ever
  called by the host and cannot call back, so there is nowhere to keep a
  snapshot. A cold start re-probes.

## Tests

`node --test plugins/codex/*.test.ts` (and `plugins/sdk/*.test.ts`), no
network. The Rust unit tests are ported: command classification, diff
mapping, transcript roles, delta accumulation (with the no-duplicate rule
for completed messages), reasoning replacement, terminal output and exit
code, run completion/failure/one-shot, usage and titles, approvals and
questions answered and withdrawn through a scripted app-server (the SDK's
`FakePeer` in place of the Python peers), plugin tool calls, subagents
(v1 spawn and `wait`, v2 activity, grandchildren, waiting, Stop order and
approval release, `cancel_task`), and the history rebuild at depth 2.
Added: initialize, options and `set_option`, prompt and steer, a refused
turn, the app-server exiting, paged history, skills, rollback, fork,
compact, sign-in, the stdin notice, and the accounts in `main.ts`.

`cargo test -p convergence-host --test e2e -- --ignored` runs the plugin in
the real plugin host against the installed CLI: a PONG prompt, a turn that
runs `cat` and calls a plugin tool, an imported thread's history, and
`initialize` with the maintenance check.

Async `agentMessage` items (`delivery: "async"`, `questions`) emit message-mode
questions on completion. Codex supplies the titles and optional choices; custom
answers are always allowed. The host sends answers through the normal prompt
path, which uses `turn/steer` while a Codex turn is active.

Versions

VersionPublishedPlugin APISizePermissionsStatus
0.2.0latestOct 5, 2026>=2 <378.1 KB4 permissionsListed

Reviews and comments

0 threads · 0 reviews

No comments yet.