Agent Workflow State Machine
Magent models the agent loop as an explicit thread -> turn -> item ledger.
This is the sole source of truth for agent workflow state. Transcript and
provider context are named read-only projections of that ledger.
Codex Alignment
Codex exposes a thread event stream at the SDK/app-server boundary:
thread.startedturn.queuedturn.starteditem.starteditem.updateditem.completedturn.completedorturn.failed
Magent keeps the same workflow boundary, adapted to Emacs:
- Provider transport still goes through
gptel-request. - The Magent loop owns tool dispatch, continuation outcomes, abort, and Emacs UI rendering. The turn layer follows Codex-style continuation: tool output drives the next sampling request, while assistant completion ends the turn; missing final assistant text fails ordinary agent turns.
- No Codex sandbox, seatbelt, bubblewrap, or shell isolation parity is introduced.
The closest Codex references are:
sdk/typescript/src/events.ts: SDK event names and turn/item stream contract.codex-rs/core/src/session/turn.rs: core sampling loop, tool continuation, history recording, and follow-up decisions.codex-rs/app-server/src/bespoke_event_handling.rs: app-server conversion from core events into thread/turn notifications.codex-rs/app-server-protocol/src/protocol/v2/thread.rs: thread status shape.
State Objects
magent-ledger.el defines the canonical ledger:
magent-thread- statuses:
not-loaded,idle,active,system-error,closed
- statuses:
magent-thread-turn- statuses:
queued,in-progress,completed,interrupted,failed,dropped
- statuses:
magent-thread-item- statuses:
pending,in-progress,completed,failed,cancelled
- statuses:
Each user prompt creates one turn. Assistant messages, reasoning blocks, tool invocations, and tool outputs are items under that turn. Reasoning items are never promoted to assistant message text, even when a provider finishes with no visible content.
The transition table is explicit:
| Object | Runtime statuses | Terminal statuses |
|---|---|---|
| thread | not-loaded, idle, active, system-error, closed |
closed |
| turn | queued, in-progress |
completed, interrupted, failed, dropped |
| item | pending, in-progress |
completed, failed, cancelled |
Turn lifecycle:
queued -> in-progress -> completed
-> interrupted
-> failed
-> dropped
Item lifecycle:
pending -> in-progress -> completed
-> failed
-> cancelled
Tool Items
Tool calls and tool results are one item lifecycle, not two persisted records.
The item starts when the model asks for the tool:
(:type tool :status in-progress :name "grep" :input (:pattern "..."))
It completes or fails when the tool result is available:
(:type tool :status completed :output "...")
This differs from the older tool-call plus tool-output split. For
prompt reuse, Magent still projects a completed tool item into gptel’s
historical (tool . PLIST) shape.
If the model asks for a tool and the result arrives later, Magent updates
the same item by call-id. If an older code path records only a tool
result, the loop creates a synthetic turn so the item still has a durable
turn parent.
Persistence
Persistence is log + snapshot, stored as three files that change at very
different rates.
<id>.jsonl: append-only on-disk event log, one JSON object per line. It records lifecycle transitions such asturn-queued,turn-started,item-started,item-content-appended,item-completed, andturn-completed. Streaming content appends here, so a save costs what changed rather than what the conversation has accumulated. Compaction retains a bounded recent tail controlled bymagent-session-log-max-events. Tool and permission audit retention is owned separately bymagent-audit.el.<id>.snapshot: materialized full thread state. It stores the current thread, turns, and items so resume does not need to replay from the beginning every time. It is rewritten only when it is missing or the log has grown pastmagent-session-log-max-events.<id>.json: the session header (identity, scope, title, status, child-agent jobs, approval overrides). It is small, and it is the only file session listing reads.
Session schema v7 writes all three:
// <id>.json
{ "id": "...", "schema-version": 7, "scope": "global", "agent-jobs": [] }
// <id>.jsonl (one event per line)
{ "seq": 1, "type": "turn-queued", "payload": { "turn": { "...": "..." } } }
// <id>.snapshot
{ "id": "...", "turns": [ { "items": [ { "...": "..." } ] } ] }
Loading requires the exact current schema. Unknown fields, missing nested
ledger fields, and other schema versions are rejected rather than converted.
Schema 6 files, which stored journal and snapshot inline, are migrated in
place on first read.
Replay semantics:
- Load
snapshotinto materialized thread state. - Attach the retained
journaltail for recovery/history visibility. - Apply only events with
seq > snapshot.last-event-seq.
This means snapshot is the fast restore point and the persisted journal
is a bounded recovery tail. The two are not interchangeable. Checkpointing
writes snapshot first and rewrites the log second, so a crash in between
leaves events that replay skips by last-event-seq.
Loop Flow
- UI submission enters
magent-runtime-api.el. magent-runtime-api.elcreates a queued ledger turn, records the completed user message item, and freezes prompt, agent, tools, skills, route, effort, thinking mode, observer, approval, and turn identity in onemagent-request-context.- When the submission actually starts,
magent-runtime-api.eltransitions the turn toin-progress. magent-agent-run-turnreuses that turn/user item idempotently instead of duplicating the user message.magent-agent-loopconsumes normalized LLM events.- Text and reasoning deltas update materialized in-progress items in the snapshot without appending one journal event per chunk. Terminal item events carry the final content.
- Tool-call events are accumulated until the provider-neutral
tool-call-batch-endevent closes the sampling batch. - Tool dispatch starts
toolitems, records approval metadata when available, and updates those same items tocompletedorfailed. - Tool output returns a continuation outcome such as
tool-output;magent-agent-run-turnowns the decision to rebuild the prompt from session history and start the next sampling request. This keeps tool execution separate from turn continuation policy. - An empty provider completion records
empty-completionmetadata and fails the turn without another model request. - Assistant completion records an assistant message item and completes the turn.
- Abort, failure, and dropped queued submissions transition the turn to
interrupted,failed, ordropped; in-progress items under an aborted turn are markedcancelled. magent-max-sampling-requestsis an optional safety guard, disabled by default. When enabled, reaching the limit fails the turn directly.
Prompt reconstruction is ledger-driven. When a current turn id is known, Magent includes ordered completed messages and terminal tool results from completed, interrupted, and failed turns, their user goals, and the current turn, so a follow-up can recover cancelled request context without letting later queued user submissions leak into the active model request.
UI Projection
The supported frontend is agent-shell. magent-agent-shell.el creates an
in-process ACP client implemented by magent-acp.el. ACP session/prompt
requests dispatch registered slash commands through magent-action.el and
submit ordinary prompts directly to magent-runtime-api.el. They stay pending
until the corresponding command invocation or ordinary runtime turn completes,
fails, or is cancelled. Command-owned turns and workflow steps use the same
runtime API. The runtime emits Magent-native observer events;
magent-acp.el converts those events to ACP session/update messages.
magent-runtime-queue.el owns queued/active turn state. The first
implementation uses one global active turn at a time. Each submission captures
the exact runtime session wrapper plus its canonical request context, so
cancellation is identity-scoped: cancelling one ACP session removes that
session’s queued work and aborts its active turn without dropping other
sessions’ queued turns.
This agent-shell + ACP path is the complete conversational UI projection and
slash-command surface. Explicit M-x wrappers enter the same
magent-action invocation lifecycle without creating another conversation
frontend. Shared runtime behavior remains frontend-neutral.
Codex Differences Still Preserved
Magent intentionally still differs from Codex in these loop-adjacent areas:
- UI is Emacs-native: supported interaction is agent-shell through an in-process ACP client.
- Provider streaming is normalized behind
magent-sampling-gptel.el, but transport remains gptel. - Codex core runs a turn as a multi-sampling loop where pending user input, mailbox items, auto-compaction, and tool follow-up can all extend one active turn. Magent keeps request serialization in the runtime queue and uses Codex-style continuation for tool results: tool execution records model-visible output, then the turn layer continues sampling until the model returns assistant completion. Magent does not implement Codex’s app-server mailbox, steering queue, stop hooks, or mid-turn auto-compaction.
- Codex thread status is app-server-facing
notLoaded/idle/systemError/active{activeFlags}. Magent keeps the same top-level status idea but also persists explicit local turn and item statuses because the Emacs session file is the durable source of truth. - Codex rollout history stores response items and turn context items.
Magent stores a ledger
snapshot + journaland derives explicit transcript, provider, compaction, and audit views for their respective consumers. - Codex provider/model plumbing is native to codex-core. Magent keeps
gptel-requestas the provider transport and normalizes gptel events before the Magent-owned loop consumes them. - Tool execution is serialized by Magent’s tool queue unless individual tool/runtime support is later added. Codex has richer per-tool runtime, approval, MCP, and process execution machinery; that is outside this workflow goal.
- Child agents are durable Magent jobs, documented in
AGENT_JOBS.org, not Codex app-server threads.
Backlog / TODO
No open TODO remains from the agent-workflow/UI refactor backlog. Future hardening candidates:
- Add per-tool renderer plugins for additional structured outputs beyond the generic file/path and grep-style link extraction.
- Add live visual smoke coverage for very large transcripts and long streaming sessions.
Design Review Summary
After this change, the main agent-loop design gap versus Codex is no longer the lack of explicit thread/turn/item lifecycle. Magent now has that lifecycle and persists it.
The remaining differences are deliberate product/runtime boundaries:
- Magent does not have Codex’s app-server lifecycle manager, subscriber model, unloaded thread cache, or active flags beyond the local ledger status.
- Magent does not have Codex’s same intra-turn input queue/mailbox semantics.
- Magent does not use Codex rollout files directly; it keeps its own
session JSON with
snapshot + journal. - Magent does not split tool call and tool output as separate source-of- truth records. A tool is one item whose status and output are updated.
Ordered messages, native replay, and plans
An incomplete Chat Completions stream can offer a retry continuation on its
normalized error event. The turn layer permits up to magent-stream-retry-limit
retries (default 2), counted toward magent-max-sampling-requests. Eligibility
requires no assistant text and no finish reason. Both clean premature EOF
and curl partial-transfer errors (exit 18, HTTP 200) qualify; explicit provider
errors and token limits are not retried. The adapter marks transport failures
before gptel’s cleanup transition so unfinished tool calls cannot execute.
The adapter restores the exact request input and
clears partial response state before restarting the same gptel FSM, after
a bounded delay. Completed tools are not executed again. Cancellation kills
the request buffer and its retry timer. A durable notice item and
sampling-retry observer event explain each retry; notices stay out of model
history. Direct Action samplers retain their own terminal-error behavior.
Assistant progress messages are completed before tool execution and the final answer occupies its own item. Providers that expose message phases keep those values unchanged. For providers without phases, pre-tool text is commentary and terminal sample text is the answer. This is a display/replay contract, not an extra reasoning or judge pass. The prompt asks for short, evidence-based updates in plain paragraphs and a self-contained final answer; hidden reasoning is independent. OpenAI-compatible Chat continuations retain any visible pre-tool text in the provider-native assistant tool-call message.
Responses output items are stored as opaque JSON in provider items, with the
backend name, endpoint identity, and model. The gptel adapter observes gptel’s decoded stream events;
it owns neither HTTP nor SSE parsing. It keeps the original output array for
native continuation, including message phases and encrypted reasoning. History
replay uses native items only for the same backend/endpoint/model, avoids duplicate
assistant messages and function calls, and uses visible text/tool history for
other routes. A call with no recorded result receives an explicit unknown-result
notice in replay, preventing a dangling call without claiming it executed.
Compaction input omits opaque provider items.
Schema v6 already supports arbitrary item types, phases, metadata, and outputs. These changes need no format migration. Existing flattened messages remain intact with unknown phase; absent historical boundaries cannot be recovered. Ordinary empty answers now fail, while direct Action samplers that explicitly allow empty results keep that contract. A completed turn means that execution ended normally, not that all code or tests succeeded.
update_plan replaces the visible plan with a nonempty list of step and
status objects. Status is pending, in_progress, or completed, with at
most one step in progress. Each validated revision is a completed plan item;
step completion is the model’s progress report and does not complete the turn.
The tool uses permission key plan, is local, and emits a plan-update observer
event. ACP maps it to sessionUpdate: plan and replays saved revisions through
the same view. There is no independent plan runner or automatic replanning.