Agent Workflow State Machine

Magent models the agent loop as an explicit thread -> turn -> item ledger. This is the sole source of truth for agent workflow state. Transcript and provider context are named read-only projections of that ledger.

Codex Alignment

Codex exposes a thread event stream at the SDK/app-server boundary:

  • thread.started
  • turn.queued
  • turn.started
  • item.started
  • item.updated
  • item.completed
  • turn.completed or turn.failed

Magent keeps the same workflow boundary, adapted to Emacs:

  • Provider transport still goes through gptel-request.
  • The Magent loop owns tool dispatch, continuation outcomes, abort, and Emacs UI rendering. The turn layer follows Codex-style continuation: tool output drives the next sampling request, while assistant completion ends the turn; missing final assistant text fails ordinary agent turns.
  • No Codex sandbox, seatbelt, bubblewrap, or shell isolation parity is introduced.

The closest Codex references are:

  • sdk/typescript/src/events.ts: SDK event names and turn/item stream contract.
  • codex-rs/core/src/session/turn.rs: core sampling loop, tool continuation, history recording, and follow-up decisions.
  • codex-rs/app-server/src/bespoke_event_handling.rs: app-server conversion from core events into thread/turn notifications.
  • codex-rs/app-server-protocol/src/protocol/v2/thread.rs: thread status shape.

State Objects

magent-ledger.el defines the canonical ledger:

  • magent-thread
    • statuses: not-loaded, idle, active, system-error, closed
  • magent-thread-turn
    • statuses: queued, in-progress, completed, interrupted, failed, dropped
  • magent-thread-item
    • statuses: pending, in-progress, completed, failed, cancelled

Each user prompt creates one turn. Assistant messages, reasoning blocks, tool invocations, and tool outputs are items under that turn. Reasoning items are never promoted to assistant message text, even when a provider finishes with no visible content.

The transition table is explicit:

Object Runtime statuses Terminal statuses
thread not-loaded, idle, active, system-error, closed closed
turn queued, in-progress completed, interrupted, failed, dropped
item pending, in-progress completed, failed, cancelled

Turn lifecycle:

queued -> in-progress -> completed
                      -> interrupted
                      -> failed
                      -> dropped

Item lifecycle:

pending -> in-progress -> completed
                       -> failed
                       -> cancelled

Tool Items

Tool calls and tool results are one item lifecycle, not two persisted records.

The item starts when the model asks for the tool:

(:type tool :status in-progress :name "grep" :input (:pattern "..."))

It completes or fails when the tool result is available:

(:type tool :status completed :output "...")

This differs from the older tool-call plus tool-output split. For prompt reuse, Magent still projects a completed tool item into gptel’s historical (tool . PLIST) shape.

If the model asks for a tool and the result arrives later, Magent updates the same item by call-id. If an older code path records only a tool result, the loop creates a synthetic turn so the item still has a durable turn parent.

Persistence

Persistence is log + snapshot, stored as three files that change at very different rates.

  • <id>.jsonl: append-only on-disk event log, one JSON object per line. It records lifecycle transitions such as turn-queued, turn-started, item-started, item-content-appended, item-completed, and turn-completed. Streaming content appends here, so a save costs what changed rather than what the conversation has accumulated. Compaction retains a bounded recent tail controlled by magent-session-log-max-events. Tool and permission audit retention is owned separately by magent-audit.el.
  • <id>.snapshot: materialized full thread state. It stores the current thread, turns, and items so resume does not need to replay from the beginning every time. It is rewritten only when it is missing or the log has grown past magent-session-log-max-events.
  • <id>.json: the session header (identity, scope, title, status, child-agent jobs, approval overrides). It is small, and it is the only file session listing reads.

Session schema v7 writes all three:

// <id>.json
{ "id": "...", "schema-version": 7, "scope": "global", "agent-jobs": [] }

// <id>.jsonl   (one event per line)
{ "seq": 1, "type": "turn-queued", "payload": { "turn": { "...": "..." } } }

// <id>.snapshot
{ "id": "...", "turns": [ { "items": [ { "...": "..." } ] } ] }

Loading requires the exact current schema. Unknown fields, missing nested ledger fields, and other schema versions are rejected rather than converted. Schema 6 files, which stored journal and snapshot inline, are migrated in place on first read.

Replay semantics:

  1. Load snapshot into materialized thread state.
  2. Attach the retained journal tail for recovery/history visibility.
  3. Apply only events with seq > snapshot.last-event-seq.

This means snapshot is the fast restore point and the persisted journal is a bounded recovery tail. The two are not interchangeable. Checkpointing writes snapshot first and rewrites the log second, so a crash in between leaves events that replay skips by last-event-seq.

Loop Flow

  1. UI submission enters magent-runtime-api.el.
  2. magent-runtime-api.el creates a queued ledger turn, records the completed user message item, and freezes prompt, agent, tools, skills, route, effort, thinking mode, observer, approval, and turn identity in one magent-request-context.
  3. When the submission actually starts, magent-runtime-api.el transitions the turn to in-progress.
  4. magent-agent-run-turn reuses that turn/user item idempotently instead of duplicating the user message.
  5. magent-agent-loop consumes normalized LLM events.
  6. Text and reasoning deltas update materialized in-progress items in the snapshot without appending one journal event per chunk. Terminal item events carry the final content.
  7. Tool-call events are accumulated until the provider-neutral tool-call-batch-end event closes the sampling batch.
  8. Tool dispatch starts tool items, records approval metadata when available, and updates those same items to completed or failed.
  9. Tool output returns a continuation outcome such as tool-output; magent-agent-run-turn owns the decision to rebuild the prompt from session history and start the next sampling request. This keeps tool execution separate from turn continuation policy.
  10. An empty provider completion records empty-completion metadata and fails the turn without another model request.
  11. Assistant completion records an assistant message item and completes the turn.
  12. Abort, failure, and dropped queued submissions transition the turn to interrupted, failed, or dropped; in-progress items under an aborted turn are marked cancelled.
  13. magent-max-sampling-requests is an optional safety guard, disabled by default. When enabled, reaching the limit fails the turn directly.

Prompt reconstruction is ledger-driven. When a current turn id is known, Magent includes ordered completed messages and terminal tool results from completed, interrupted, and failed turns, their user goals, and the current turn, so a follow-up can recover cancelled request context without letting later queued user submissions leak into the active model request.

UI Projection

The supported frontend is agent-shell. magent-agent-shell.el creates an in-process ACP client implemented by magent-acp.el. ACP session/prompt requests dispatch registered slash commands through magent-action.el and submit ordinary prompts directly to magent-runtime-api.el. They stay pending until the corresponding command invocation or ordinary runtime turn completes, fails, or is cancelled. Command-owned turns and workflow steps use the same runtime API. The runtime emits Magent-native observer events; magent-acp.el converts those events to ACP session/update messages.

magent-runtime-queue.el owns queued/active turn state. The first implementation uses one global active turn at a time. Each submission captures the exact runtime session wrapper plus its canonical request context, so cancellation is identity-scoped: cancelling one ACP session removes that session’s queued work and aborts its active turn without dropping other sessions’ queued turns.

This agent-shell + ACP path is the complete conversational UI projection and slash-command surface. Explicit M-x wrappers enter the same magent-action invocation lifecycle without creating another conversation frontend. Shared runtime behavior remains frontend-neutral.

Codex Differences Still Preserved

Magent intentionally still differs from Codex in these loop-adjacent areas:

  • UI is Emacs-native: supported interaction is agent-shell through an in-process ACP client.
  • Provider streaming is normalized behind magent-sampling-gptel.el, but transport remains gptel.
  • Codex core runs a turn as a multi-sampling loop where pending user input, mailbox items, auto-compaction, and tool follow-up can all extend one active turn. Magent keeps request serialization in the runtime queue and uses Codex-style continuation for tool results: tool execution records model-visible output, then the turn layer continues sampling until the model returns assistant completion. Magent does not implement Codex’s app-server mailbox, steering queue, stop hooks, or mid-turn auto-compaction.
  • Codex thread status is app-server-facing notLoaded/idle/systemError/active{activeFlags}. Magent keeps the same top-level status idea but also persists explicit local turn and item statuses because the Emacs session file is the durable source of truth.
  • Codex rollout history stores response items and turn context items. Magent stores a ledger snapshot + journal and derives explicit transcript, provider, compaction, and audit views for their respective consumers.
  • Codex provider/model plumbing is native to codex-core. Magent keeps gptel-request as the provider transport and normalizes gptel events before the Magent-owned loop consumes them.
  • Tool execution is serialized by Magent’s tool queue unless individual tool/runtime support is later added. Codex has richer per-tool runtime, approval, MCP, and process execution machinery; that is outside this workflow goal.
  • Child agents are durable Magent jobs, documented in AGENT_JOBS.org, not Codex app-server threads.

Backlog / TODO

No open TODO remains from the agent-workflow/UI refactor backlog. Future hardening candidates:

  • Add per-tool renderer plugins for additional structured outputs beyond the generic file/path and grep-style link extraction.
  • Add live visual smoke coverage for very large transcripts and long streaming sessions.

Design Review Summary

After this change, the main agent-loop design gap versus Codex is no longer the lack of explicit thread/turn/item lifecycle. Magent now has that lifecycle and persists it.

The remaining differences are deliberate product/runtime boundaries:

  • Magent does not have Codex’s app-server lifecycle manager, subscriber model, unloaded thread cache, or active flags beyond the local ledger status.
  • Magent does not have Codex’s same intra-turn input queue/mailbox semantics.
  • Magent does not use Codex rollout files directly; it keeps its own session JSON with snapshot + journal.
  • Magent does not split tool call and tool output as separate source-of- truth records. A tool is one item whose status and output are updated.

Ordered messages, native replay, and plans

An incomplete Chat Completions stream can offer a retry continuation on its normalized error event. The turn layer permits up to magent-stream-retry-limit retries (default 2), counted toward magent-max-sampling-requests. Eligibility requires no assistant text and no finish reason. Both clean premature EOF and curl partial-transfer errors (exit 18, HTTP 200) qualify; explicit provider errors and token limits are not retried. The adapter marks transport failures before gptel’s cleanup transition so unfinished tool calls cannot execute. The adapter restores the exact request input and clears partial response state before restarting the same gptel FSM, after a bounded delay. Completed tools are not executed again. Cancellation kills the request buffer and its retry timer. A durable notice item and sampling-retry observer event explain each retry; notices stay out of model history. Direct Action samplers retain their own terminal-error behavior.

Assistant progress messages are completed before tool execution and the final answer occupies its own item. Providers that expose message phases keep those values unchanged. For providers without phases, pre-tool text is commentary and terminal sample text is the answer. This is a display/replay contract, not an extra reasoning or judge pass. The prompt asks for short, evidence-based updates in plain paragraphs and a self-contained final answer; hidden reasoning is independent. OpenAI-compatible Chat continuations retain any visible pre-tool text in the provider-native assistant tool-call message.

Responses output items are stored as opaque JSON in provider items, with the backend name, endpoint identity, and model. The gptel adapter observes gptel’s decoded stream events; it owns neither HTTP nor SSE parsing. It keeps the original output array for native continuation, including message phases and encrypted reasoning. History replay uses native items only for the same backend/endpoint/model, avoids duplicate assistant messages and function calls, and uses visible text/tool history for other routes. A call with no recorded result receives an explicit unknown-result notice in replay, preventing a dangling call without claiming it executed. Compaction input omits opaque provider items.

Schema v6 already supports arbitrary item types, phases, metadata, and outputs. These changes need no format migration. Existing flattened messages remain intact with unknown phase; absent historical boundaries cannot be recovered. Ordinary empty answers now fail, while direct Action samplers that explicitly allow empty results keep that contract. A completed turn means that execution ended normally, not that all code or tests succeeded.

update_plan replaces the visible plan with a nonempty list of step and status objects. Status is pending, in_progress, or completed, with at most one step in progress. Each validated revision is a completed plan item; step completion is the model’s progress report and does not complete the turn. The tool uses permission key plan, is local, and emits a plan-update observer event. ACP maps it to sessionUpdate: plan and replays saved revisions through the same view. There is no independent plan runner or automatic replanning.