MiniMax Code Is Now Open Source: Inside the Harness Behind a Coding Agent

On September 18, we released MiniMax Code CLI v0.4.12 and opened its source under the MIT license. The CLI is a core component of the MiniMax Code client. With this release, developers can inspect the harness that runs it. Our launch announcement has the release details.

GitHub: github.com/MiniMax-AI/minimax-code
MiniMax Code CLI, from the launch announcement.

MiniMax Code CLI, from the launch announcement.

The announcement also shared a FrontierHarness Eval run: MiniMax Code passed 23 of 30 tasks, a 76.7% pass rate, with a median runtime of 4 minutes 33 seconds among successful tasks. The evaluation report documents the results and comparison conditions.

76.7% pass rate, 4m 33s median among successful tasks. This is a numerical comparison of that run, not an official leaderboard ranking.

76.7% pass rate, 4m 33s median among successful tasks. This is a numerical comparison of that run, not an official leaderboard ranking.

Effective cost per passed task, including failed attempts and repriced at benchmark rates: approximately $1.83.

Effective cost per passed task, including failed attempts and repriced at benchmark rates: approximately $1.83.

This article follows the implementation. Suppose a tool edits a file, then the process exits before committing the result to history. What will the session know when it resumes? Or suppose the model has already read a tool's schema when a plugin update arrives. Which implementation should execute the call?

Answering these questions means tracing runtime state and the order of operations. In the launch announcement, we invited developers to inspect tool execution and permission handling. Here, we walk through how MiniMax Code assembles capabilities, admits tool calls, manages context, and acknowledges history and subtask results.

1. What happens when a task enters MiniMax Code?

Consider a task: fix an edge case in a function, check its callers, and run the tests. The model might search the repository, read several files, propose a change, and use the test output to decide what to do next. One user request can require many model requests and tool calls.

MiniMax Code has three entry points: interactive mcode, script-oriented mcode exec, and mcode acp for compatible clients. All three use the same in-process runtime through CliService and local applications. At startup, createEmbeddedRuntimeHost() creates a local host, waits for it to become ready, and obtains CliService. This path does not depend on the Desktop HTTP service. Model requests and some tools can still access remote services.

A Session holds state that must survive across turns: conversation history, session configuration, and task relationships. When the runtime admits an execution, it creates a Turn with its own identity, cancellation signal, and terminal state. Model requests and tool steps run within that Turn.

LocalAgentHost coordinates the execution. It prepares the environment, assembles capabilities through agent-runtime, and hands the model-and-tool loop to PiTurnRunner, which uses the vendored Pi components. The Session system owns history storage and recovery.

At each iteration, the runtime prepares a model request. If the model asks to use a tool, the runtime checks and executes the call, then commits the result to history for subsequent requests. Search results, file contents, and test output gradually enter the context this way. When the model produces its final response, or execution is interrupted, the Host still has history and Turn state to finalize.

A Turn can therefore contain multiple model requests, and a Session can span multiple Turns. Although PiTurnRunner can be reused, each runTurn() creates a fresh Agent, event bridge, event queue, and history cursor. Mutable execution objects belong to the current Turn; information needed by later Turns belongs in the session layer.

The rest of this article follows that loop: preparing model input, executing and recording proposed actions, then extending the same process to child agents.

2. What must the harness prepare before a model request?

To fix our example function, the model needs the task requirements, findings so far, and the tools available to it. That input changes as work proceeds. Newly read files must be included, test logs may be large, and plugin configuration can change while the task is running.

Two constraints apply. The tool descriptions in the request must match the capabilities used by this Turn. The complete request must also fit the model's context window. MiniMax Code handles these through capability assembly and context management.

Fix the Turn's capabilities before building requests

The extension Registry normally moves through new → initializing → ready. Initialization errors put it in failed. Each extension registers capabilities through an owner-scoped interface, and registration is allowed only within its initialization window. A partially initialized registry cannot be treated as ready.

On entering assembleTurn(), the Registry synchronously snapshots which extensions are enabled before the first await. A call to setEnabled() during assembly takes effect on the next assembly. Hooks are also created as independent closures over the current Turn's ctx, preventing concurrent assemblies from sharing the wrong context.

Plugin capabilities are leased by revision. The Turn capability lifecycle acquires a view containing plugins, skills, runtime tools, and hooks for each admitted execution. Capturing the view and registering the in-flight execution happen without an intervening await, within one synchronous JavaScript critical section. The execution releases its reference when it ends.

A plugin update therefore cannot silently swap the capability version after the model has read its tool descriptions. Newly admitted Turns can use the new version. This fixes the capability view; files and remote services accessed by those tools can still change.

Tool descriptions also consume context. The MCP disclosure policy checks its feature switch, model allowlist, and effective window before estimating the description size of configured MCP tools. At the threshold, eligible tools are placed behind a search index and a deferred-call registry, while other tools remain directly exposed. The default threshold is 15% of the window; model eligibility and configuration overrides still apply.

Deferred disclosure saves the space occupied by a full tool catalog. It also adds a discovery step, and finding the right tool depends on retrieval quality.

As files and test output accumulate, will the next request still fit?

Once capabilities are available, the runtime converts history into model input. MiniMax Code's local footprint measurer estimates tokens across the messages, system prompt, and tool definitions, and calculates the UTF-8 serialized byte size of that request representation. These local measurements can differ from the provider's exact accounting of the final wire payload.

The budget reserves room for output and continued execution. Let C be the context window, O the output limit, R the reserve, and S the safety margin. The shared defaults are R = 16,384 and S = 2,048.

When C − O − S ≥ R, budget calculation uses the full output allowance O. Otherwise, it uses min(O, R). Calling this value O_eff, the input limit is:

L = max(1, min(floor(0.95 × C), C − R, C − O_eff − S))

Automatic compaction starts at an earlier threshold:

T = min(L, max(1, C − min(2 × R, floor(C / 4))))

For a window of 131,072 and an output limit of 16,384, these defaults give L = 112,640 and T = 98,304. Context management starts before the input reaches its hard limit. Different models or settings require recalculation.

If our task accumulates large amounts of source code and test logs, the runtime's ToolResultArchiver selects archive candidates using output watermarks, candidate size, protection for recent turns, and expected savings. It first makes a plan without changing state, then materializes the archive. When a read tool is available, large older outputs can be replaced with archive references while their originals remain available locally.

An archive candidate must pass validation: at least one of its token count or byte size must decrease, its token count must fit the input budget, and its byte size must satisfy the applicable limit. Without an explicit byte ceiling, the candidate's serialized size must not exceed the original.

If the archive candidate cannot satisfy admission directly, execution moves to the LLM checkpoint path. The generated checkpoint is validated, history is replaced, and the request is measured again. If it still exceeds the limits, the result is POST_ADMISSION_FAILED. Compaction is complete only when the next request, including the prompt and tool definitions, satisfies admission.

The request used to generate a checkpoint can itself be too large. Input-reduction paths handle tool output, media, and intermediate history. History reduction preserves tool-interaction closure so that tool calls and their results remain paired. Even when the result fits, compaction can lose task information. Evaluation should check for missed constraints and repeated investigation as well as size reduction.

An in-process usage anchor helps decide when to compact. It takes an eligible assistant usage report as an anchor, adds an estimate of subsequent growth, and uses the larger of that estimate and the current local estimate. The anchor is tied to the session scope, provider, API, model, system-prompt/tool fingerprints, history epoch, and anchored message identity. Changing models, tools, or history requires checking whether it still applies.

Comparisons between pre- and post-compaction candidates still use fresh estimates for each. The anchor informs the trigger decision; it does not replace candidate validation. With its capability view and current context prepared, the runtime can continue requesting model output.

3. How does the harness execute an action and record its result?

The model has read the code and decides to edit a file. The runtime now needs to establish two separate facts: whether the operation may execute at this moment, and whether its result has entered the history that subsequent execution will use. A tool returning output is only part of that process.

Recheck execution eligibility before starting the tool

The tool-control chain first checks for duplicate tool-call IDs within the Turn. For delegated calls carrying a host tool_ref, it also resolves the actual execution target. The call then passes through safety/plugin pre-processing, tool policy, and extension hooks before permission evaluation. Downstream checks receive the actual target, preventing authorization from checking only the wrapper tool.

Suppose the tool is waiting for user approval when the user cancels the task. By the time approval arrives, the original Turn may have ended. After permission is granted, the chain therefore attempts openToolResultTail(). It succeeds only if the matching active Turn is still running.

The Turn controller checks sessionId, turnId, leaseId, admission sequence, execution reason, and signal identity. A session can execute many Turns over time; a callback from an old Turn must not take control of a new one. The lease rejects stale execution within this runtime.

Actual filesystem and network access is also constrained by the enabled sandbox backend. The current V2 default registers the macOS srt-macos backend. Isolation on other platforms depends on the corresponding backend's support.

Commit tool results to canonical history

After the file edit succeeds, the model needs its result before proceeding to tests. If the process exits before history is committed, the file on disk can disagree with what the resumed session knows. MiniMax Code separates execution events from canonical history: events describe the running process, while canonical history is the authoritative history used to continue execution.

The Session system stores and restores that history. The Host uses a shared committed-history writer for ordinary appends, replacements, compaction, and reconciliation. History commits happen throughout execution; the compaction described above uses this path too.

A commit proceeds in this order:

  1. Capture a semantic snapshot and validate the session, Turn, and operation identity.
  2. Perform the durable mutation, then reread canonical history within the same per-session operation lane.
  3. Validate the reread result and obtain the committed revision.
  4. Wait for required HistoryCommitted projections and, when applicable, compaction lifecycle completion.
  5. Return acknowledgement to the caller.

The snapshot prevents the caller from changing the submitted object during an asynchronous commit. Rereading gives downstream components the version confirmed by storage, rather than the array the caller intended to write.

On retry, the writer identifies an operation by sessionId + operationId and compares its semantic fingerprint. It reuses the corresponding execution only when both identity and content match. Reusing an identity with different content triggers conflict handling. This replay registry is bounded and process-local; it does not guarantee exactly-once execution across processes or external services.

Additional input carried through tool-result handling follows a similar distinction between taking and acknowledging: the chain claims the input first and acknowledges it after history commits. Closing a Turn is deferred while consumption is still in flight. Unconsumed user steering has a requeue path, and already delivered machine input must not be dropped during finalization.

Required state projections must finish before history is acknowledged. Usage accounting and diagnostics are best effort: failures are recorded without blocking acknowledgement. External side effects need separate handling. If a remote write succeeds but its response is lost, the target service still needs an idempotency key or the application needs compensation. Local history cannot automatically undo the write.

Once tool results are committed, the loop returns to input preparation: build context containing the new results and request the model again. Our example continues editing and testing until it can finish. On exit, the Host confirms the runner's terminal state, reconciles history, runs turn_end handlers, and finally commits and acknowledges the Turn's terminal state. The model may have stopped generating while the Host is still doing this work.

4. How does this extend to multiple agents?

The same repair task can be divided: the main agent edits the function while a child agent checks its callers or investigates related code. The child has its own session and execution, and the main agent needs to receive its findings at the appropriate point. Input preparation, tool execution, and history acknowledgement still apply within each execution.

MiniMax Code records an independent task ID, owning session, parent task, status, and output reference for background tasks. States include queued, running, stopping, succeeded, failed, canceled, and lost. The TUI displays delegated agents through parent-child session relationships.

Delegation adds another question: when has the parent actually consumed a completed result? Task completion, notification delivery, and result reading can happen at different times. The parent can also fail after reading the output.

deliveredAt records delivery state; an automatic completion notification does not mark the result as consumed. On the current Host path, a consumption candidate is recorded only after a successful read of terminal task output. Reading progress from a running task does not suppress the later completion notification. The Host confirms those candidates only after the parent's Turn ends as completed and its terminal state has been committed.

For example, the child finishes investigating callers, the main agent reads its findings through task_output, and the main agent subsequently fails. Permanently marking the result consumed as soon as the read returns could remove an opportunity to remind the parent when it resumes. Deferring confirmation until the Turn has committed its completion ties the read to the parent's completed execution.

That ordering also leaves a failure window: the Turn may already be committed when consumption confirmation or subsequent settlement fails. Such failures produce diagnostics; they do not retroactively turn a committed Turn into a failed one. The steps are not one atomic transaction across all components.

After a runtime restart, recovery checks task execution ownership. Tasks without a valid executor are settled as lost instead of remaining running indefinitely. The parent can continue from restored session history, but a vanished child execution process needs separate handling.

Closing thoughts

Completing a coding task requires suitable model context and a working sequence of tool execution, history commits, and subtask delivery. By opening these implementations, we hope more developers can examine the tradeoffs and use, test, and improve them in their own environments.

If you are interested in agent harnesses, take a look at MiniMax Code on GitHub. A Star helps support the project. We'd also love to hear what you find when you try it; issues and feedback are welcome. Thank you for your support.