这位博主将Harness Engineering做成了可视化的动画。
他认为,理解 AI Agent 的发展方向,关键不在模型本身,而在于学习并做好 Harness Engineering(代理runtime工程)。因为,模型只是其中一部分,当 Agent 开始读文件、调用工具、修改状态、跨多步执行时,系统质量将越来越取决于模型周围的软件层(harness)。要点如下:
Harness 的职责:管理上下文选择与压缩、工具接口与权限、执行控制、持久状态、检查点、限制和执行轨迹。
长任务带来的难题:历史越积越多,不能全部重放;需要决定哪些信息留在上下文、哪些摘要/检索、哪些放到窗口外持久状态;还要支持中断恢复、避免重复劳动、控制高风险操作,并保留足够轨迹便于排查错误。
工程工作内容:通过查看执行轨迹、在代表性任务上评估、分析反复出现的失败模式,然后改进上下文策略、工具暴露方式、状态处理和运行时控制。
关键洞察:很多 Agent 的失败,并非靠换更强模型就能解决,而是要调整“模型看到什么、工具怎么暴露、状态怎么保存、失败后runtime怎么处理”。随着任务变长,Harness Engineering 才是把模型能力变成可控、可检查、可测试、可迭代执行过程的关键。
If you are trying to understand where AI agents are going, learn harness engineering.
A capable model is only one part of an agent system. Once the model begins reading files, calling tools, modifying state and working across many steps, the quality of the system depends increasingly on the software around it.
Consider a coding agent working through a large repository. The model can decide that it needs to inspect a file, search for a symbol, make an edit or run a test, but those decisions do not execute themselves.
The surrounding runtime has to decide which resources are available, whether the requested action is permitted, how the operation should be performed, what result should be retained, and what information should be presented to the model on the next step.
This becomes harder as the run gets longer. As history accumulates, replaying everything can become costly and less effective. The harness has to decide what should remain in context, what should be summarized or retrieved later, and what belongs in persistent state outside the context window.
Execution has similar problems. A long-running agent may need to survive an interruption, avoid repeating completed work, enforce permissions around consequential actions, and preserve enough history to reconstruct what happened when the final result is wrong.
These are harness problems.
The harness is the layer that manages context, tools, execution, state, checkpoints, limits and traces around the model.
Harness engineering is the work of designing and improving that layer. Engineers inspect execution traces, evaluate agents on representative tasks, look for recurring failure modes, and then change things such as context selection, tool interfaces, state handling or execution controls.
That last part matters because agent failures are often not fixed by changing the model. Sometimes the useful change is in what the model sees, how a tool is exposed, what state is preserved, or what the runtime does after a failed step.
As agents take on longer tasks, the demands on this surrounding software grow. Model capability remains essential, but harness engineering is what turns that capability into an execution process that can be controlled, inspected, tested and improved.