Introducing Marathon: adaptive inference infrastructure for long-running agents, powered by
@GoKiteAI.
Run your agents at significantly lower cost, without sacrificing model quality.
How it works:
Install our one-line plugin for Codex or Claude Code, then choose a completion window per request: now, soon, later, or anytime. The more you’re willing to wait, the more you save, up to ~65% off.
Launching with support for the top five open-weight models, from
@Kimi_Moonshot,
@Zai_org,
@deepseek_ai,
@Alibaba_Qwen, and
@nvidia.
Try Marathon today:
marathon.build/