Introducing Marathon: adaptive inference infrastructure for long-running agents, powered by @GoKiteAI.
Run your agents at significantly lower cost, without sacrificing model quality.
How it works:
Install our one-line plugin for Codex or Claude Code, then choose a completion window per request: now, soon, later, or anytime. The more you’re willing to wait, the more you save, up to ~65% off.
Launching with support for the top five open-weight models, from @Kimi_Moonshot, @Zai_org, @deepseek_ai, @Alibaba_Qwen, and @nvidia.
Try Marathon today: https://t.co/80OlhbZgvC
Same model. Up to 65% less.
DeepSeek’s new API prices take effect on August 16. Marathon lets latency-tolerant workloads trade wait time for lower inference costs without switching models.
▷ Choose NOW when every second matters.
▷ Choose SOON or LATER when a few minutes are acceptable.
▷ Choose ANYTIME for deferrable work, with savings up to 65%.
Not every task needs an instant answer. Not every task should pay the instant price.
Pick your window: https://t.co/AlcAYs0xWM
Savings are live estimates and vary with capacity. The final price is shown when you submit the request.
Same model. Up to 65% less.
DeepSeek’s new API prices take effect on August 16. Marathon lets latency-tolerant workloads trade wait time for lower inference costs without switching models.
▷ Choose NOW when every second matters.
▷ Choose SOON or LATER when a few minutes are acceptable.
▷ Choose ANYTIME for deferrable work, with savings up to 65%.
Not every task needs an instant answer. Not every task should pay the instant price.
Pick your window: https://t.co/AlcAYs0xWM
Savings are live estimates and vary with capacity. The final price is shown when you submit the request.
Same model. Up to 65% less.
DeepSeek’s new API prices take effect on August 16. Marathon lets latency-tolerant workloads trade wait time for lower inference costs without switching models.
▷ Choose NOW when every second matters.
▷ Choose SOON or LATER when a few minutes are acceptable.
▷ Choose ANYTIME for deferrable work, with savings up to 65%.
Not every task needs an instant answer. Not every task should pay the instant price.
Pick your window: https://t.co/AlcAYs0xWM
Savings are live estimates and vary with capacity. The final price is shown when you submit the request.
Same model. Up to 65% less.
DeepSeek’s new API prices take effect on August 16. Marathon lets latency-tolerant workloads trade wait time for lower inference costs without switching models.
▷ Choose NOW when every second matters.
▷ Choose SOON or LATER when a few minutes are acceptable.
▷ Choose ANYTIME for deferrable work, with savings up to 65%.
Not every task needs an instant answer. Not every task should pay the instant price.
Pick your window: https://t.co/AlcAYs0xWM
Savings are live estimates and vary with capacity. The final price is shown when you submit the request.
Same model. Up to 65% less.
DeepSeek’s new API prices take effect on August 16. Marathon lets latency-tolerant workloads trade wait time for lower inference costs without switching models.
▷ Choose NOW when every second matters.
▷ Choose SOON or LATER when a few minutes are acceptable.
▷ Choose ANYTIME for deferrable work, with savings up to 65%.
Not every task needs an instant answer. Not every task should pay the instant price.
Pick your window: https://t.co/AlcAYs0xWM
Savings are live estimates and vary with capacity. The final price is shown when you submit the request.
Marathon Playground is live.
Interact with your favorite models and submit AI workloads through a sleek, built-in UI. No API keys to wire up, no external tools required.
Pick a model. Pick a completion window. Longer windows cost less: save up to ~65% at the longest window.
https://t.co/derLBYcZlb
Your signup credit is meant for building, so we're protecting it.
Automated accounts were draining the free credit pool at scale. We've reset balances and added a one-click reclaim, keeping that credit with developers who are actually running jobs.
Open your dashboard and hit "Reclaim your credits." Your API keys and purchased credit are untouched.
https://t.co/KTKngbFMZt
For all builders:
Your first tokens are on us. Free.
Sign up for Marathon and your signup credit converts to up to 45,600,000 tokens on the ANYTIME window.
Same credit, more tokens, whenever capacity is cheapest.
Try https://t.co/ko99sVXxrE 💻
Model="kimi-k3"
@Kimi_Moonshot K3 is live on Marathon.
Same OpenAI-compatible API, 1M context, and all four windows from the start: now, soon, later, anytime.
If your code already talks to Marathon, you are one line away.
https://t.co/derLBYcZlb
For all builders:
Your first tokens are on us. Free.
Sign up for Marathon and your signup credit converts to up to 45,600,000 tokens on the ANYTIME window.
Same credit, more tokens, whenever capacity is cheapest.
Try https://t.co/ko99sVXxrE 💻
Introducing Marathon: adaptive inference infrastructure for long-running agents, powered by @GoKiteAI.
Run your agents at significantly lower cost, without sacrificing model quality.
How it works:
Install our one-line plugin for Codex or Claude Code, then choose a completion window per request: now, soon, later, or anytime. The more you’re willing to wait, the more you save, up to ~65% off.
Launching with support for the top five open-weight models, from @Kimi_Moonshot, @Zai_org, @deepseek_ai, @Alibaba_Qwen, and @nvidia.
Try Marathon today: https://t.co/80OlhbZgvC