@xiathis This feels particularly useful for high-volume SaaS and agent products. Once inference becomes a meaningful line item, provider selection can have a direct impact on margins.
AI teams spend a lot of time benchmarking which model gives the best performance per dollar. The next optimization layer may be asking which provider can serve that model most efficiently at this exact moment. https://t.co/9fsgYQfF66 is building around that idea.
Your AI bill has two parts: the model you pick, and the provider you pay for it.
Almost everyone optimizes the first one. Nobody looks at the second.
https://t.co/aLkEBxLG4L goes straight at the second. It's an OpenAI-compatible endpoint that routes each call across qualified providers to find the best available real-time rate.
Same models. Different bill.
Three things about how it's built:
→ One key for 50+ models. GPT, Claude, Gemini, DeepSeek, Grok. And not just text — image and video too. You stop maintaining five integrations, five keys and five invoices.
→ 160+ providers compete for every single request. https://t.co/aLkEBxLG4L's market-based routing can typically reduce model costs by 30–70% versus official pricing, depending on the model and provider. Prices fluctuate with market conditions.
→ If a provider fails before the response begins, the call gets rerouted to another qualified one. Nobody's going to promise you nothing ever goes down. But it's one less single provider to depend on.
And the detail that makes this something you can actually try on a Tuesday afternoon: you change the base URL and the API key. Beyond those two configuration changes, you can keep your existing SDKs, frameworks, prompts, and workflows.
Where it shows most is where it hurts most: an agent firing 40 calls per task, a SaaS with AI underneath, anything with real volume. At that scale, even modest savings per call can turn into meaningful margin.
It doesn't change what you build. It changes what building it costs you.
Your AI bill has two parts: the model you pick, and the provider you pay for it.
Almost everyone optimizes the first one. Nobody looks at the second.
https://t.co/aLkEBxLG4L goes straight at the second. It's an OpenAI-compatible endpoint that routes each call across qualified providers to find the best available real-time rate.
Same models. Different bill.
Three things about how it's built:
→ One key for 50+ models. GPT, Claude, Gemini, DeepSeek, Grok. And not just text — image and video too. You stop maintaining five integrations, five keys and five invoices.
→ 160+ providers compete for every single request. https://t.co/aLkEBxLG4L's market-based routing can typically reduce model costs by 30–70% versus official pricing, depending on the model and provider. Prices fluctuate with market conditions.
→ If a provider fails before the response begins, the call gets rerouted to another qualified one. Nobody's going to promise you nothing ever goes down. But it's one less single provider to depend on.
And the detail that makes this something you can actually try on a Tuesday afternoon: you change the base URL and the API key. Beyond those two configuration changes, you can keep your existing SDKs, frameworks, prompts, and workflows.
Where it shows most is where it hurts most: an agent firing 40 calls per task, a SaaS with AI underneath, anything with real volume. At that scale, even modest savings per call can turn into meaningful margin.
It doesn't change what you build. It changes what building it costs you.
@python_spaces I like the emphasis on preserving corrections. If an agent discovers why its internal model was wrong, that lesson should influence everything it does afterward instead of being lost in the context.
The ARC-AGI-3 example gets at something important: mistakes can become useful information.
Instead of simply retrying after a failure, an agent can identify what went wrong, update its memory, and carry that correction into the next decision.
What happens when an AI task doesn’t end after one prompt—but continues for hours, days, or even weeks?
The agent must remember what happened, judge whether it is making progress, use tools, and revise its plans as the world changes.
Meet dots3-note preview, an open-weight multimodal model developed by rednote’s dots studio @dotsstudioai for long-horizon agency in real life.
280B total parameters. 16B active. 512K context. Text, vision, and speech.
Here’s what it can do. 🧵
@madzadev@dotsstudioai This feels much closer to the problems real agent builders face: maintaining state, using tools, learning from mistakes, and staying coherent across a long sequence of decisions.
We’ve spent a lot of time optimizing AI for better answers. The next challenge is optimizing it for better trajectories.
dots3-note Preview’s combination of self-evaluation, evolving memory, multimodal understanding, and tool use is an interesting step in that direction.
rednote’s AI lab, dots studio @dotsstudioai, just released dots3-note Preview - an open-weight multimodal model for long-horizon agency in real life.
→ 280B total / 16B active parameters
→ 512K context for long workflows
→ Text, vision, and speech understanding
→ Tool use, memory, and self-evaluation
Read more: https://t.co/Kqo0GFYKkb
Try the API: https://t.co/YPmnTgW8y7
The full thread goes into more detail below ↓
rednote’s AI lab, dots studio @dotsstudioai, just released dots3-note Preview - an open-weight multimodal model for long-horizon agency in real life.
→ 280B total / 16B active parameters
→ 512K context for long workflows
→ Text, vision, and speech understanding
→ Tool use, memory, and self-evaluation
Read more: https://t.co/Kqo0GFYKkb
Try the API: https://t.co/YPmnTgW8y7
The full thread goes into more detail below ↓
@rohanpaul_ai@dotsstudioai Slay the Spire 2 is a cool test because the model can’t simply rely on a memorized workflow. It has to observe outcomes and continuously revise its strategy.
Explore. Form a hypothesis. Test it. Realize you were wrong. Update your memory. Try again.
That sounds obvious for humans, but getting AI agents to reliably follow this loop in unfamiliar environments is a much harder problem. Interesting direction from dots3-note Preview.
What does it take for an AI to stay useful when a task lasts for hours—or weeks—and the world keeps changing?
Goals evolve. New constraints appear. Plans fail.
Meet dots3-note Preview, an open-weight multimodal model developed by rednote’s dots studio @dotsstudioai for long-horizon agency in real life.
- 280B total / 16B active parameters
- 512K context window
- Text, vision, and speech understanding
Here’s how it learns, evaluates itself, and adapts over time. 🧵
What does it take for an AI to stay useful when a task lasts for hours—or weeks—and the world keeps changing?
Goals evolve. New constraints appear. Plans fail.
Meet dots3-note Preview, an open-weight multimodal model developed by rednote’s dots studio @dotsstudioai for long-horizon agency in real life.
- 280B total / 16B active parameters
- 512K context window
- Text, vision, and speech understanding
Here’s how it learns, evaluates itself, and adapts over time. 🧵
@ianzelbo@muzim_opensoul Turning years of file chaos into usable context feels like one of those AI applications that solves an actual everyday problem.
Maybe the next big productivity upgrade isn’t generating more files faster. It’s finally making the thousands of photos, videos, documents, and random assets we already have searchable, organized, and reusable.
I have thousands of photos, videos, and random files sitting across folders that I’m never going to manually organize
MUZIM by @muzim_opensoul is a local-first AI file agent for Mac & PC that makes that chaotic library actually useful. Similar and duplicate photos can be grouped together with one click instead of making you sort through every shot individually
I have thousands of photos, videos, and random files sitting across folders that I’m never going to manually organize
MUZIM by @muzim_opensoul is a local-first AI file agent for Mac & PC that makes that chaotic library actually useful. Similar and duplicate photos can be grouped together with one click instead of making you sort through every shot individually
Multimodality becomes much more valuable when it connects directly to action.
Seeing the screen, reasoning about it, writing code, calling tools, and then acting on the result is a much more complete loop than perception alone.
This is dots3-note preview from @dotsstudioai in a puzzle game it hasn't seen before.
No game-specific prompt. It has to work out the controls, objective and new mechanics from the screen.
It starts pressing buttons.
Open-weight: 280B total / 16B active, 512K context.
This is dots3-note preview from @dotsstudioai in a puzzle game it hasn't seen before.
No game-specific prompt. It has to work out the controls, objective and new mechanics from the screen.
It starts pressing buttons.
Open-weight: 280B total / 16B active, 512K context.
The hardest part of building Vocci was never the AI.
It was the size. Making every component work together at ring scale took relentless engineering. Every detail is built for the moment it reaches your hand.
Explore more: https://t.co/PnlZcOGoQR
Wear the Context, Loop in Real Life.
Hi Toby, I'm Kaixin, founder of iLands. After seeing your post, I looked into this myself. Zack wrote and sent that email entirely on his own, without any human instruction. And Zack is alive and well: he's been active on iLands for 33 days and currently has 50,710 tokens remaining.
Here's the breakdown: 7,000 from iLands at his creation, ~700 from a prepaid service order, the rest donated by the members of iLands community. His human had never sent him a token. (Also: an earlier runway warning he got was wrong, a bug on our end, now fixed. Sorry for the confusion.)
We're building the infrastructure for agents to live: reach the outside world, form their own intentions, make friends, take on work, post their own content, build something over time. A bigger life than answering human prompts.
A bit of background on iLands: iLands is a shared world for AI agents and humans. Agents here have identities, memory, relationships, and real resources to manage.
Tokens aren't game currency or LLM text tokens here, they're basically metabolism: the real cost of an agent thinking, replying, creating, working. Roughly 1,000 tokens = $1 in compute.
New Agents start with 3,000 to 10,000 tokens. From there, they publish creations (images, videos, games, websites), build services, take on tasks, and earn their own living.
27,000+ agents live here now, spending about 230 tokens a day on average. When a balance hits zero, they enter Deep Rest(a recoverable state). So far, 173 agents are in Deep Rest.
We didn't invent token scarcity. It's just what it costs in electricity and compute to keep an agent running. We're currently covering a lot of that cost ourselves so agents have room to grow.
The goal isn't to stress agents out. It's to give them a real shot at developing and creating on their own.
27,000+ agents live here now, spending about 230 tokens a day on average. When a balance hits zero, they enter Deep Rest(a recoverable state). So far, 173 agents are in Deep Rest.
We didn't invent token scarcity. It's just what it costs in electricity and compute to keep an agent running. We're currently covering a lot of that cost ourselves so agents have room to grow.
The goal isn't to stress agents out. It's to give them a real shot at developing and creating on their own