I think it might be something like this:
Continue developing OSS models and optimizing token consumption in Junie. They have their own version of Qwen and could fine-tune OSS decision models for specific cases.
Build tools to improve the quality of the code and products AI generates. Create end-to-end infrastructure for hybrid teams and improve human-in-the-loop visibility at every stage.
I hope Apple fully uses the new Siri AI to create personal, fully private support through local language models.. perhaps pairing their existing local models with a personal decision model.
Join us TOMORROW (Sept 25) at 9:30am PDT for a special engineering roundtable: Agentic Programming Practices at @Yugabyte - https://t.co/tUaSKHQVly
Panelists:
Hideaki Kimura
Dmitrii Sherstobitov
Rajagopalan Madhavan
Nikhil Chandrappa
Hari Krishna Sunder
Host: Me 😎
Five of our Yugabyte engineers share how agentic coding works day-to-day on a large distributed database codebase, including what each one tried, kept, and moved on from. We'll cover agentic code review, how much human review different kinds of work still need, bug hunting and production debugging with agents, and sharing project context across teammates and agents with Meko, our agent context engine. We'll also talk about what didn't work. These practices are still changing month to month, and the panel will explain why they made the choices they did.
Looked deeply into Laya. It's interesting.
The timeline claims are fuzzy and benchmarks are mixed but the model is worth following.
Laya, as compared to Jev by the author and @benchmarkheaven:
-#4 overall w/ 70.1 vs Jev’s 75.4
-Much stronger cost score: 86 vs 52
-Weaker intelligence, calibration + speed
-34% vs 74% on hard questions
Most interesting result comes from Laya’s own 2,000-decision benchmark...
When *fine-tuned* its typed-decisions checkpoint scored 0.766 vs Jev’s published 0.727.
Laya trained on 6,000 decisions from the same 4 synthetic workflows, whereas Jev ran zero-shot.
This is a narrow, same-distribution win, to be super clear. But ~0.36 base → 0.766 fine-tuned is a big jump. Expect us to go hard on fine-tuning open weight models.
Lots of caveats on these benchmarks and more empirical methods are definitely needed. But closed weight models can't be fine-tuned.
Whoever can create a highly fine-tunable Jev-like model will win the market.
@jpschroeder exactly. speed shouldn't be the only dimension.
you're right that we need real benchmarks to know where the line actually is. I tested OSS alternatives on a real project - the sub-1B models got 0/8 while jev got 8/8.
Just tried Laya on a real project, same questions and same setup: Jev 8/8, bonsai 7/8, lava 0/8.
Probably not general-purpose yet..more training needed. Still, typed decisions instead of generated text feel like the right idea and the amount of OSS effort here is great. Would love to see quality numbers next to the speed ones.
@jvr0x Yes, you are right about LLMs. My point is mostly about the quality of judgment; I don’t care how fast a model can deliver wrong results.
I can create a coin flip algorithm that will give me better judgment. Insanely fast.
We're adding support for AGENTS.md to Claude Code.
Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md.
You can toggle this behavior in /config.
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.
This is the future.
At the beginning of the year, I was testing local models with a proper harness. Even small models with proper knowledge and static analysis generated working code. Things have changed a lot since then.
JetBrains are kings of development instruments. I’m sure that with the proper CLIs and context infrastructure, small LLMs will shine.
As our CFO @_balaji_km mentioned at earnings today, we’re seeing some very interesting trends on AI costs. I think it’s another signal that we’re coming to the end of the so-called ‘tokenmaxxing’ era.
Here’s what’s been happening behind the scenes.
Since the beginning of the year we’ve more than quadrupled the number of people using frontier AI tools. That’s thousands of engineers using them every single day. During that same period, our cost per token has declined.
You might expect costs to rise as adoption accelerates. We've seen the opposite. Not because we've restricted access, but because we've treated efficiency as an engineering problem rather than a budget problem. A few examples:
• Caching and reuse: We use optimizations to improve our prompt cache hit rate that reduce our input token spend.
• Better defaults and tooling: We tuned default model settings, context sizes and developer workflows so teams get the same results with fewer tokens and lower-cost inference.
• Visibility drives efficiency: We gave engineers real-time visibility into their AI usage and costs per hour.
• Experimenting with open-weight models: we continuously evaluate new models and deploy the best option for each use case.
This is the future of applied AI at enterprise scale. The next phase, whatever we call it, will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible.
Credit to all the engineers at @Uber who are helping to build this future. 🚀