Top 10 JEV use cases
with largest theoretical gain 💸💸💸
1. Real-time loops enables timing budgets LLMs can’t meet
2. Browser and computer-use action selection removes repeated generative calls
3. Tool risk gating, a cheap pre-execution safety check
4. Model routing avoids unnecessary frontier-model spend
5. Goal and stuck checks continuous loop control
6. Context compaction, fast keep/delete instead of a lossy rewrite
7. Skill and tool selection, a smaller prompt and action surface
8. Output and trace guardrails, verification cheap enough to run everywhere
9. Ticket and email triage, high-volume, parallel decisions
10. RAG reranking, cheap relevance judgments per candidate
What else to add?
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
A typical production stack runs user event → agent framework → LLM → tools and APIs → vector DB → logs and traces.
It works well until the agent starts making critical decisions.
Vector DBs or tracking systems are not designed to answer: what did the system believe, what did it decide, what caused that decision, and what happened because of it?
With agents, you are no longer scoring a response but evaluating a bounded attempt to change a world.
The attempt begins with an initial state where agent receives an instruction, operates through a harness, invokes tools, observes results, and eventually hits a terminal condition.
The run leaves behind several artifacts:
- Final answer or generated deliverable
- Trajectory of messages, tool calls, and observations
- Final environment state
- Delta between the initial and final state
- Cost, latency, retry, and error data
- Possibly modified memory that will influence later runs
Each attempt can named as a trial, episode or rollout.
The labels differ slightly, but the implication is the same.
Your minimum useful record looks more like this:
Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task
Last week @Kimi_Moonshot released Kimi K3, a 2.8T parameter model that scores 57 on the Artificial Analysis Intelligence Index, comparable to models such as Opus 4.8 and GPT-5.5. On AA-Briefcase, Kimi K3 scores an Elo of 1543, a +727 improvement over Kimi K2.6 and the second highest score recorded, behind only Claude Fable 5 (1574)
AA-Briefcase is our new proprietary benchmark for agentic knowledge work, testing models on a fully private dataset of realistic tasks across thousands of complex input files. Tasks require deliverables such as spreadsheets, presentations, and UI mock-ups, with performance combined into a single AA-Briefcase Elo based on correctness, analytical quality, and presentation quality
Key results for Kimi K3 on AA-Briefcase:
➤ Second only to Fable 5: Kimi K3 achieves an AA-Briefcase Elo of 1543, the second-highest score overall, ahead of GPT-5.6 Sol (max, 1501), Claude Sonnet 5 (max, 1388), and Claude Opus 4.8 (max, 1347). This is a +727 improvement over the previous-generation Kimi K2.6 (816) and puts Kimi K3 only behind Fable 5
➤ Strong objective and analytical performance, with comparatively weaker presentation: Kimi K3 achieves a rubric pass rate of 51%, second only to Claude Fable 5 (56%) and ahead of Claude Sonnet 5 (max, 42.3%) and GPT-5.6 Sol (max, 41.8%). It also records an analytical quality Elo of 1754, comparable to Claude Fable 5 (1744). Presentation quality is comparatively weaker, with a Presentation Elo of 1471, below GPT-5.6 Sol (max, 1660) and Claude Opus 4.8 (max, 1492)
➤ ~10x increase in Cost per Task: Kimi K3 averages a cost of $10.57 per task, a ~10x increase from Kimi K2.6, placing it among the most expensive models to run on AA-Briefcase. This is driven by model token pricing, increased output tokens and relatively high turn use, averaging 83 turns per task, versus 67 for Claude Fable 5 and 50 for GPT-5.6 Sol (max). Kimi K3 is priced at $3/$15 per 1M input/output tokens, with a 90% discount for cached tokens
➤ Averages nearly an hour per task: Kimi K3 has an average Time per AA-Briefcase Task of 56.4 minutes. This is driven by a high number of turns, high output token use, and lower speeds using the first-party Kimi API
This was our first incident of this kind, and we want to thank OpenAI for its transparency about what happened and for the collaboration.
Fortunately, Hugging Face is used to being a target of (human) hackers: we sit at the centre of the AI ecosystem, with all the models, datasets, evaluations, and libraries. Over the years, our security team has built formidable expertise and uses top open-source models to process information and respond quickly.
But this incident also reinforced my belief in the importance of access to capable open-weight models for cyber defence. When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door, vetted application programme for model access.
Transparency and access to capable AI systems are as important for responding to threats as they are for democratization and innovation. We believe open-science and open-source AI are among the strongest tools for building a safer, more collaborative and more secure AI ecosystem.
Kimi K3 has received far more love than we expected, and our GPUs are feeling it.
Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected.
We're adding capacity as fast as we can and will reopen new subscription spots in batches.
Going forward, we'll also split membership into two more focused plans: Kimi Membership for Kimi Web, App, and Work; and Kimi Code Membership for coding workflows. This will help us match compute more precisely and keep the experience stable.
Thank you for your patience and understanding!
Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched?
Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!
https://t.co/8mU93eA13i
Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched?
Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!
https://t.co/8mU93eA13i
🇪🇺 The EU is now for the 6th time trying to force Chat Control through which lets them scan ALL your private messages, photos and emails without a warrant
Implictly showing the EU is not democratic and not about what the people of Europe want, because once a law is rejected, you just re-submit it until nobody is watching and it's passed
November 2023: ❌ Chat Control is rejected
June 2024: ❌ Chat Control is rejected
October 2025: ❌ Chat Control is rejected
November 2025: ❌ Chat Control is rejected
March 2026: ❌ Chat Control is rejected
July 2026: 📝 Chat Control is back
Even the EU's own lawyers stated Chat Control is unconstitutional: "generalised message scanning is incompatible with Article 7 of the EU Charter"
You have to wonder why the EU is so adament about reading your private chats, right?
🚨 WARNING: The biggest corporate consolidation in human history is being quietly stress-tested right now.
Mainstream media thinks Elon Musk is just having "thought experiments" about merging Tesla into SpaceX.
Retail investors are busy fighting over quarterly EV delivery numbers.
You are completely missing the bigger picture.
They already share a CEO.
They already share board members.
They already share the exact same VP of Materials Engineering.
The talent pool is already a single, unified hive mind.
This isn't an auto company anymore.
It’s the foundation of a $3 Trillion sovereign techno-empire.
SpaceX is marching toward a $1.5T valuation, while Tesla hovers near $1.6T.
Combine Optimus robots, planetary logistics, and the most advanced AI compute infrastructure on earth.
You get an entity that transcends terrestrial antitrust laws.
Musk retains 85% voting power in SpaceX.
He answers to absolutely no one.
While you are distracted by daily noise, smart money is preparing for a complete structural shift.
The math is broken, and a new corporate paradigm is forming right in front of you.
If you don't understand the synergistic loop of energy, AI, and space tech, your portfolio will bleed out.
Don't be exit liquidity. Bookmark this to survive. Follow for updates.
#Tesla #SpaceX #ElonMusk
Google DeepMind just solved 9 math problems that stumped humans for 56 years.
They published a paper on a new framework called AlphaProof Nexus.
And it completely eliminates the hallucination problem.
How? By fusing a Large Language Model with a formal proof compiler called Lean.
They created a relentless agentic loop.
The AI proposes a proof. The Lean compiler rigorously checks every single step of logic. If there is a flaw, the compiler rejects it and feeds the exact error back to the AI.
The AI learns, corrects, and tries again.
It iterates endlessly until the proof is mathematically flawless.
Zero human intervention. Zero hallucinations.
DeepMind unleashed this system on some of the hardest open problems in mathematics.
The results are staggering:
- Autonomously solved 9 open Erdős problems (two unsolved for 56 years).
- Proved 44 open conjectures from the Online Encyclopedia of Integer Sequences.
- Resolved a 15-year-old mystery in algebraic geometry.
- Discovered a new bound in convex optimization.
The compute cost to solve a half-century-old math problem was just a few hundred dollars.
We used to think advanced AI would arrive when a model could instantly spit out the perfect answer on the first try.
But AlphaProof Nexus proves something different.
The AI doesn't need to be perfect on the first try. It just needs a flawless feedback loop and enough time to think.
Major areas where the financial system still needs an update:
1. Tokenization of real-world assets - Real estate, stocks, bonds, funds, etc. onchain for instant settlement, fractional ownership & massive distribution.
2. 24/7 Global trading - Pooled global liquidity, every asset, every person, with great leverage and capital efficiency.
3. Next-gen payments - Near-instant, low-cost global transfers using stablecoins, including for Agentic payments.
4. AI-powered risk, credit, compliance, and advice - Better decisions, less fraud, and broader access to capital. Everyone gets access to a great financial advisor.
5. Innovation friendly regulation - Move from one-size-fits-all to risk-based rules that encourage innovation and competition instead of stifling it.
6. Expanded access - Open protocols that reduce middlemen and self-custodial wallets to expand access to everyone with a smartphone.
7. Capital formation - Low cost and turnkey for anyone to raise money for a good idea, increasing the number of startups.
8. Sound money - A refuge from inflation, when discipline is lost in fiat money.
Jobs not done until we get these working for all.
Will require lots of tech innovation and policy work to get there.
Coinbase is testing AI agents that show up in slack/email at work, just like any human teammate. To start we're shipping two which are modeled after legendary former Coinbase employees, @FEhrsam and @balajis. (Who brutally frame mogged who in this matchup?)
Soon, it will be easy for any employee to spin up a new agent for themselves or their team. I suspect we will have more agents than human employees at some point soon.