Chinese students just found the best way to use JEV for any LLM or AI agent - released a PDF research
the shift: I pasted it into Claude and GPT - and cut my costs by~63х
here’s what they found across 44 benchmarks:
1 → 7,193 responses, 10 types of failure. Jev was tested on hallucinations, prompt injections, data leaks, and other AI failures
2 → One simple question worked: 0.886 median AUROC, beating trained baselines on 25 of 31 benchmarks without task-specific training
3 → Context beat clever prompting - give Jev the source or rule it needs to check the answer against
4 → Keep the probability, not just "yes" or "no" - Fitting a threshold on 10 labeled examples raised median F1 from 0.706 to 0.793
5 → Among the 50% most confident decisions, median accuracy reached 0.933 - send uncertain cases for another review
6 → Jev even helped uncover labeling errors in three benchmarks. Sometimes the test’s "correct answer" was the problem
7 → 11.4 questions per call, with 0.31-second median latency - on 19 benchmarks, checking cost $0.30 vs $18.96 with LLM judges - roughly 63× cheaper
the result: It will made your setup CHEAPER and FASTER than what 95% of people are running
Copy the Jev setup researchers tested across 44 benchmarks - then read the full Jev architecture ↓
Today, we’re releasing Kalypta, the first app to block AI notetakers in your meetings.
Granola? Wisprflow? Cluely? No more.
With Kalypta, you become inaudible to AI.
Your call continues normally.
> be anthropic
> say ai might take over internet
> researchers: “it might kill humans.”
> dario: “we need independent evaluators”
> choose metr
> metr gets funded by the same network funding anthropic
> which also funds ai safety orgs
> which also funds doom research
> which also funds policy work
> which also holds major wealth linked to anthropic
> which also funds tarbell
> tarbell pushes doom into time, the verge, science, la times, etc.
> anthropic makes more money
> network gets richer
> more fund goes into ai doom
> ai doom pushes more regulation
> regulation favors “responsible” frontier labs
> anthropic gets more power
> repeat
the doom funds the network → the network amplifies the doom → anthropic benefits from both
...showered with likes
Brits have lost all sense if respect for private property and personal decision making and now wonder why their country is going to shit
“Task-CoEvolve Efficient Harness Optimization via Adaptive Validation Task Selection”
Harness optimization usually wastes tons of compute re-evaluating every candidate on every validation task, even when many tasks are already too easy or too hard to be useful.
So this paper keeps sampling the tasks where candidate harnesses disagree most, then corrects for that sampling to estimate full-set performance.
On Terminal-Bench 2.1, it gets nearly the same final performance with only 20% of the evaluations, cutting search cost by 67-80%.
https://t.co/fJ8T4Ef0YG
Google just quietly dropped something big for agent builders
It's called SAM: Sovereign Agent Mesh
A peer-to-peer network where AI agents auto-discover each other, authenticate every packet and call tools across the mesh without a central server
Think of it like BitTorrent, but for your agents
Right now every agent stack you build is a walled garden. Claude Code talks to its MCPs. Codex talks to its MCPs. If you want agents running on your laptop, your VPS, and your Mac Studio to share tools with each other, you're duct-taping SSH tunnels and REST endpoints together like it's 2015
SAM makes that whole layer disappear
You run a sam-node on each machine and the nodes find each other automatically. Every connection is cryptographically identified, so your agents plug in and start calling tools across the whole network
The identity is portable too. Same node, same keys, whether it's on GCP, your laptop, or a Raspberry Pi in your closet
If you're serious about running agents across multiple devices, get on the testnet this weekend:
1. Install sam-node (binary or Docker)
2. Point it at https://t.co/qEWBOml0xU
3. Connect Claude or Gemini to your local MCP endpoint
4. Watch your agent discover tools running on other people's nodes
Still early. Not officially supported by Google.
This is where the agent stack is heading. Get on it early
Repo: https://t.co/1zIBGp0BgV
This is Quantum Chaos inside a Bunimovich stadium.
The motion is governed by the Schrödinger equation. The whirlpools in the flow are Quantum phase vortices.
There's a simple answer to which LLM is best for meta-analysis.
It's not Claude, ChatGPT, Grok, or Gemini.
It's Kimi, because it doesn't have any compunctions about using Sci-hub or any other tool to access papers.
The H100 hour just hit $3.08, its highest in a month.
A three-year-old chip is supposed to be depreciating fast. Instead it is at a monthly high, and the compute behind it is still in demand.
Depreciation is an assumption. This is the price at https://t.co/TYz89lPhU3