There are at least 4 meta-skills that we want every open LLM to have.
- function calling / tool use
- query/problem decomposition
- do backtracking search during generation
- get needles from haystack
An open LLM fine-tuned to be reliable for all above will be really powerful.
@NielsRogge the interface isn't the hard part! (having conviction was though)
the efficiency was hard but not the hardest part
by far the hardest part is making it smart+robust+reliable 😭 when I first started in this direction, I thought it would take me a week
I optimized SQLite 3. The optimized build achieves a verified geometric-mean speedup of 1.59x over the pristine latest trunk across four benchmarks:
- 2.06x on the official speedtest1 (~30k statements)
- 1.90x on TATP (OLTP mix: 400k transactions, 100k subscribers)
- 1.30x on the Star Schema Benchmark (13 queries x 2, 1.5M-row lineorder)
- 1.25x on kvtest (blob I/O: 40k x 10 KB, seq + random + update)
All 1,032,940 cases in the full SQLite test suite pass, and every benchmark run produces checksum-verified identical results. For context: the SQLite team has spent nearly 20 years tuning this code; their own measurements show ~3.5x total CPU improvement since 2008, earned a few percent at a time. Tested so far on standard Linux provided by GCP.
An agent called KISS Sorcar did it in under 8 hours for under $150 in API cost. I wrote no code: 1 main short prompt plus a couple of steering prompts.
Why I trust the results:
1. No cheating in the speedup. The biggest win is defaulting to WAL journaling, the configuration SQLite's own docs recommend. Re-measured durability-neutral (synchronous=FULL: committed transactions survive power loss exactly as strongly as the baseline), it is still 1.54x faster.
2. Adversarial testing. A separate attacking agent tried to break the changes: a 37-script differential SQL corpus (recursive CTEs, window functions, triggers, UPSERT, JSON, FTS5, rtree, corrupt inputs) compared byte-for-byte against a pristine build under ASan/UBSan, plus WAL-file corruption, multi-process mptest, fd-exhaustion, symlink, and read-only-media attacks.
3. Security hardening and independent review. Two hardening rounds with kimi-k3 as the model: in-tree fuzzers (fuzzcheck over all 8 corpora, sessionfuzz) came back clean; 14 hostile-WAL corruption scenarios - no crashes; OOM injection and page-size sweeps; one real bug found and fixed.
1.59x is what an agent can honestly find in one of the most heavily optimized codebases in the world overnight.
Optimized SQLite Repository: https://t.co/WlQ7Znlk3z
Blog: https://t.co/K51LOdWxpN
Recently met @srush_nlp and he started giving me an impromptu lecture on how targeted on-policy self-distillation works.
I asked him if I could record it on my iPhone.
The basic idea is this: if the model made a mistake at some point in the rollout (for example, calling a tool that doesn't exist), we want to discourage this specific error, but we don't want to just learn from the final reward, because it's a very noisy signal spread out over the whole trajectory.
So we have another model read this trajectory and figure where the error was made. It simply inserts some hint tokens to the part of the trajectory right above where the mistake was made.
Now with these injected hint tokens, have the model run a forward pass. You're not having to regenerate a new rollout - aka no new decode required.
The hint causes the model to assign lower probabilities to the error tokens. You then trains the original model to match these new probabilities, teaching it to downweight that specific mistake.
@AmazonHelp I have not received any response from your team. More surprisingly, I received a call from same harassing delivery person asking for OTP to cancel the order. What is going on?
@AmazonHelp Haven't heard from you since 2 days. The order status is the same and the delivery person still harassing. Do you need any more information to act?
@AmazonHelp I contacted support on call, they assured that item status will be set to delivered but nothing happened. On chat with support again and they have no information about previous call.
Getting AI to gen the MOM from 5-word notes: genius.
Use AI to summarize the 200-pg book so you sound like you read it: strategic.
Ask AI to rewrite it so it doesn't sound like AI wrote it: deeply human.
Using AI to add pg numbers to kids' summer assignment PDF: priceless
AI Writes Great PRs. That Doesn’t Mean They’re Safe to Ship
I think we are finally zeroing in on what the top engineering skills are for the upcoming cobuild-with-AI generation.
@GenAI_is_real Great post and experience report. I wonder why you think of the problem in terms of different kinds/names of agents vs a workflow that routes to different .md files depending on the query. I think latter is an even simpler mental model
I'm beginning to suspect that a key skill in working effectively with coding agents is developing an intuition for when you don't need to closely review every line of code they produce. This feels deeply uncomfortable!
In this video, I explain:
- turn detection, naive vs semantic
- tradeoffs: responsiveness, latency, and false positives - How turn detection is framed as a machine learning (ML) problem
- compare popular implementations
- metrics for turn detection
https://t.co/c5kNUCVdAq
Semantic Turn Detection Explained
Voice AI agents are only as good as their conversations. One crucial element makes or breaks that experience is turn detection. Naive Turn detection based on VAD can make voice agents appear sluggish. 🧵