If you know me from twitter, you might've seen me write a lot of tech articles on DevTo. A long time, and a lot of things changed (in my life and in the world) since I last wrote anything, but, well, let's try.
Obviously, the comeback post is about AI 🏜️
https://t.co/mygDrJ4cGV
I think the best use case for Jev so far is to act as a judge of which tool the LLM should call. Instead of spending a lot of time and money on reasoning models to decide which tool to call, we could delegate it to Jev.
Then it replies back quickly with the selected tool.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
For anyone curious how Jev works, I made a visual explanation using @claudeai :)
This is based on the Qwen2.5-RLCD model which @harshagundal released on @huggingface
The idea is to replace autoregressive LLM generation by a single Transformer decoder (of a pre-trained LLM), which processes the context + JSON schema only once. The keys and values of those tokens are cached.
Next, for each field of the JSON schema, we:
1. pass its field suffix tokens through the Transformer decoder again (reusing the KV-cache)
2. obtain a final hidden state, which we pass through the language modeling head
3. we obtain scores, also called logits, for all tokens in the vocab of the LLM
4. we only look at the scores of the tokens we care about for the given field, and pass those through a softmax to obtain probabilities which sum to 1
5. we take the token with the highest probability.
The benefits of this are that:
1. it's fast (we don't need to generate the JSON schema token by token)
2. it's 100% valid JSON (we don't need to rely on the model to generate a valid schema)
@RealGalego Entender algoritmos e estruturas de dados agora é ainda mais importante, codar talvez não. Conceitos de computação mesmo como você mencionou são muito importantes
If you’re a CS student, I’d suggest writing your thesis around SDCs (silent data corruptions) and efficient algorithmic fault tolerance.
We’re quickly going to need to accept the (un)reliability of computing systems again!
Every indicator (density, lower voltage gating, raw scale) is moving towards SDCs becoming a more serious issue over the next decade. And let’s not forget when all of this stuff starts to go in space and we have to deal with cosmic events on “not-super-rad-hard” CPUs+GPUs!
No, you can’t solve it all in hardware. It’s prohibitively expensive; and the long tail of “potential SDCs” is just too long. One bad apple ruins your pie. One bad GPU can inject repeated corruptions into otherwise noise-tolerant workloads (*cough* AI training).
If you continuously accept a corrupted value from say…a bad tensor core, that bad value can quickly spread across the whole cluster. If it happens again and again and again and no one notices; it’s like putting a little bit of spin on a bowling ball. Eventually, the overall trajectory ends up widely different!
When you start to imagine this stuff being in space (cosmic bit flips), and quantized(!), each remaining bit carries more critical information. Long term we’re going to have to accept clever, low-overhead software algorithms.
Accepting say, a ~3% perf loss can be *absolutely* worth it if it increases your odds of detecting a mercurial core enough!
Somehow a 100-line SQL query is considered unmaintainable, but splitting the same logic across five services and a message queue is considered architecture.
Introducing Projects, a new way of working in Cursor.
Rather than creating a chat for every task, you work with a coordinator agent in a single, persistent thread.
Like @bot, your agent is always on, proactively manages work with subagents, and improves over time.
Introducing WalShadow: sub-second Postgres replication to ClickHouse from physical WAL.
No logical replication. WalShadow decodes the same physical WAL stream Postgres replicas use and writes ClickHouse-native blocks directly into ClickHouse. In our benchmarks: ~200 ms commit-to-visible latency and 289K rows/sec sustained, keeping pace with the source Postgres.
Open source and available today: https://t.co/BE2qDichVu Blog: https://t.co/540PF6jvhn
I hereby declare the "learn to code" era officially dead:
Big declines in the number of people studying computer science in the last year or two 📉
Chart from this week’s edition of our newsletter on AI and the labour market https://t.co/y4yKn6GTl1