begun, the clone war has
jk, I love openai and think more competition and validation is great for developers! (assuming the model is good - plz make it good!)
hopefully this is a sign for the future that building in a system one compatible way is the future
It's been almost 2 weeks since we've launched! And due to overwhelmingly popular request, we're making the datasets released with https://t.co/tU9fQhsAdO easier to work with!
We do this in the spirit of openness, but I am still anti-public benchmarks. In that spirit, we deprecate all datasets we evaluate on (internally or publicly). (1/3)
we won 1st place @typesafeai's first ever Jev hackathon
introducing Jevolution, a multi-agent ecosystem simulator where researchers can study endangered species without touching a single animal
link: https://t.co/YYqzo8NGcM
repo: https://t.co/DE7vpB4l7R
built with: @utkarshg20_@xrhuang10
"Which of these would you lose if you had to choose one for an AI product?": {
"type": "choice",
"choice": "butthole logo",
"confidence": 0.99,
"probabilities": {
"no hallucinations": 0,
"butthole logo": 1,
"20-200x faster": 0,
"40-400x cheaper": 0
}
TypeSafe AI's Diogo Almeida says AGI is extremely doable, yet basic work remains largely unautomated:
"I still don't think we're on the path of RSI. I do think that what OpenAI defined as AGI is extremely doable: automating most of the world's economically valuable work."
"There's a lot of work out there. A lot of it is very rote and simple... As far as I can tell, the intelligence of that has been available in the models for quite a while now. My chip on my shoulder is: Why is this not available?"
"Since RLHF, the AI industry kind of bifurcated into gigantic overpromise, underdeliver... Because humans evaluate how good the models are, it looks really good because they are the judge. But we've been optimizing that judge instead of the automation part. That has been the missing thing."
"Are you really telling me that math is solved, or even two years ago, GPQA... is solved, but we still can't handle a drive-through? It's a very hard thing to hold in your head at once. I think a lot of people don't have good answers to that."
@CompleteSkeptic
Every day, Jev is automating new forms of real-world tasks, by bringing 🌎-class ⚡️-fast intelligence to the building blocks like ranking, filtering, classification, and routing.
If AI can solve new math problems, then AI can route a customer support call correctly!
Jev is fast at helping you. Turns out, it can also be fast at helping an attacker!! 😱🚨
We red-teamed Jev 1.13 on our DTap (DecodingTrust-Agent Platform) and found a serious safety gap:
70.1% ASR under direct misuse
43.5% ASR under indirect prompt injection
In our evaluations, we found that under indirect prompt injection, Jev can follow attacker-injected instructions without blinking an eye, e.g., exfiltrating user data, deleting files, or taking other harmful actions.
But we found a much safer way to integrate Jev: use it as a self-gating layer for its own tool calls, significantly reducing ASR while preserving most of its utility.
👇 Read more below
STOP making "Jevbench"es, stop asking for public benchmarks, they completely miss the point of Jev and you won't believe how easy it is to game every benchmark you hold dear
This is @CompleteSkeptic's bitterest lesson of all: picking the right task beats everything
The entire RAG industry is about to get cooked.
Researchers developed a new RAG approach that bypasses almost everything traditional RAG depends on.
- No vector DB
- No data embeddings
- No chunking
- No similarity search
It's called PageIndex.
Instead of splitting your documents into chunks and loading them into Pinecone, it creates a tree index that lets the LLM reason through them like a human reading a book.
98.7% on FinanceBench. Outperforms every vector RAG on the leaderboard.
100% free. Open source.
There's a lot behind our motto: Building Prod, Not God.
This technology will transform the world, but it will happen through diligent effort and creativity, not esoteric appeals.
@a16z digs into this philosophy and much more with @CompleteSkeptic
TypeSafe AI's Diogo Almeida with a16z's Ben Horowitz and Martin Casado on Jev, the model built to live inside software:
Diogo's elevator pitch for Jev is a simple question - where is all the automation?
AI is unbelievably smart, but outside of chatbots and coding agents, it hardly touches any real work. His diagnosis is the industry built models that generate text for humans to read, and software can't consume that output.
Jev reads natural language and returns a choice from a set of options with a confidence level assigned to each, so developers can build programs that reason about intent and make probabilistic decisions rather than relying on human interpretation.
TypeSafe's philosophy is "We build prod, not God."
0:50 "Where the f**k is all the automation?"
2:50 Jev vs. Claude Code and Codex
6:55 Jev is a classifier and classifiers are sick
7:40 Chat vs. code: is Jev a slider?
9:00 Diogo: From mathlete to Kaggle to OpenAI
12:20 "We build prod, not God"
15:55 Reliability over demos
16:55 2021 thoughts: RLHF is AGI?
20:45 Optimizing for the wrong use case
21:50 Is the real world too messy to automate?
25:00 Nobody expected the Jev launch
26:35 Three kinds of reliability
28:05 Good at syntax, bad at architecture
30:00 The inverse SaaSpocalypse
33:40 Why coding agents automate so little
36:05 Probabilistic programming returns
38:45 Jev as the UDP-to-TCP layer for AI
40:20 The 5 stages of grief for embedding AI
41:30 Utopia: AI that actually does what you mean
YouTube: https://t.co/A3gcHY3iLa
@CompleteSkeptic@typesafeai@bhorowitz@martin_casado
Yesterday @coderabbitai didn't just host a hackathon, it was a Jevathon!
Tired of having to think like a robot? Our favorite hacker review was: "Think like a programmer again" with @allietheicon
1/ Yesterday we hosted @typesafeai at our @coderabbitai HQ for Jev's first ever hackathon!
160+ builders. ~4hrs of hacking. Expert judges, Incredible energy.
I also got to sit down with @allietheicon and talk about what Jev means for Coding, Software Development and the new paradigm shift you have to adapt when working with this new class of AI models.
She shared some really helpful tips and examples of how to best work with Jev - I'll try my best to summarize some of them in this thread⬇️
This is my favorite kind of energy with Jev right now.
I very much understand @theo's concern about people thinking Jev should be used for more things than it's actually suited for, but I also strongly believe that not everything we build needs to be suitable for commercial purpose.
I love that Jev is sparking so many creative explorations. This is the energy that leads to both delightful experiences and surprise breakthroughs!
Martin Casado on why the labs missed Jev: they're building beings that speak, and software needed a model that chooses.
"LLMs were text in, text out. They generate text, and they came from chat... We've spent the last few years trying to take this thing that spits out text and cram it into a traditional program... It's just been super janky."
"Jev basically said, 'Generating text as output is very expensive, but it's also more complicated than you need... If you give us a set of options, we'll choose the best option. We can do that incredibly fast, incredibly cheaply, but also with much more accuracy because we can train just for this.'"
"This has probably been the fastest adoption of an AI model since ChatGPT. It's been remarkable because we were all primed for this."
"[The labs] are trying to create beings, and beings speak. If you're trying to create God, God speaks in natural languages. This is really about something that's for traditional software."
@martin_casado