We're pushing updates to cookbooks, best practices, and use cases as fast as we can, but the community is 1000x-ing us here.
The Factorio-ing is real too; some of the highest leverage additions are automating your automation, like teaching your agents how to use Jev directly in your applications. One example in the wild: https://t.co/hTcW4QMU59
tl;dr Jev makes web search more context-efficient. I'm open sourcing a thing that does this for you: https://t.co/JcHYc06r57
So, I ran out of Codex tokens last week, then ran out on my backup account. Same for my Claude tokens.
So I did a little context usage audit, and one of the biggest culprits was web research -- I "read" a lot of academic papers. Like, all of them.
I figured TypeSafe's new Jev model could help me make something more efficient than burning expensive Astra and Fable tokens parsing through web results in the raw. Plus I needed a search CLI for my Pi/Kimi workflows.
So I made this, and it's pretty dope!
3 months ago, an anonymous account commented "Y'all are sleeping on the most important talk of this conference". The speaker was @CompleteSkeptic and his talk was on RLHF's deal with the devil, and the argument behind Jev.
Diogo argued that training AI to produce answers people prefer comes with a trade-off. You get polished, coherent responses. You lose the unusual alternatives through mode collapse.
The problem is that the unusual alternative can be the correct one.
Three months later, @typesafeai's Jev, a model designed for decisions, is now generally available to everyone.
If Jev has caught your attention over the weekend (the internet is on fire with it after all!), Diogo’s @AICouncilConf talk is worth watching. It lays out the argument behind what his team is building and why they chose this direction.
Diogo's talk: https://t.co/627kG7eMOh
Text classification in MotherDuck just got ~50x faster at ~1% of the cost.
prompt_jev() is a SQL function powered by Jev, TypeSafe's new system one model. 100k rows: 40s, $0.50, frontier-LLM accuracy. The LLM took 32 min and $37.
Read on:
https://t.co/XIE2cw6rUS
@deepfates damn feels like a hack though. Of course the distribution of internet text would think a thing thinking to itself is more likely to have qualia
can @typesafeai Jev count the r's in strawberry?
no.
47% says 3, 47% says 2. a coin flip, same as every LLM.
On 70% on 168 test words, it undercounts doubled letters like everyone
then I gave it the letters instead of the word: ["s","t","r","a","w","b","e","r","r","y"]
168/168. same model, same question, 260ms
@CompleteSkeptic is this expected?
One of the most fun parts of this company is @CompleteSkeptic getting to go off on any topic he feels like either nerding out about or shitposting on. Always a terrific listen.
Jev and the System One Model: RLCD, intelligence/$, reliable AI, & the end of chat-first AI https://t.co/H2bZXCENyW
@typesafeai CEO @CompleteSkeptic explains why AI can solve extraordinarily hard problems yet still fail to automate basic work, why Jev is built for reliable decisions inside software instead of chat, why TypeSafe rejects public benchmarks and refusals at the API layer, why data and the right task matter more than brute-force compute, how System One Models could reshape coding agents and software, and why even with $1 billion he wouldn’t pre-train a model from scratch.
I was honestly surprised by the number of absolutely sick computer use demos that hit the TL right after we launched; we know that we're going to dominate in that space, but user creativity always shines brighter than you can predict.
@browserbase was so on this that they hosted a Jev Day hackathon less than a week after our launch, and @kylejeong has a thorough guide on what Jev's for and how to get Jev working for you.
Wait till y'all see what the image model can do 😎
@Siddharth87@browserbase Editorial with semantic search goes hard. I already use Jev, e.g., for targeted edits in anti-AI writing patterns across any auto-generated content
@tenobrus “Engineers want control again” misses the point; obviously agents can and will write system one calls into code themselves. What’s been missing is appropriate primitives to apply intelligence exactly where and how it’s needed to dramatically improve overall intelligence per $