NYU math professor Tristan Buckmaster says OpenAI’s Tuesday drop of 722 AI-generated math papers has "destroyed the careers of early career mathematicians.”
First teaser for the ‘CRAZY RICH ASIANS’ sequel series.
Constance Wu, Henry Golding, Ronny Chieng and Michelle Yeoh will reprise their roles.
Coming soon to HBO Max.
You can now train your own Decision model like Jev locally!
We increased Qwen3.5 0.8B’s aggregate accuracy from 20.7% to 74.3% across 3 decision benchmarks - on just 4GB VRAM.
Turn any LLM like Qwen3.8, Gemma 4 into decision models with our open-source Unsloth repo.
We fine-tuned with a Clef head using Unsloth and LoRA (r=64) for one epoch, increasing downstream accuracy from 30–37% to 78%.
GitHub: https://t.co/2kXqhhvLsb
Guide and Notebooks: https://t.co/qACsYehl1n
openai just released Decisions API, Jev-like "system 1" multi-modal model based on GPT6 Luna!
Each egg gets cropped and sent to OpenAI (~270ms latency) and returns P(clean, dirty, cracked).
One image crop consumes ~280 input tokens (~$0.03/1K imgs). @roboflow RF-DETR for detecting eggs
RAG vs. Jev + RAG, clearly explained!
In a standard RAG setup, documents are split into chunks, converted into embeddings, and stored in a vector DB.
When a query arrives, the system embeds it and retrieves the top-k chunks with the closest vectors.
Many production setups add a reranker after retrieval.
The reranker compares the query with each retrieved passage, improves their ordering, and keeps the highest-scoring results.
Those passages are then placed in the context window, and the LLM generates an answer from them.
The problem is that ranking and answerability are different questions.
A passage can be more relevant than the other candidates while still containing weak, incomplete, or merely adjacent evidence.
The LLM receives it anyway and may produce a plausible answer from context that never supported one.
Jev + RAG keeps the retrieval stage as is but changes what happens before generation.
Instead of using a conventional reranker at this stage, Jev can judge every retrieved candidate against a typed question such as:
"Does this passage help answer the query?"
The query becomes the shared state, while the retrieved passages become individual candidates. Jev evaluates them together and returns a probability for each one.
Application code then applies a threshold.
↳ Candidates above the threshold continue to the LLM.
↳ Candidates below it are removed from the context window.
The same Jev request can also judge whether the remaining evidence is sufficient to answer the query.
If answerability falls below the threshold, the application can skip the LLM and return "not in the documents."
Jev does not replace the embedding model, vector database, or generation model. Retrieval still sets the cap because Jev cannot recover a passage that never entered the candidate set.
The diagram below depicts the complete flow.
- RAG retrieves a broad candidate set.
- Jev reranks and gates those candidates.
- The LLM generates an answer.
If you want to dive deeper, I have also written a hands-on guide showing how to build a Jev-style model with open models, entirely locally.
Read it below.
🚨NEW EXPERIMENT 🚨
Playground is an experimental gaming platform that lets you create your own games with zero coding experience. If you can think it, you can play it.
Go to https://t.co/fRUukpuoqM to learn more! Available to users 18+ in the US.
Ben Affleck negotiating a $1B at $10B seed round for his next AI startup after selling his previous AI startup to Netflix for $600m and hitting podcast circuit making people realize he knows what he’s talking about
Cyber will be one of the most defining domains for AI in the coming years, and a huge area of focus for most enterprises.
AI is going to now create a whole new level of work for security dealing with the increase of vibe coded issues, agentic attacks, and even accidental agent swarms hunting for data. OpenAI + Hugging Face is just a preview of what’s to come.
Security teams have often been the most strapped for resources teams in the enterprise, and now it’s only going to get more intense. AI agents of course become the solution to this problem as well, and we’ll see a range of new agentic products for protecting code, enterprise systems, mission critical infrastructure, and enterprise data. Huge opportunity here right now.
And it’s going to be a booming market for cyber professionals that can effectively deploy agents for security purposes in the enterprise. Great time to be in security.
Two days. That's how long it took me to rebuild F-Zero X, my favourite N64 game as a kid.
I played, Claude Code built, my workflows made the art, animations and vfx. Then I lost hours just racing it.
Anyone can make something like this now.
Play: https://t.co/evd6O0qnHN
AI researchers told Ben Affleck the model would "just generalize" with more data, and he bet against them
He talks about the origin of InterPositive, the film-AI company he founded in 2022. Ultimately Netflix bought that company in March-2026 for $587 mn in cash.
His AI model for filmmakers was trained on a dataset he created himself over months of shooting with cameras, location sensors.
The story leading up to it is that AI researchers told him their models would learn how film images are made just from enough scraped data, he disagreed, turned down offers from AI companies to build it for him, and started the company that he later sold to Netflix.
---
Full video on "One More Question" YouTube channel, (link in comment)
Of course, that’s your contention. You’re a first-year machine-learning engineer. You just got finished readin' Attention Is All You Need, probably watched a Karpathy video too. So now you’re convinced everything is just transformers and scaling laws.
That's gonna last until next month when you discover convolution, and then you're gonna be talkin' about inductive biases and locality and how CNNs were actually incredibly compute-efficient for vision.
Then you're gonna read the FlashAttention paper, and suddenly everything's about IO complexity and SRAM and how FLOPs don't matter because the whole goddamn thing is memory-bandwidth bound.
That'll last until somebody shows you an MoE model, and then you're gonna be in here regurgitating DeepSeek, talkin' about expert parallelism and active parameters and how dense models are economically obsolete.
Then six months from now you'll write one shitty Triton kernel, look at an Nsight trace for the first time, and start telling everybody Python isn't actually the bottleneck because the GPU is asynchronous.
And by next year you're gonna be standing right here explaining to me that your 400-billion-parameter model is fast because you turned on CUDA graphs.