If you've wondered where I've been recently, I gave an interview. This is a very difficult time for me. I'm trying to bring my mom home, and I could use your help spreading the word. Please watch, and please share.
interview: https://t.co/rloipKmoWW
website: https://t.co/Z7pubzoP0A
Super cool work from @VMSuriyakumar & team!
They achieve 100% recall on CSAM-specialized LoRAs without ever prompting models to generate harmful content. Super impactful work!
Can we determine whether a model has been trained to generate harmful content without ever producing a single image?
In our new paper, which won an Outstanding Paper Award at the AI4GOOD workshop at ICML 2026, we answer yes.
We introduce Evaluation without Generation, a new paradigm for auditing models without producing outputs, and Gaussian Probing, the first scalable method for detecting harmful specialization directly from model weights.
We demonstrate it by detecting CSAM-specialized LoRAs, a task where standard evaluation is impossible because generating even a single image is illegal.
📄https://t.co/nWKnlQnqVm
🧵
Can we determine whether a model has been trained to generate harmful content without ever producing a single image?
In our new paper, which won an Outstanding Paper Award at the AI4GOOD workshop at ICML 2026, we answer yes.
We introduce Evaluation without Generation, a new paradigm for auditing models without producing outputs, and Gaussian Probing, the first scalable method for detecting harmful specialization directly from model weights.
We demonstrate it by detecting CSAM-specialized LoRAs, a task where standard evaluation is impossible because generating even a single image is illegal.
📄https://t.co/nWKnlQnqVm
🧵
Excited to share our recent work on how we approach distilling investor taste into models to hill climb on financial tasks!
Had a great time working with @kevinbzhu@XiaoEmily41333@rohanalur@ddkang
on this blog post :)
Sorting which financial docs are worth an analyst's time is surprisingly hard for frontier LLMs. With an expert-labeled dataset and on-policy distillation, Bridgewater fine-tuned a model to do it reliably and cheaply.
https://t.co/gyYzXq15zd
We’ve raised 75m in new funding from Sequoia and Spark Capital—partnering with @sonyatweetybird, @MikowaiA, and @YasminRazavi, all of whom are deeply supportive of our long-term mission. We’ve also brought on angels & advisors including @karpathy, @tszzl, and @_milankovac_.
-----
Our early results with FDM-1 moved computer use from a data-constrained regime to a compute-constrained one; this latest round of funding unlocks several orders of magnitude of compute scaling for that work. With the FDM model series we have a path to scale agentic capabilities through video pretraining, and we expect to achieve superhuman performance on general computer tasks in the same way that current language models have superhuman performance on coding tasks.
We’re also now able to invest in the blue-sky research necessary to our long term mission of building aligned general learners. To realize the civilizationally transformative impacts of AI, models must generalize far out of their training distributions, actively exploring and building skills in new environments. This capability represents a substantial shift from the current paradigm of model training. We believe that current alignment techniques are insufficient to predictably and safely steer a model with human-level learning capabilities, and so we’re doing work to study small versions of this problem in controlled environments to develop a science of alignment for general learners.
We’re a team of 6 people in San Francisco. We’re hiring world-class researchers and engineers to help us achieve our mission. If that’s you, please get in touch.
Unclear if builders aren’t buying the “hype around AI”; I think more so people are growing more skeptical about hastily done, low impact projects made with AI. I will say that I think majority of projects done at hackathons nowadays are in some part vibecoded (including some very impressive projects)
Absolutely! I think that the crux of my argument is that AI is a super powerful tool + can definitely be used to push forward progress on genuinely impactful projects, but current reward structures are misaligned with this and as such, many people are using AI to accelerate the development of slop instead.
Part of the reason I joined Cerebras was because I felt that the team was full of people who genuinely wanted to solve big problems rather than chasing shallow wins :)
one of my favorite parts of this job is getting to sit down with people like Sarah Su
she's a junior at MIT studying AI, organizing HackMIT, interning at @cerebras, and she has a really unique perspective on why her generation isn't buying into the AI hype the way you'd expect
Sarah dives into her first-hand experience with AI in the classroom, declining hackathon projects, and what the next generation of MIT CS/AI students are excited to build.
Guest @SarahhhSuuu
Producer @alyciazcary
Computer use models shouldn't learn from screenshots.
We built a new foundation model that learns from video like humans do. FDM-1 can construct a gear in Blender, find software bugs, and even drive a real car through San Francisco using arrow keys.