we, fortunately, saw this coming a few years back and bought our own gpus and self hosted them (in my shed). We've saved close to 10x the rental cost over time.
now, the downside is that I got woken up at 2am because the AC filter needed changing and the room got too hot.
But we have more experiments to run than we can find compute for (thank you to everyone who has reached out).
ps. we're always buying hardware, so if you have some to sell, hmu
Our models study the workspace through thousands of rollouts before deployment.
When solving new tasks, they produce better responses and get to the right answer faster.
With study, we're able to outperform Opus 4.8 X-high while using 3.3x fewer tokens.
RL is hitting a ceiling with human feedback. What if the world itself becomes the signal?
Join us at the RLxF: RL from World Feedback 🌍 workshop at ICML 2026 @icmlconf tomorrow (July 10th)!
Web page: https://t.co/cN0itnL1yI
The best teams of the next decade will be small enough to fit in one room.
we're building @renaibuild, the infrastructure for autonomous companies. we're a small team with strong opinions on who we work with. agents do most of our chores, so every human we add to the team has to count.
we are looking for people who:
- think in loops and agentic systems
- who can smell AI slop from a mile away
- have high ownership and high agency
- are neurotic about the tiniest of details and hunt for clarity
we're not looking for people who outsource their thinking to a chatbot or need a defined process to get things done.
your work is your resume. show us what you've built and we'll move fast. hiring across growth, content and engineering.
links below.
Introducing the RL from World Feedback workshop with a list of great co-organizers and speakers!
Don't hillclimb on benchmarks. Hillclimb on the world! (Yes, I only give talks or organize workshops with "World" in it)
Note: RL means many things including rejection-sample SFT.
RL is hitting a ceiling with human feedback. What if the world itself becomes the signal? Introducing RLxF: RL from World Feedback 🌍 workshop at ICML 2026 in Seoul, South Korea!
Speakers: David Silver, @chelseabfinn, @jessezhang, @robertarail
Web page: https://t.co/tHdX6PIebL
Update: Our ICML 2026 RLxF workshop accepts submissions in NeurIPS format! Consider submitting your NeurIPS papers to the workshop.
- Full paper (up to 9 pages)
- Short paper (2-4 pages)
Deadline: May 13, AoE
Web page: https://t.co/cN0itnL1yI
OpenReview: https://t.co/dzpZTJ4qRn
Running a single simulation of Borg, @Google's massive compute cluster takes between 1 and 18 hours.
Incredible work from @GoogleResearch (@yashakha, @SagiPerel, @XingyouSong et al.) replaces it with a 60M parameter encoder-decoder trained from scratch, achieving 0.99 rank correlation and 100x lower MSE than tabular regression! 🧵
(1/n) Can the likelihoods of flow-matching/diffusion models be learned by another model, so we never have to compute the costly change-of-variable eqn?
In our NeurIPS 2025 paper, BoltzNCE, we show that this is possible and use it to recover the Boltzmann distribution
🧵👇🏽
Add-on: from our experience with Regression Language Models which absorb evaluations via weight updates - it takes at least *thousands* of (x,y) pairs to learn any meaningful world model of the objective.
This makes sense because information theoretically, you have lots of changing string features which naturally creates a curse of dimensionality for predicting the objective value.
If we assume that in-context learning has the same order of sample-efficiency, then eg Gemini would need *billions* or *trillions* of token lengths to do end-to-end optimization.
Maybe this is too pessimistic - but as long as we stick to the current paradigm of “fixed weights at inference”, we’ll be stuck to using evolutionary wrappers for LLM-based optimization.
"Net effect, 1 text model replaces many bespoke predictors across languages, graphs, and hardware." -- Thats the vision, try out universal regression!
Paper: https://t.co/fgfFQOSmD5
Code: https://t.co/CtTQQrH28y
Datasets/Model: https://t.co/T3ZtxZGHGt
New Google+Cornell paper shows 1 compact language model can read code and predict memory, latency, and accuracy across languages and hardware.
A 300M model hits 0.9+ on APPS memory and leads classic neural architecture search predictors.
The task is code to metric regression, predict memory or runtime from code without running it.
Past systems rely on hand tuned features per language or graph, and they break when code changes.
This model reads raw code or ONNX graphs with a T5Gemma encoder and predicts numbers digit by digit.
Sequential prediction lets 1 model learn many tasks and capture tradeoffs like accuracy versus latency.
Digit tokenization avoids normalization across mixed scales and beats a mean squared error head.
High rank correlation means it ranks candidates well, like picking the lowest memory solution.
Language pretraining and synthetic regression pretraining speed training and raise accuracy while removing brittle feature engineering.
Net effect, 1 text model replaces many bespoke predictors across languages, graphs, and hardware.
----
Paper – arxiv. org/abs/2509.26476
Paper Title: "Regression Language Models for Code"
Language Models are surprisingly precise guesstimators! We believe this is a key step in the ‘AI Scientist’ vision, serving as a cheap, uncertainty aware feedback that can consume any data (unstructured notes, code, even multimodal!)
Regression Language Models (RLMs) can consume code, server logs, and even graphs to predict outcomes of code. Think predicting accuracy of models with 200+ nodes before the model is even trained, latency of triton kernels on GPUs, server-scale hardware utilization with nearly no feature engineering.
Paper: https://t.co/JMMEyXthuq w/ @XingyouSong@mohsaied
Code: https://t.co/kTkAhe6zPE
Datasets: https://t.co/UsyI8TNg8e
There is a toxic type of person to avoid— the kind obsessed with hierarchy, constantly ranking everyone: net worth, IQ, number of publications,academic titles & citations, cycling FTP, 5K running speed, education, etc.
They are miserable & want you to be miserable too.