You could have been doing this with Scala/JVM, C#/.NET for last 10 years at the least if you had the attention span longer than a gnat^Hrubyprogrammer.
Caveat being higher memory usage than Rust, but WAY better readability; readability is a wash with a good IDE (not vsc**e, lol)
Wrote about Nikole Hannah-Jones’s latest essay, upper middle class parenting, and the delusions of moral, individual choices. Also a defense of teachers. https://t.co/opISPPHH9w
Here is an interesting use case of HyperLogLog you probably have not seen before - predicting whether adding more RAM to your database will actually help.
I was going through Neon's engineering blog and found that their autoscaler uses HLL to estimate the Working Set Size (WSS) of a database. It counts unique data pages accessed over time and, from that estimate, decides whether scaling up memory actually makes sense.
Quick background, if you are unfamiliar: WSS is the amount of data actively being accessed during a given time window. If your database is 1 TB but only 5 GB is being queried regularly, you do not need massive memory - just enough to hold that 5 GB working set.
This is where HLL fits in nicely. Instead of tracking every unique page exactly (which is expensive in both memory and CPU), HLL gives you an approximate count of distinct pages accessed with very low memory usage (a few KB), fixed space regardless of input size, and error rates around 1-2%.
Over a sliding time window, the system observes page accesses, feeds page IDs into an HLL, and gets an estimate of unique pages accessed - that is your approximate WSS.
The scaling logic is now simple: if the WSS fits in the local file cache, more RAM helps. But if the workload is too large, adding memory changes nothing - you end up paying for resources that cannot move the needle on cache hit rate.
So what it answers is this one interesting question - "Will this memory actually be used by hot data, or is it wasted?"
Probabilistic data structures are fun, ngl - somehow I keep finding them in the least expected places :)
When talking about Ehrlich today, spare a moment to note that he set the template that the modern environmentalist movement still uses:
apocalyptic forecasts that don't bear out, but which lead countless people to make disastrously bad personal & public-policy decisions.
the magic of cowork and openclaw and other AI products is that they replace our giant row of infinite browser tabs
And lol - no, don't feel guilty, I have too many tabs too. AI makes it so that every workflow that required 4 browser tabs and a spreadsheet is getting collapsed into one AI-native experience
Just as one quick example-
think about how you used to research a person or a company: LinkedIn tab, X tab, Google tab, notes doc, slack open. now one prompt does it in 10 seconds. the "tab count" of a workflow is basically a proxy for how much AI can compress it
if your product eliminates 6 tabs and a copy-paste loop, users will like it. If you can create a whole series of these workflows then your users will absolutely love it. Thus the biggest opportunities are workflows where people currently alt-tab 20+ times per task. Sales, recruiting, research, compliance, procurement. Boring? yes. Massive? also yes. But this is why these agentic tools are going to crush
AI doesn't need to be superintelligent to be wildly useful. it just needs to be good enough to close the tabs
Cuda to any other modern GPU abstractions (RocM, oneAPI/sycl) is relatively child’s play for Claude code and codex these days.
3-4 months ago I tried getting cuda code to port to Tenstorrent TTNN abstractions…that was much harder but got several kernels working with iterations and guiding the agent manually
Performance is another matter…you can make progress if you can guide the agent with examples, you need to have a good depth on the architecture you are targeting though..
The next step I’m excited about is skipping the intermediate abstractions entirely..why not CUDA to GCN straight?
Take it one step further, why not “math equations to direct hardware assembly “ ? Why do agents need to talk Python or CUDA with hardware?
These are exactly the questions we are researching and solving for at Oxmiq
I'm a virologist and wanted to clear up something I think vaccine "choicers" might miss in the latest data.
Of over 2200 cases of measles in 2025, 93% of were in people who were unvaccinated. You might think that was their choice. But 575 were in kids under 5, and...
NEW: UC San Diego has released a new report documenting a “steep decline in the academic preparedness” of its freshmen.
The number of entering students needing remedial math has exploded from 1/100 to 1/8.
They’ve had to create a second remedial class covering elementary and middle school math skills in addition to the one covering gaps from high school.
🧵
Empire of AI by @_karenhao is by far the most accurate telling of the era when I was at OpenAI, which was an important few years – from the first commercial step to shortly after the launch of ChatGPT.
There is one important piece that is incorrect: the portrayal of @sama
He’s presented as some machiavellian and reckless leader and the facts don’t support that.
I joined OpenAI when we were about 100 people and purely a research lab. As head of product, I helped transition OpenAI from a research org to one deploying our research as products. During this time a number of large and complex decisions were worked through. There were no easy and obvious solutions to any of these and many of these decisions were seemingly at odds with past decisions.
Complex situations often look very different to people and there were dynamics at OpenAI during this time that made everything more challenging – from the org’s structure to philosophical belief structures and much in between.
The weirdness of OpenAI at this time appealed to me – the unusual structure felt like it created space for something different and the differing beliefs (while exhausting at times) felt necessary for navigating genuinely novel territory.
But that same weirdness created real tensions as we worked through three major challenges. First, the Microsoft partnership: how do we take billions from a tech giant without compromising independence and our mission? Second, productization: how do we go from a research lab to shipping products without abandoning our original purpose? Third, deployment: how do we deploy AI research fast enough to matter while being careful enough to be responsible?
In the moment, none of these had obvious answers. The right path forward was uncertain, and reasonable people disagreed – often strongly – about what we should do. Led by Sam, we worked through each of these tensions carefully and deliberately. With the fullness of time and the ability to see how things actually played out, I believe the evidence shows we reached the right decisions on all three.
When negotiating the early Microsoft deal the entire term sheet was shared with everyone at the org. We’d add questions and comments and then Sam would host an endless meeting where we’d talk through the questions, discuss the spirit of what we cared about, gather feedback on what missed the mark, etc. Each iteration of the term sheet, month after month, progressed like this. Some opposed the partnership, but their voices were always heard and attempts to address their concerns were made. In hindsight, a deal of this sort was required – there was no other viable path – but Sam ensured that our independence and our mission were preserved while spending time working through everyone’s concerns.
The first product roadmap spent considerable time articulating why shipping product supported our mission and how we could do so safely. I spent significant time working through my colleagues’ concerns about productization because getting buy-in across the org on the why was essential to doing it right. With Sam’s full support, we consistently slowed down our product work and made decisions that hurt our business and metrics. We refused to allow entire use-cases we felt we couldn’t handle responsibly. We learned what was required – technically and operationally – to comfortably support select use-cases and prioritized that work. We fired some of our biggest customers because we were concerned about misuse. We didn’t get everything right during this era, but we did an excellent job identifying, sizing, and mitigating risk while building one of the most widely-used products in history. This wasn’t luck, it was the result of the deliberate, sometimes frustrating culture Sam insisted we work through.
On deployment, many of us believed that deployment was essential to the safety strategy (not separate and something to fear). Learning to deploy the research responsibly would require practice, and the time to practice was when the stakes were lowest. And so we embraced an iterative deployment strategy. While other labs struggled with misuse and PR crises, we consistently deployed without major incidents and we learned and improved with each model release. We all understood that being able to shape the norms and standards of AI was critical to our mission. Sam argued that writing policy memos could only go so far and we’d be in a much stronger position to define norms aligned with our values if we were consistently the first to deploy responsibly. His argument proved more correct than many of us realized at the time.
One question I’ve reflected on a lot is why brilliant, well-intentioned people have such different views of this era and Sam’s leadership. I have respect for many who have framed Sam’s leadership negatively, and count many of them as friends, and so it’s somewhat uncomfortable to share my conclusion. Over the years, when I’ve listened to people share examples of what they saw as problematic behavior, I’ve noticed that it often traces back to one of these dynamics: someone who lost an internal debate and attributed it to bad faith rather than legitimate disagreement; someone who struggled to accept that complex situations made previous plans untenable; someone unfamiliar with how large organizations with multiple stakeholders actually function; or someone who pursued power and lost.
I don’t say this to dismiss the substance of these perspectives – the concerns about Microsoft, productization, and deployment were real. But I think these underlying dynamics shaped how people interpreted complex, ambiguous situations. When I joined I was told we’d only ever be 200 people. For reasons I understood, we had to abandon this idea. I didn’t feel lied to or misled. I understood we were navigating novel territory where plans had to evolve. Not everyone experienced it that way, and I understand why. But those different experiences don’t mean Sam was acting in bad faith. With several years of distance, I believe the major decisions from that era have held up remarkably well. That doesn't mean we got everything right or that the concerns weren't legitimate – but it does suggest Sam was navigating these tensions with more wisdom than many give him credit for.
Transformers are great for sequences, but most business-critical predictions (e.g. product sales, customer churn, ad CTR, in-hospital mortality) rely on highly-structured relational data where signal is scattered across rows, columns, linked tables and time.
Excited to finally share what I have been working on over the last year: a Foundation Model architecture which brings the power of Transformers to relational domains, enabling large-scale pretraining and zero-shot generalization in enterprise settings. 🧵1/n
Diviseema, the Krishna river Delta region, known for it's canals, lush green fields, where the Krishna meets the Bay of Bengal at Hamsaladeevi.
Not as widely known as Konaseema, in fact most assocciate it rather with that notorious 1977 cyclone that devastated the region.
The scaling laws of AI are brutal and limiting diminishing returns to AI improvement. They dictate that to get linear gains in machine learning accuracy, you need to exponentially increase training data / compute / model size.
We've known about these since circa 2000, long before LLMs or transformers. They're common across multiple ML approaches. They appear to be fairly fundamental to our current suite of methods.
And they point to dimishing returns and saturation of AI quality (at least with current approaches). Exponentially increasing inputs --> linear gains. In the real world, progress plateaus or at least slows to linear.
They also create a dampening function that makes dreams of recursively self-improving AI look far-fetched. We should expect that every quanta that you improve an AI by leads to less than that quanta of further improvement in the next generation. The system rapidly plateaus, rather than achieving takeoff.
Will this change some day? Maybe! But that will require algorithmic breakthroughs that are difficult to predict.
"AI isn't replacing radiologists" good article
Expectation: rapid progress in image recognition AI will delete radiology jobs (e.g. as famously predicted by Geoff Hinton now almost a decade ago). Reality: radiology is doing great and is growing.
There are a lot of imo naive predictions out there on the imminent impact of AI on the job market. E.g. a ~year ago, I was asked by someone who should know better if I think there will be any software engineers still today. (Spoiler: I think we're going to make it). This is happening too broadly.
The post goes into detail on why it's not that simple, using the example of radiology:
- the benchmarks are nowhere near broad enough to reflect actual, real scenarios.
- the job is a lot more multifaceted than just image recognition.
- deployment realities: regulatory, insurance and liability, diffusion and institutional inertia.
- Jevons paradox: if radiologists are sped up via AI as a tool, a lot more demand shows up.
I will say that radiology was imo not among the best examples to pick on in 2016 - it's too multi-faceted, too high risk, too regulated. When looking for jobs that will change a lot due to AI on shorter time scales, I'd look in other places - jobs that look like repetition of one rote task, each task being relatively independent, closed (not requiring too much context), short (in time), forgiving (the cost of mistake is low), and of course automatable giving current (and digital) capability. Even then, I'd expect to see AI adopted as a tool at first, where jobs change and refactor (e.g. more monitoring or supervising than manual doing, etc). Maybe coming up, we'll find better and broader set of examples of how this is all playing out across the industry.
About 6 months ago, I was also asked to vote if we will have less or more software engineers in 5 years. Exercise left for the reader.
Full post (the whole The Works in Progress Newsletter is quite good):
https://t.co/ON3GwlI3mi
There is significant unmet demand for developers who understand AI. At the same time, because most universities have not yet adapted their curricula to the new reality of programming jobs being much more productive with AI tools, there is also an uptick in unemployment of recent CS graduates.
When I interview AI engineers — people skilled at building AI applications — I look for people who can:
- Use AI assistance to rapidly engineer software systems
- Use AI building blocks like prompting, RAG, evals, agentic workflows, and machine learning to build applications
- Prototype and iterate rapidly
Someone with these skills can get a massively greater amount done than someone who writes code the way we did in 2022, before the advent of Generative AI. I talk to large businesses every week that would love to hire hundreds or more people with these skills, as well as startups that have great ideas but not enough engineers to build them. As more businesses adopt AI, I expect this talent shortage only to grow! At the same time, recent CS graduates face an increased unemployment rate, though the underemployment rate — of graduates doing work that doesn’t require a degree — is still lower than for most other majors. This is why we hear simultaneously anecdotes of unemployed CS graduates and also of rising salaries for in-demand AI engineers.
When programming evolved from punchcards to keyboard and terminal, employers continued to hire punchcard programmers for a while. But eventually, all developers had to switch to the new way of coding. AI engineering is similarly creating a huge wave of change.
There is a stereotype of “AI Native” fresh college graduates who outperform experienced developers. There is some truth to this. Multiple times, I have hired, for full-stack software engineering, a new grad who really knows AI over an experienced developer who still works 2022-style. But the best developers I know aren’t recent graduates (no offense to the fresh grads!). They are experienced developers who have been on top of changes in AI. The most productive programmers today deeply understand computers, how to architect software, and how to make complex tradeoffs — and who additionally are familiar with cutting-edge AI tools.
Sure, some skills from 2022 are becoming obsolete. For example, a lot of coding syntax that we had to memorize back then is no longer important, since we no longer need to code by hand as much. But even if, say, 30% of CS knowledge is obsolete, the remaining 70% — complemented with modern AI knowledge — is what makes really productive developers. (Even after punch cards became obsolete, a fundamental understanding of programming was very helpful for typing code into a keyboard.)
Without understanding how computers work, you can’t just “vibe code” your way to greatness. Fundamentals are still important, and for those who additionally understand AI, job opportunities are numerous!
[Original text: https://t.co/nqzPC6eUpR ]
What if we could evolve AI models like organisms in nature, letting them compete, mate, and combine their strengths to produce ever-fitter offspring?
Excited to share our new work: “Competition and Attraction Improve Model Fusion” presented at GECCO’25🦎 where it was a runner-up for best paper!
Paper: https://t.co/ihfdriOPNw
Code: https://t.co/lOW2ghb5bj
Summary of Paper
At Sakana AI, we draw inspiration from nature’s evolutionary processes to build the foundation of future AI systems. Nature doesn’t create one single, monolithic organism; it fosters a diverse ecosystem of specialized individuals that compete, cooperate, and combine their traits to adapt and thrive. We believe AI development can follow a similar path.
What if instead of building one giant monolithic AI, we could evolve a whole ecosystem of specialized models that collaborate and combine their skills? Like a school of fish 🐟, where collective intelligence emerges from the group.
This new paper builds on our previous research on model merging, which follows such an evolutionary path. We started by using evolution to find the best “recipes” to merge existing models (our Nature Machine Intelligence paper: https://t.co/zqs7kGL4Gl). Then, we explored how to maintain diversity to acquire new skills in LLMs (our ICLR 2025 paper: https://t.co/00Buf1051A). Now, we're combining these ideas into a full evolutionary system.
A key limitation remained in earlier work: model merging required manually defining how models should be partitioned (e.g., by fixed layer or blocks) before they could be combined. What if we could let evolution figure that out too?
Our new paper proposes M2N2 (Model Merging of Natural Niches), a more fluid method, which overcomes this with three key, nature-inspired ideas:
1/ Evolving Merging Boundaries 🌿: Instead of merging models using pre-defined, static boundaries (e.g. fixed layers), M2N2 dynamically evolves the “split-points” for merging. This allows for a far more flexible and powerful exploration of parameter combinations, like swapping variable-length segments of DNA rather than entire chromosomes.
2/ Diversity through Competition 🐠: To ensure we have a rich pool of models to merge, M2N2 makes them compete for limited resources (i.e., data points in a training set). This forces models to specialize and find their own “niche,” creating a population of diverse, high-performing specialists that are perfect for merging.
3/ Attraction and Mate Selection 💏: Merging models can be computationally expensive. M2N2 introduces an “attraction” heuristic that intelligently pairs models for fusion based on their complementary strengths—choosing partners that perform well where the other is weak. This makes the evolutionary search much more efficient.
Does it work?
The results are fascinating: This is the first time model merging has been used to evolve models entirely from scratch, outperforming other evolutionary algorithms. In one experiment, starting with random networks, M2N2 evolved an MNIST classifier that achieves performance comparable to CMA-ES, but is far more computationally efficient.
Does it scale?
We also showed that M2N2 can scale to large, pre-trained models: We used M2N2 to merge a math specialist LLM with an agentic specialist LLM. M2N2 produced a merged model that excelled at both math and web shopping tasks, significantly outperforming other methods. The flexible split-point was crucial here.
Does it work on multimodal models?
When we applied M2N2 to text-to-image models, we merged several models by adapting them only for Japanese prompts. The resulting model not only improved on Japanese but also retained its strong English capabilities—a key advantage over fine-tuning, which can suffer from catastrophic forgetting.
This nature-inspired approach is central to Sakana AI’s mission to find new foundations for AI based on collective intelligence. Rather than scaling monolithic models, we envision a future where ecosystems of diverse, specialized models co-evolve, collaborate, and combine, leading to more adaptive, robust, and creative AI. 🐙
We hope this work sparks more interest in these under-explored ideas!
Published in ACM GECCO’25: Proceedings of the Genetic and Evolutionary Computation Conference. DOI: https://t.co/5eSwhvs5tQ
If you have been following the GPT-5 rollout, one thing you might be noticing is how much of an attachment some people have to specific AI models. It feels different and stronger than the kinds of attachment people have had to previous kinds of technology (and so suddenly deprecating old models that users depended on in their workflows was a mistake).
This is something we’ve been closely tracking for the past year or so but still hasn’t gotten much mainstream attention (other than when we released an update to GPT-4o that was too sycophantic).
(This is just my current thinking, and not yet an official OpenAI position.)
People have used technology including AI in self-destructive ways; if a user is in a mentally fragile state and prone to delusion, we do not want the AI to reinforce that. Most users can keep a clear line between reality and fiction or role-play, but a small percentage cannot. We value user freedom as a core principle, but we also feel responsible in how we introduce new technology with new risks.
Encouraging delusion in a user that is having trouble telling the difference between reality and fiction is an extreme case and it’s pretty clear what to do, but the concerns that worry me most are more subtle. There are going to be a lot of edge cases, and generally we plan to follow the principle of “treat adult users like adults”, which in some cases will include pushing back on users to ensure they are getting what they really want.
A lot of people effectively use ChatGPT as a sort of therapist or life coach, even if they wouldn’t describe it that way. This can be really good! A lot of people are getting value from it already today.
If people are getting good advice, leveling up toward their own goals, and their life satisfaction is increasing over years, we will be proud of making something genuinely helpful, even if they use and rely on ChatGPT a lot. If, on the other hand, users have a relationship with ChatGPT where they think they feel better after talking but they’re unknowingly nudged away from their longer term well-being (however they define it), that’s bad. It’s also bad, for example, if a user wants to use ChatGPT less and feels like they cannot.
I can imagine a future where a lot of people really trust ChatGPT’s advice for their most important decisions. Although that could be great, it makes me uneasy. But I expect that it is coming to some degree, and soon billions of people may be talking to an AI in this way. So we (we as in society, but also we as in OpenAI) have to figure out how to make it a big net positive.
There are several reasons I think we have a good shot at getting this right. We have much better tech to help us measure how we are doing than previous generations of technology had. For example, our product can talk to users to get a sense for how they are doing with their short- and long-term goals, we can explain sophisticated and nuanced issues to our models, and much more.
You spend more time on social media than you intend to, because time flows faster on these platforms, causing you to lose hours in what feels like minutes. This is no accident; it’s a result of a decades-long plot to steal your time.
My new essay.
https://t.co/7tnGcPxSK2
A lot of midwit chatter about AI and coding is very confused.
Effectively using LLM chat, coding assistants, and agents is a new category of skill.
These are the clear points about the state of the art models and tooling:
1. Coding assistants and agents are very good at many coding tasks, reasonably good at some, and bad at others
2. Knowing this difference and having the judgement to decide when and how to use coding assistants and agents is a massive strength.
3. LLM chat interfaces are also often insightful and respond with great (derived) responses about: architecting, engineering decisions, navigating team dynamics, etc. They work great if you are *already reasonably discerning* about these matters.
4. Agents, chat, and assistants are good at brainstorming and taking you from “0 to 1,” especially when working with problems and programming languages that are at the periphery of your core competence.
That's pretty much it.
There is going to be a divergence in the next few years between people who have the skill to augment their work well with assistants, agents, etc. and those who cannot.
UC Berkley has two free courses on LLM Agents for foundational and advanced levels. it also has some of the best lecturers from DeepMind, Meta, and top universities.
basically covers all you need to know about agents from the best resources out there.