1/ Can Large Language Models (LLMs) truly reason? Or are they just sophisticated pattern matchers? In our latest preprint, we explore this key question through a large-scale study of both open-source like Llama, Phi, Gemma, and Mistral and leading closed models, including the recent OpenAI GPT-4o and o1-series.
https://t.co/2tv8Pp9MSz
Work done with @i_mirzadeh, @KeivanAlizadeh2, Hooman Shahrokhi, Samy Bengio, @OncelTuzel.
#LLM #Reasoning #Mathematics #AGI #Research #Apple
This morning I had my first visceral “🤯” moment with AI for ~2 years
🧵on o1 and cryptic crosswords:
My test for new models is a set of cryptic crossword clues that aren’t online (my granny wrote them). Every model so far has been completely useless at them… but o1 gets them
Even as a developer your value is not in coding. Your value is in being a translator of business requirements and a communicator to those who come after you.
Don't generate your comments. Rather write the comments where you realise your greatest value, then generate the code 2/2
Today I heard someone extol the use of GenAI for commenting after coding.
Don't. Please.
If I want that, I can run the code through an LLM myself. An LLM won't tell me the things I most need to know - that business rule, that subtle corner, that edge case. 1/2
Interesting.
Curious to see what, if any, truly emergent behaviour was seen in these multi-agent systems. How much of that cooperation was prepared for in prompts or training?
Introducing Project Sid: the first simulations of 1000+ truly autonomous agents collaborating in a virtual world, w/ emergent economy, culture, religion, and government
Humans are the only species to land the moon, because we can cooperate at a vast scale
Can AI do the same?
On one hand, this blows my mind. On the other, burning half the planet training a diffusion model in order to recreate DOOM with it has a whiff of lighting a cigar with a $100 note 💵🙈
Brought back sweet memories playing this game with my friends at an impressionable age.
Wow, diffusion models (used in AI image generation) are also game engines - a type of world simulation.
By predicting the next frame of the classic shooter DOOM, you get a playable game at 20 fps without any underlying real game engine.
This video is from the diffusion model.
We're seeing some deserved pushback against the Generative AI hype. But. Given the right use case it is already an incredibly powerful productivity multiplier. And we're only getting started. https://t.co/GnDTNvIJMh
Just had a *mind blown* moment with Lindy AI
I accidentally enabled the AI agent I built on Friday that validates inbound leads, sends hyper-personalized emails, and negotiates on my behalf...
... and it WORKED
For context, I'm new to the platform and was tinkering with it yesterday evening, not realizing that I turned the triggers on.
It was pretty simple flow, and I wasn't even done:
1. The AI agent triggers when someone enters our sponsor form @TheRundownAI
2. It looks through the form input and validates the lead through looking at their LinkedIn and their goals
3. Uses GPT-4 to send a hyper personalized email to them based on their goals, our schedule, and our offer
4. It loops in a human operator when there's a higher level task (calls, in-person meetings, anything it doesn't know, etc.)
At first glance on Saturday morning I thought my inbox got hacked.
The AI agent was literally responding as me (Rowan) and got me a couple calls booked for next week.
I turned the agent off for now because I want to finish tinkering with the prompts before we do a full test flight.
But once I figure it all out, and if it actually works, it could take over ~50% of the tedious work my partnerships team has to do.
This means us mere humans can focus more on the unscalable stuff: client success, follow up calls, in-person relationships, and the stuff that actually matters like improving our newsletter.
It also works around the clock, which *should* increase satisfaction for the leads.
Once we do start a test flight, I'll be monitoring it extremely closely to make sure it only increases lead satisfaction before implementing it at scale.
More updates coming soon.
I learnt how naive we had been; I learnt BASIC, 6502 and Z80 assembler, built my own C compiler... geek stuff.
That memory came back just now; so I opened a prompt and typed the same question of all those years ago:
please teach me basic
And here it is; how far we've come! 2/2
I remember the first time I saw a personal computer in the shops when I was a kid. It was switched on, and anyone could type something.
Someone wrote, "PLEASE TEACH ME BASIC"
SYNTAX ERROR
> _
"It doesn't work like that," I thought. I started reading up on how it DID work. 1/2
A thousand times this.
There's a whole string of things that must've gone wrong for the Crowdstrike outage to happen, but this is the biggie. No risk mitigation; on the surface it seems they either are unaware of best practice or cannot be bothered with it.
When you are deploying an auto update that can brick a customer's computer, you roll it out very slowly over the course of several days, starting with a small random sample (the canary)
If it breaks the canary, you stop the roll out
This has been standard practice for decades
⚡️ Excited to share that I am starting an AI+Education company called Eureka Labs.
The announcement:
---
We are Eureka Labs and we are building a new kind of school that is AI native.
How can we approach an ideal experience for learning something new? For example, in the case of physics one could imagine working through very high quality course materials together with Feynman, who is there to guide you every step of the way. Unfortunately, subject matter experts who are deeply passionate, great at teaching, infinitely patient and fluent in all of the world's languages are also very scarce and cannot personally tutor all 8 billion of us on demand.
However, with recent progress in generative AI, this learning experience feels tractable. The teacher still designs the course materials, but they are supported, leveraged and scaled with an AI Teaching Assistant who is optimized to help guide the students through them. This Teacher + AI symbiosis could run an entire curriculum of courses on a common platform. If we are successful, it will be easy for anyone to learn anything, expanding education in both reach (a large number of people learning something) and extent (any one person learning a large amount of subjects, beyond what may be possible today unassisted).
Our first product will be the world's obviously best AI course, LLM101n. This is an undergraduate-level class that guides the student through training their own AI, very similar to a smaller version of the AI Teaching Assistant itself. The course materials will be available online, but we also plan to run both digital and physical cohorts of people going through it together.
Today, we are heads down building LLM101n, but we look forward to a future where AI is a key technology for increasing human potential. What would you like to learn?
---
@EurekaLabsAI is the culmination of my passion in both AI and education over ~2 decades. My interest in education took me from YouTube tutorials on Rubik's cubes to starting CS231n at Stanford, to my more recent Zero-to-Hero AI series. While my work in AI took me from academic research at Stanford to real-world products at Tesla and AGI research at OpenAI. All of my work combining the two so far has only been part-time, as side quests to my "real job", so I am quite excited to dive in and build something great, professionally and full-time.
It's still early days but I wanted to announce the company so that I can build publicly instead of keeping a secret that isn't. Outbound links with a bit more info in the reply!
Introducing Dream Machine - a next generation video model for creating high quality, realistic shots from text instructions and images using AI. It’s available to everyone today! Try for free here https://t.co/rBVWU50kTc #LumaDreamMachine
It eally wants to be proved out on present-day architectures and scales (the authors had budget constraints, apparently) but interesting early performance per compute results using image an transformer on ImageNet data and GPT-2 on OWT data. 2/2
Interesting paper just released (https://t.co/cmprZUa01Q) about using tensor networks to replace dense linear layers in visual models and LLMs. I've seen a few promising approaches such as a Tucker decomposition, but this paper use a generalisation of tensor train. 1/2
A lot of quantum computing news focuses purely on hardware, or new algorithms. There's been awesome progress in 2024!
However we need more work to scale quantum computing.
Here's some areas in the quantum technology industry that I’ve been thinking about:
1. We've been seeing that some algorithms work better on certain types of hardware. It seems better results are coming from trapped ion systems than superconducting for QML. However, superconducting is doing better than others. That means not only one hardware modality will win. VC funds wouldn't invest in multiple hardware companies before, which was generally hard for the industry, and now is clear is a losing game.
2. Many cryogenic companies are popping up and getting large funding rounds. Current technology hasn’t changed much over the years and is expensive—with each cryogenic fridge costing around a million dollars— it’s time to improve the systems, not only for more coherence but for accessibility in costs.
3. Beyond the fridges themselves, the need for 2 coaxial cables per qubits complicates the control systems as the number of qubits increases. And the problem of pure area here. Even building a 1000-qubit fridge is challenging. How do we do 10,000 superconducting qubits? To scale further, companies are exploring modular quantum systems. Others are looking into new control technologies to reduce the amount of cables.
4. Speaking of modularity, that means we may want a superconducting chip connected to a trapped ion system with photonics with classical PQC modules and quantum networks. That means we need to work on infrastructure. How do we build that? And how will these systems coordinate and talk to each other? There are hardware and software infrastructure challenges here to solve.
5. Companies focusing on specific industries are seeing results today. Universal simulators are valuable, but some companies have decided to go deep into a specific industry. Much like the logic of building application-specific chips, building application-specific quantum simulators means we can build larger simulators for the same computing power as universal chips - getting to some “quantum advantage” even today.
6. What does quantum advantage mean? It used to be about purely speeding up problems. Now, as I mentioned, we are seeing results in cost savings, if not speed, with simulators. Customers don’t care whether it’s a quantum computer, GPU, or a person doing math on paper behind the tech, as long as it works. Even if a quantum computer isn't faster, if it's cheaper, it's a win. So, we are reframing the benefits of quantum technology.
With advancements in hardware, algorithms, and infrastructure, the focus is on building a versatile, interconnected quantum ecosystem.
There’s a lot to do!
Both slightly insane and absolutely wonderful: run a full GPT style pipeline (nanoGPT) in a spreadsheet as an educational tool. https://t.co/WmGGoiXVUl