We are announcing a long-term strategic partnership with NVIDIA. NVIDIA is making a substantial investment in SSI that will let us 10x our compute in the next 12 months. We reached the point where our research is worth scaling and with this partnership we will be able to. We are honored by NVIDIA’s conviction.
As a professor, I am often asked: Should students use AI coding agents for their research?
I have also been trying to understand this question myself, but until now, I had mostly learned about students’ use of AI indirectly, through conversations with them or by seeing the final results. It is rare to have a long, uninterrupted period in which I can closely observe how they actually work, what they delegate to AI, and how their decisions evolve from day to day.
The past three weeks at the Telluride Neuromorphic Workshop offered exactly that opportunity. Students propose their own innovative science and engineering projects, form teams, and work toward a demoable final presentation over three weeks. Here are a few things I noticed:
1. Productivity increased dramatically with AI. This was evident across both engineering and scientific projects. The ability to prototype an idea quickly has never felt so powerful or direct until I sat down with the students and observed their coding pipelines. For instance, one integrated-circuit group developed an asynchronous microcontroller with custom instructions, connected it to event-based sensors, deployed it on an FPGA, and produced a runnable demo with C code and a polished, configurable web interface. Completing all of this in three weeks would have been almost unthinkable a few years ago. Contrary to what people sometimes presume, many amazing ideas still emerge from discussion, while the Claude tokens are burning in the background.
2. Students are under peer pressure to use AI to keep up with the rapid pace of development. I noticed that students were more likely to use coding agents when other members of their teams were already using AI to finish their parts. Once some students begin moving much faster, the pressure to produce comparable results becomes very real.
3. The time budget shifts from debugging to planning. I noticed that students spent significantly more time planning their next steps while leaving much of the tedious plumbing work to coding agents. I saw far fewer of the typical frustrations caused by a “stupid bug,” which allowed students to focus more directly on system design and the actual problem they were trying to solve.
4. AI can be more focused on engineering solutions than scientific exploration. Students sometimes found that coding agents concentrated narrowly on solving the problem stated in the prompt, while overlooking instructions to branch out or explore less conventional possibilities. Instead of substantially modifying an established solution or pursuing a risky new direction, coding agents often preferred a safer implementation that was more likely to work. This is useful for engineering, but it can also discourage the kind of uncertain exploration that leads to genuinely novel research.
5. The biggest concern is understanding. Students increasingly delegate entire tasks to AI, which sometimes leads them to present systems or results that they do not fully understand. This is clearly harmful because it removes both the intellectual ownership and the educational value of academic research. A working demo is not enough if the student cannot explain why it works, where it may fail, and which technical decisions shaped the final result.
My conclusion is that AI in research is no longer optional. We need to embrace this change, but we must also separate two objectives that are often conflated: education and research project execution. Students benefit most from AI when they already understand the fundamentals, and when they use it to remove repetitive engineering work so that they can focus more deeply on the core scientific questions.
Students should still own the questions, the decisions, and the interpretation, even if AI writes much of the plumbing. My lab may soon become more AI-enabled. The challenge will be making sure the agents accelerate the research without becoming the researchers
#AI #Research #Engineering #Education
Congrats, @Kimi_Moonshot!
『 Kimi’s Four Commandments』have circulated in the Chinese AI community for months. Many people add their own fifth for comic relief, but the original four were the core. Here's an English translation --- since apparently none can be taken for granted.
The brilliant @nitishsr and I were both with @geoffreyhinton when Russ first joined UofT as faculty and we absolutely jumped at the opportunity to add Russ as co-advisor for our PhD.
I've always admired Russ and his work even in undergrad, reading some of his papers 5 times over! Deep Boltzmann Machines were the representative work in GenAI in 2009.
So when I graduated in 2015, we followed the blueprint laid out by our friends of DNNResearch (Geoff, Ilya, and Alex) and that led to our very own Perceptual Machines Inc. (w/ @nitishsr, @rsalakhu)
Words can not capture all the memories, excitements, and up and downs of those years, truly blessed to have studied and worked with both of you @nitishsr, @rsalakhu !!
Big shoutout also goes to the amazing @Ahmad_Al_Dahle for giving us the opportunity and believing in us very early on.
Charlie @tang_1c was one of my very first PhD students at Toronto. I still remember those days when together with Nitish @nitishsr, we were pitching our research to Apple execs. Great memories and fun times!
Russ's first PhD and first startup co-founder here..
Kids these days would forget that his 2006 Science paper is what kicked off the *entire* Deep Learning revolution.
Still vividly rem that one time him settling for my sunnyvale apt air 🛏️ as we prepared long into the night for our deck to the Apple execs - Russ worked incredibly hard in pursuit of greatness in both academia and industry.
Lastly, Russ was and still is the most approachable and gregarious super-star AI/ML prof of them all, no matter if you were a colleague, investor, mag 7 CEO or simply a random ML conf attendee. #character
I’ve been asked several times whether Zhilin Yang, the founder of @Kimi_Moonshot was my PhD student. The answer is yes and he is absolutely brilliant.
But I’ve been incredibly fortunate to work with so many outstanding PhD students over the years. So I thought I’d brag a little about them and their career paths (of course there are also many MSc and undergraduate students, sorry if I missed anyone):
Founders / Founding Team Members
Devendra Chaplot @dchaplot PhD, Founding Member Thinking Machines / Mistral
Zhilin Yang PhD, Founder & CEO, Moonshot AI
Jimmy Ba @jimmybajimmyba MSc/PhD, Co-founder xAI
Hubert Tsai PhD, Co-founder Spuree, Apple
Nitish Srivastava @nitishsr PhD, Co-founder Perceptual Machines; Co-founder Vayu Robotics
Charlie Tang PhD, Co-founder Perceptual Machines, DE Shaw.
Professors
Paul Liang @pliang279 PhD, MIT
Ben Eysenbach @ben_eysenbach PhD, Princeton University
Ruosong Wang @RuosongW PhD, Peking University
Bhuwan Dhingra @bhuwandhingra PhD, Duke University
Roger Grosse @RogerGrosse Postdoc, University of Toronto
Alexander Schwing Postdoc, UIUC
Research Scientists
Shuyan Zhou @shuyanzh36, Postdoc, Meta Superintelligence Lab
Tiffany Min @SoYeonTiffMin PhD, Microsoft AI
Murtaza Dalal @mihdalal PhD, Tesla AI
Minji Yoon @MinjiYoon90 , PhD, Microsoft AI
Shrimai Prabhumoye PhD, NVIDIA AI, Mistral
Haitian Sun @sun_haitian PhD, Google DeepMind
Emilio Parisotto PhD, Google DeepMind
Lisa Lee PhD @rl_agent, Google DeepMind
Manzil Zaheer @ManzilZaheer PhD, Google DeepMind
Jamie Kiros PhD, Google Brain, OpenAI
Yuri Burda PhD, OpenAI, Anthropic
Cody Severinski PhD, Amazon
I became Senior VP at a multi-million dollar company at age 26. My salary was $600k. This was in 2018.
How did I do it?
It wasn’t hustle culture. No 5:30am wakeups, cold showers, or productivity hacks.
What got me there was a relentless focus on impact. Every project I touched, every deck I built, every presentation I gave MOVED THE NEEDLE.
Always, I asked myself: what is the single most valuable contribution I can make to the company right now? And I did that. If people disagreed, I convinced them otherwise.
I kept this up for three years before the CEO (my dad) finally recognized my results and promoted me to SVP.
There are no gimmicks. There are no shortcuts.
To get ahead, create value.
In an era of AI overtaking humans in more and more domains, computer/AI ranking of sports teams is still quite different from humans, due to a low sample size.
#AI#SportsPicks
An OpenAI executive expects the “entire economy” to become an “Reinforcement Learning Machine.” This implies that AI might train on recordings of how professionals in all fields handle day-to-day work on their devices. Details on this new era of AI training:
• AI developers are training models on carefully curated examples of answers to difficult questions.
• Data labeling firms are hiring experienced professionals in niche fields to complete real-world tasks using specific applications that the AI can watch.
Read more: https://t.co/XrlVGLzqFL
Notice that her face showed disgust even *before* she tasted the soup!
AGI tests along these lines could be created both for video gen AIs and video understanding AIs.
for sure, network structure is often very important for abstractions or learning invariant features/representation. for example, when trying to model faces under different lighting conditions, it's more natural to have multiplicative layers (aka tensor factorization) because the way the light interacts with 3d objects IRL. I consider this a different network structure than your classic MLP layers.
https://t.co/PwoPNUEV3o
Very nice written article!
It's interesting how DAgger-esque methods have been deployed in practical self-driving: e.g.
- loyal FSD fans sending tesla disengagement corner cases
- road testing collects safety driver takeovers, sends it for human labeling, then retrain the models
In RL, PPO's advantage function can be seen as the teacher/expert (\pi*), but not as good because it is a bootstrapped estimate of the future expected discounted reward for a state-action.
Obviously the two can be combined to potentially reduce sample complexity or avoid unsafe explorations.
New research on fixing speech native foundation models: EchoX. (what i'm reading)
ELI5 ➡️ Speech native foundation models like GPT4o (advanced voice mode) is nicer than modular ASR->LLM->TTS systems but are not as good on knowledge and reasoning due to the *acoustic-semantic gap* !
The problem is that standard speech2speech training will try to match the acoustic characteristics of the data, such as prosody, timbre, pitch, etc. however, these are not the best objectives for generating semantically meaningful tokens.
The paper's big idea is that instead of using raw audio tokens as targets during training, they will use a TTS system to generate audio tokens of specific style as targets instead. Basically this process collapses certain variations in the original raw audio (due to prosody/timbre)
They show good results using only *6000* hours of speech data, compared millions of hours used in other systems.
👏Zhang et al. from the CUHK.
Customer interactions with Taco Bell’s AI has gone viral recently, for the *hilarious* reasons, leading to the pulling back of deployments
💡Why Taco Bell's voice AI failed: (and how to fix it)
❌Failure modes:
1. Sub-human language understanding
2. Adversarial attacks (mostly for fun by pranksters)
3. Sounds too robotic
4. High latency
1. Sub-human level language understanding:
- In one instance the AI asks if the customer wanted a drink even after the customer requested a mountain dew. This comes from the fact that either ASR failed or the previous request wasn't registered (e.g. not matched exactly to a menu item).
- Another time a customer says "no sauce" and the system continues with "we have sauce like ranch..."
2. Adversarial attacks:
They not security related, but rather pranks where the human asks for implausible items or "1 million beers" and the AI becomes confused.
3. Sounding too robotic:
AI always starts with "Hi, welcome to taco bells, what can i get started with you today?", where as human staff might go with something like "what can i get you?" or "hey, there, welcome!", which are more personal, and shorter, not wasting 1-2 seconds of time.
4. High latency:
Almost all deployed systems today are a modular or cascade system, where audio input is first processed by Automatic Speech Recognition (ASR), then a LLM with customized prompts handles the NLP reasoning and logic, followed by a text-to-speech (TTS) system for speech.
This is akin to a human plugging their ears until the other human finishes speaking and then start to process the utterance. High latency is guaranteed.
High latency causes customers to become impatient and wanting to talk to a human.
* The main issue is that these AI systems are just the classical Finite-State Machines *dressed up* nicely with the latest voice/speech technology
✅how to fix it:
1. Let LLMs handle more of the reasoning, go away from finite-state machines and scaffolds, which are lipstick on a pig (the pig here is the classic control flow or flowchart programming paradigm)
2. Adversarial attacks:
Better reasoning and intent estimation (of the human user) will solve this. i.e. the AI needs to be trained to include the possibility of customers whose intention is to make outlandish requests.
3. Train on real human-human data, not just prompt-to-TTS approach done by AI engineers.
4. Move away from modular voice AI, and towards real-time streaming ASR/TTS, or end2end duplex models. These models do not wait for the other's turn to be over before starting to process, this is how humans behave in a conversation.
video credits: @FoodlesCA YT