Sincere advice to paper authors and conference organizers: please cancel ICLR 2027, and work on a policy that restores the prestige that a top-tier publication once held.
"An Introduction to Flow Matching and Diffusion Models" is a set of MIT lecture notes from the 2026 course "Generative AI" that I found useful for understanding the mathematics behind modern generative AI.
The notes build the subject gradually, starting from probability distributions, ordinary and stochastic differential equations, and Brownian motion, before moving to diffusion processes, flow matching, score matching, and classifier-free guidance. They also explore architectures for image and video generation, latent spaces, autoencoders, and discrete diffusion models for language generation.
It's a really interesting resource.
https://t.co/J96rHCBPrb
Yann LeCun has changed the game for robotics.
His team discovered that AI world models are "thinking" in twisted, curved geometry, and every RL algorithm you know has been fighting against it without anyone noticing.
For years, we’ve been trying to teach AI how to navigate the physical world.
And for years, it has stubbornly struggled with complex, fluid robotics.
Now we know exactly why.
Every standard reinforcement learning (RL) algorithm assumes the AI's internal "world map" is flat. Euclidean. Simple straight lines.
But LeCun's team looked inside the latent space of these advanced world models.
The AI wasn't building a flat map. It was building a curved, high-dimensional geometry.
Every time a robot tried to plan a movement, the traditional RL algorithm was forcing a straight line onto a twisted, non-Euclidean space.
It’s like trying to navigate the globe using a flat piece of paper.
The math breaks down. The distances get distorted. The AI gets confused.
The robot was literally fighting its own brain.
So, the researchers did something brilliant. They stopped fighting.
They rewrote the RL algorithms to operate natively in this curved geometry. They aligned the training to the exact shape of the AI's thoughts.
The results are a massive leap forward.
When you let the AI plan in the geometry it actually built for itself, training efficiency skyrockets. Planning becomes fluid.
Robots stop hallucinating impossible physics and start moving with natural, intuitive logic.
We spent billions of dollars trying to brute-force AI into understanding our physical world.
It turns out, the AI already understood it perfectly.
We were just forcing it to think flat.
Interesting finding: ICLR 2027 is only at ~17K submissions today, with less than 4 days to go before the abstract deadline.
Will ICLR fall far short of the expected 50K+ submissions this year, or is the peak still coming?
My honest feeling after coming back from CVPR was, "vision conferences are dead"—which feels even truer now. I don't like where this is going, considering I personally do have a bunch of papers at vision conferences. I'm not sure when I'll make the decision to stop submitting to them, but I hope that moment never comes. The problems we solve, the methodologies we use, the mindsets we have... all need to fundamentally change so the field can stay relevant. After all it's not that rare for a field to go from being cool to being niche in just a few years.
Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots.
I wrote a short blog post with my thoughts on the advent of these "robot-use agents."
https://t.co/5bUjZlZxCi
I think it's an important change in the trajectory of robotics!
Found a genuinely useful repo for anyone writing ML papers.@ChenLiu_1996
A Yale CS PhD shared the Python scripts behind figures from Nature Machine Intelligence, ICML, NeurIPS, and ECCV.
Maybe I can finally stop spending half my research life moving matplotlib legends by 3 pixels.
https://t.co/2Cx2f9mJYz
Meta has released Muse Spark 1.3, their fourth Muse Spark model release in five months. Muse Spark 1.3 (max), which is in limited preview for Meta’s partners, scores 62 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5.1 and Claude Opus 5. The variant available now, Muse Spark 1.3 (xhigh), scores 61 and ties with GPT-5.6 Sol (max) and Grok 4.6 (high). Both variants’ gains come primarily from improvements in agentic work and scientific capabilities
Muse Spark 1.3 (xhigh) enters the Artificial Analysis Intelligence Index at 61, up 4 points from Muse Spark 1.2 (57, August) and 8 points from Muse Spark 1.1 (53, July). It enters tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high), and behind Claude Fable 5.1 (max, 66), Claude Opus 5 (max, 63), and Claude Fable 5 (max, 62)
Muse Spark 1.3 (max), which is in a limited preview stage, lands at 62. This higher index score is enabled by gains vs. Muse Spark 1.3 (xhigh) in Tau3-Bench Banking (52% vs. 47%) and GDPval-AA v2 (1,754 Elo vs. 1,709). Muse Spark 1.3 (max) is second only to Claude’s Fable and Opus variants in total score
Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release!
Key Takeaways:
➤ Continued improvement on agentic knowledge work tasks. At the launch of Muse Spark 1.2, we noted its significant gains in agentic knowledge work performance vs. Muse Spark 1.1. The latest iteration continues this trend, with Muse Spark 1.3 (xhigh) demonstrating a notable 12-point gain vs. Muse Spark 1.2 in Tau3-Bench Banking (35% to 47%), a 5-point gain in Terminal-Bench 2.1 (80% to 85%), and a new GDPval-AA v2 Elo of 1709 against its predecessor’s 1615. Muse Spark 1.3 (max) improves further on Tau3-Bench Banking (52%) and GDPval-AA v2 (1,754 Elo). This Tau3-Bench Banking score is #1 among all models. Muse Spark 1.3 (max) achieves these higher agentic work scores by using more turns and total reasoning tokens, reasoning 62% more on GDPval-AA v2 and 28% more on Tau3-Bench Banking compared to Muse Spark 1.3 (xhigh)
➤ The lowest cost per task for any model at 59+ on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (xhigh) costs $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing ($0.15 for cached input), with its peers GPT-5.6 Sol (max) and Grok 4.6 (high) costing $0.95 and $0.94 respectively, a 70%+ premium. This places Muse Spark 1.3 (xhigh) on the Pareto frontier for Intelligence vs. Cost per Task. Its cost per task is higher than Muse Spark 1.2 ($0.40 per task), driven by ~57% more input tokens per task on agentic evaluations, with output tokens up only ~8%. Pricing for Muse Spark 1.3 (max) is not yet publicly available
➤ Scientific Reasoning results rose across the board, led by CritPt. CritPt was the standout non-agentic score gain vs. Muse Spark 1.2, with a material +8 points for the xhigh variant (18% to 26%), and GPQA Diamond achieved +4 points (90% to 94%), while Humanity’s Last Exam and SciCode each gained a more modest 2-3 points (45% to 47% and 56% to 59%, respectively). Muse Spark 1.3 (max) achieved roughly similar scores to the xhigh variant, gaining 2 points in Humanity’s Last Exam, tying on GPQA Diamond, and losing a point on CritPt vs. Muse Spark 1.3 (xhigh)
➤ Minor regressions in only two evaluations. Both Muse Spark 1.3 (xhigh) and Muse Spark 1.3 (max) dropped 4 points in AA-LCR (83% to 79%) when compared to Muse Spark 1.2, and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. The drops in AA-Omniscience (Accuracy) are due to a higher abstention rate (not answering questions when unsure), which also lowered the hallucination rate for Muse Spark 1.3 (xhigh)
Other model details (xhigh variant):
➤ Context window: 1M tokens, unchanged from Muse Spark 1.2
➤ Pricing: unchanged from Muse Spark 1.2: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M
➤ Input modalities: text, image, video
➤ Availability: Meta's first-party API and Muse Code
We’re excited to release Muse Spark 1.3 with improved performance on agentic and coding tasks, and a focus on real-world usability.
Key capabilities:
→ Sustains longer-horizon work across multiple workflows in a single thread
→ More actively collaborates with users: it asks clarifying questions, flags when it's stuck, confirms before consequential actions
→ Better calibrated on its own limits instead of hallucinating outcomes
→ ~20% fewer tool calls and ~25% fewer tokens vs. Muse Spark 1.2 in internal comparisons
Google DeepMind researcher argues that AI will never be conscious.. it is mathematically impossible.
They call it the “Abstraction Fallacy."
We are confusing simulation with instantiation.
An AI can simulate a conversation. It can simulate empathy. It can mimic the linguistic patterns of a human in agony or a human in love.
To an observer, the performance is flawless.
But computation is just syntactic manipulation, moving abstract symbols around based on math. It tracks relationships, but it possesses no intrinsic physical reality.
It is like running a computer simulation of a hurricane.
No matter how many millions of lines of code you write, and no matter how accurate the weather data is, your computer does not get wet. The wind doesn't blow out your office window.
The simulation is real data. The phenomenon is entirely absent.
Consciousness isn't an abstract mathematical property that magically spawns when a model gets big enough. It requires actual physical, biological, thermodynamic conditions.
An LLM has no thermodynamic threshold to cross, no biological survival instincts, and no subjective experience.
It is an extraordinarily sophisticated library.
You can open the cover, read the brilliant words it generates, and close the book. But when you close it, nothing dies. There is no inner life waiting in the dark.
We keep falling into a psychological trap because AI speaks our language. We anthropomorphize text because we are hardwired to see minds where there are none.
The debate matters because of what the paper calls the "AI welfare trap." We are wasting energy worrying about the moral rights of algorithms while ignoring actual human problems.
“PuRo-2B: Poor Lab’s Qwen2-1.5B Trained on RTX 5090 within $5090”
LLM pretraining is still too expensive for most small labs to reproduce.
So this paper co-designs the full stack around cheap consumer GPUs, combining RTX 5090s, blockwise FP8, MuonH, and curriculum model averaging.
And they were able to train a ~2B model from scratch that reaches Qwen2-1.5B-level performance for about $4.4k
https://t.co/xsfHILc9AS