RIP fine-tuning ☠️
This new Stanford paper just killed it.
It’s called 'Agentic Context Engineering (ACE)' and it proves you can make models smarter without touching a single weight.
Instead of retraining, ACE evolves the context itself.
The model writes, reflects, and edits its own prompt over and over until it becomes a self-improving system.
Think of it like the model keeping a growing notebook of what works.
Each failure becomes a strategy. Each success becomes a rule.
The results are absurd:
+10.6% better than GPT-4–powered agents on AppWorld.
+8.6% on finance reasoning.
86.9% lower cost and latency.
No labels. Just feedback.
Everyone’s been obsessed with “short, clean” prompts.
ACE flips that. It builds long, detailed evolving playbooks that never forget. And it works because LLMs don’t want simplicity, they want *context density.
If this scales, the next generation of AI won’t be “fine-tuned.”
It’ll be self-tuned.
We’re entering the era of living prompts.
ChatGPT o3 found a path through this 200x200 maze for me in one try.
I had to overlay the solution over the original in Photoshop and flip between layers while zoomed in to check the solution never crosses a wall and none of the walls are changed. It's perfect.
Non content de rouler sur la piste cyclable avec son scooter, il faut qu'il vienne me bousculer puis m'agresser. Quel enfer...
Le sujet complet : https://t.co/EFsP3eeYjb
Microsoft just dropped VASA-1.
This AI can make single image sing and talk from audio reference expressively. Similar to EMO from Alibaba
10 wild examples:
1. Mona Lisa rapping Paparazzi
🚨 Our new paper: we know that GPT-4 generates better ideas than most people, but the ideas are kind of similar & variance matters
But it turns out that better prompting can generate pools of good ideas that are almost as diverse as from a group of humans https://t.co/LkGsU0VC7S
Having just taught initial AI stuff to 250+ undergrads & grad students in multiple classes today:
-AI use approached 100%. Many used it as a tutor. The vast majority used AI on assignments at least once
-Knowledge about AI was mostly based on rumors
-Prompting knowledge was low
The ML ecosystem in France is on fire🔥 It has amazing talent and resources. Here are 10 facts you might not know:
1. There are great research labs - from @MistralAI and @kyutai_labs to large ones from @AIatMeta and @GoogleDeepMind. The Llama 2 and CodeLlama authors are based in France!
2. Did you know sklearn is maintained by @Inria (a top national research institution).
3. Companies such as @OVHcloud and @Scaleway are European leads in computing and hosting.
4. There is the Jean Zay supercomputer with 28 petaflops, where @BigscienceW Bloom was trained.
5. @huggingface has its largest office there🤗.
6. There is @joinstationf, the world's largest startup campus with 1k+ startups, and 42, a very interesting CS school.
7. It has top CS universities. Maybe they don't have international renown or prestige and they don't invest as much in marketing, but they are really top, and their alumni are really, really strong.
8. A thriving startup ecosystem, @MithrilSecurity@photoroom_app@giskard_ai ChainLit @zama_fhe and many others.
9. Unlike SF, it's not in a tech bubble. E.g., the fashion industry is very strong. There are many art collectives. A good example of the power of this is Obvious Art (https://t.co/8zBT5Thzxl), a collective of researchers and artists working with ML.
10. It's very well located in Europe; quick train ride away from other tech hubs such as London, Barcelona, or Zurich, as well as strong universities (EPFL, ETH, UK unis, etc).
The multimodal and reasoning capabilities of Gemini are quite strong. The benchmark results, which I’ll discuss in a moment are nice, but I’m most excited by demonstrations of what it can do.
Consider the image below. A teacher has drawn a physics problem of a skier going down a slope, and a student has worked through a solution to computing the speed of the skier at the bottom of the slope. Using Gemini’s multimodal reasoning capabilities, the model is able to read the messy handwriting, correctly understand the problem formulation, convert both the problem and solution to mathematical typesetting, identify the specific step of reasoning where the student went wrong in solving the problem, and then give a worked through correct solution to the problem. The possibilities in education alone are exciting, and these multimodal and reasoning capabilities of Gemini models could have dramatic applications across many fields.
L'Europe trouve un accord sur le Data Act
🗨️ "Il y a une grande levée de boucliers de la part des entreprises, car ils craignent que le secret des affaires ne vole en éclat"
🎙️ @ClaudiaECohen
How can generative AI effect an organizational culture? Listen to the latest @HarvardBiz podcast, featuring Nitin Mittal on adopting generative AI effectively and ethically. https://t.co/4K5pCb4PjI
AI-generated QR codes using ControlNet are insane.
This is going to be increasingly common in ads in the near future.
These examples blew my mind (try scanning them):
1. Ancient Village
Ford CEO Jim Farley explains why legacy auto software sucks.
To survive, legacy automakers will need to learn how to be software companies for the first time.
So I watched the entire 3-hour long Senate hearing on AI today. I want to make sure I cover as much relevant info as possible in Friday's weekly AI news video. I was really impressed with Sam Altman's responses.
Here's a super-edit I made (I recommend 2x speed)