Data, AI, and AI agents are all deeply transforming telco. In this interview with @TelecomTV, I detail how our CEO @Cheydema’s very well-received ‘Trust the future’ strategy will enable us to generate €600 million in value per year with data and AI by 2028. That’s building upon our very strong achievement of over €300 million in value from data and AI generated in 2025.
Our new strategy includes:
🔸 Sovereign and Trusted AI - For decades @Orange has served the needs of many major European enterprises as well as our governments and militaries. We offer a compete solution for these customers ranging from public cloud to truly sovereign both in our SecNumCloud Azure ‘Bleu’ data center with @CapGemini and @Microsoft, as well as in our own @orangebusiness data centers in France. Our OB CEO @AlietteML will have many more exciting announcements to share in the weeks to come at the Orange Business Partner Day.
🔸 AI in Customer Intimacy - A significant proportion of the €300 million of value in 2025 was thanks to a reinvention of our customer management systems and processes to provide far more relevant and optimized marketing to current and perspective customers We are also transforming our contact centers, leading to far higher efficiencies, and more importantly, improved customer satisfaction, retention, and upsell.
🔸 AI for Network - We are saving many tens of millions of euros thanks to AI in network CAPEX planning, but we’re also using AI to transform how our NOC and field services operate to provide a far more resilient and dynamic network. This work leverages our many long-standing partnerships with companies like @nokia and @Huawei but also includes many new innovative companies in this space. I am proud to be working with our brilliant Group CTO @lleboucher and my AI for Networks engineering VP Olivier Simon on this multi-year transformation.
🔸 AI for Innovative Growth - Led by our @orangeafrica CEO @YasserShaker_ , our #MaxIt super app is used by over 23 million people monthly. Now, thanks to our gifted AI research team we are enabling the over 80 million people in our African countries whose regional languages are not understood by any AI to be able to interact with MaxIt in languages such as Wolof and Bambara.
🔸 Excellence at Scale - Orange has created Live Intelligence, the world’s leading telco B2B solution for trusted AI. We have over 100,000 Orange employees using it, which includes the latest AI models from @AnthropicAI , @GoogleAI ,@MistralAI , and @OpenAI with full support for prompt engineering, RAG, and MCP, as well as true AI agents. The solution was so successful internally, we’re now selling it externally B2B through Orange Business.
A huge thank you to all of the Orange AI research and software engineering teams that have made all of this a reality and an equally heartfelt thank you to all of our brilliant partners around the world.
2026 is going be another great year for Data and AI at Orange. Trust the future!
Almost all AI model and agent progress is downstream from evals. Open weights post training for specific domains comes down to evals. Agent improvements in the applied AI layer is all about evals. Agentic enterprise deployments that actually can augment work is all about evals. It’s all evals.
This will become a core competency of any enterprise in the future. The companies that are able to best understand their own (and/or customers) workflows and how well agents participate in that work will be in the best position to actually drive real automation.
Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor.
It’s happening faster than we thought, and the implications deserve greater attention. https://t.co/OVVPJO7VQx
We should view the history of physics as a long-running program synthesis task. Kepler and Newton were searching the space of possible symbolic models to find the simplest one that would best satisfy available observations.
im excited about agent harnesses because i think are the first stable agent abstractions we can build on top (which is why we're investing so much in deepagents)
we always wanted to run llms in a loop and have them call tools (remember autoGPT? that's all that was)
but the models weren't good enough so we built chains and other architectures as a proxy
as the models got better, the "right" way to build the most agentic systems changed pretty dramatically which meant the frameworks (like langchain) had to change pretty dramatically to keep up
but now the models are good enough where "running the model in a loop calling tools" is actually starting to work
there is still a lot to build (async subagents, multiagent systems) but the foundation is more solid. rather than frameworks that are moving like quicksand, i genuinely think these agent harnesses are the first stable building block
we can build on top of them going forward, instead of rearchitecting them
I made a Claude Code skill that turns any arxiv paper into working code.
Every line traces back to the paper section it came from & any implementation detail the paper skips will be flagged, and not assumed.
open sourcing it -
https://t.co/sSio4JfpIo
New conceptual guide: 🔄 The agent improvement loop starts with a trace
Tracing is the foundational primitive for improving agents.
A trace gives you the full behavioral record of what an agent actually did. From there, teams can enrich traces with evals and human feedback, turn recurring failures into test cases, validate fixes before shipping, and repeat.
This guide breaks down the full improvement loop and why reliable agents are built through trace-centered iteration, not one-off debugging.
Read more → https://t.co/UnXqStm4OY
Thank you @apple for so many insanely great tools and experiences over the past 50 years.
Also thank you to so many Apple alumnae friends and @generalmagicmov Magician coworkers.
You taught me how to build things as well as what it was like to work with Steve.
🙏 @tfadell@andyhertzfeld@baldy@kevinlynch and especially the late @billatk
- Drafted a blog post
- Used an LLM to meticulously improve the argument over 4 hours.
- Wow, feeling great, it’s so convincing!
- Fun idea let’s ask it to argue the opposite.
- LLM demolishes the entire argument and convinces me that the opposite is in fact true.
- lol
The LLMs may elicit an opinion when asked but are extremely competent in arguing almost any direction. This is actually super useful as a tool for forming your own opinions, just make sure to ask different directions and be careful with the sycophancy.
If you are looking for an example of AI tools such as Claude Code driving higher velocity SDLC at company-wide scale, look no further than @AnthropicAI itself.
Bravo @bcherny and team.
73 product releases in 52 days. That's not a launch cadence — that's a different kind of company.
I tracked every Anthropic release from Feb 1 to Mar 23 by going through @bcherny, @trq212, @noahzweben, @felixrieseberg, @lydiahallie, @amorriscode, @feldman, @dickson_tsai, and @claudeai. Built a calendar with first-announcement attribution.
Look at the acceleration. February had bursts with gaps between them. March 9 onward is almost every single day — Code Review, Channels, Dispatch, Computer Use, back to back.
The individual features get coverage. The shipping velocity doesn't. It should.
Our great @orangebusiness CEO @AlietteML just launched the first trusted AI agents in Europe in our groundbreaking partnership with @LangChain.
LangChain and LangGraph agents will run on our #LiveIntelligence platform with on-premise LangSmith observation all running on GPUs hosted entirely in our latest sovereign Orange data center in France.
Excited for this state of the art solution for our customers requiring trusted AI solutions in France and beyond.
Thanks to the brilliant LangChain CEO @hwchase17 as well as the core Orange Trusted AI team including @dr_ujavaid , @miguelalva, @AnaBildea, and of course ‘The Father of Dinootoo’ Joachim Flechaire and our exec sponsors @bruno_zerbib and Aliette.
Bravo !
#TrustTheFuture #TrustOrange
This rocks. Great to see the Anthropic team moving so quickly to respond to new ideas like those that came out of the Claw ecosystem, well at the same time playing to their strengths. Bravo !
We're shipping a new feature in Claude Cowork as a research preview that I'm excited about: Dispatch!
One persistent conversation with Claude that runs on your computer. Message it from your phone. Come back to finished work.
To try it out, download Claude Desktop, then pair your phone.
This is a perfect solution for French and other European and African enterprises who have local regulatory restrictions that their data must live on infrastructure that’s more ‘trusted’ than the public cloud. Orange has been serving those customers with those needs for decades.
And now we have a state-of-the-art AI agentic and LLMaaS solution for them…
Thanks, Harrison! This partnership opens the door to trusted AI agents all across France and Europe for our @orangebusiness customers with sensitive AI workloads but who don’t want to settle for less than state-of-the-art agents. This is also a perfect fit with @orange’s new corporate #TrustTheFuture strategy. It’s great to be working with you and your brilliant team on this.
cc @AlietteML@Cheydema@bruno_zerbib
“Timing is very important. You need to pick hard problems to solve and be ambitious with them. But you've also got to pick the right time when the world and the context that you're in is the right kind of environment for those ideas to flourish.”
In his official Nobel Prize interview, Demis Hassabis discussed how his aspirations as a young gaming programmer were ahead of their time.
Watch our official interview: https://t.co/2ovRqsSAtc
Three days ago I left autoresearch tuning nanochat for ~2 days on depth=12 model. It found ~20 changes that improved the validation loss. I tested these changes yesterday and all of them were additive and transferred to larger (depth=24) models. Stacking up all of these changes, today I measured that the leaderboard's "Time to GPT-2" drops from 2.02 hours to 1.80 hours (~11% improvement), this will be the new leaderboard entry. So yes, these are real improvements and they make an actual difference. I am mildly surprised that my very first naive attempt already worked this well on top of what I thought was already a fairly manually well-tuned project.
This is a first for me because I am very used to doing the iterative optimization of neural network training manually. You come up with ideas, you implement them, you check if they work (better validation loss), you come up with new ideas based on that, you read some papers for inspiration, etc etc. This is the bread and butter of what I do daily for 2 decades. Seeing the agent do this entire workflow end-to-end and all by itself as it worked through approx. 700 changes autonomously is wild. It really looked at the sequence of results of experiments and used that to plan the next ones. It's not novel, ground-breaking "research" (yet), but all the adjustments are "real", I didn't find them manually previously, and they stack up and actually improved nanochat. Among the bigger things e.g.:
- It noticed an oversight that my parameterless QKnorm didn't have a scaler multiplier attached, so my attention was too diffuse. The agent found multipliers to sharpen it, pointing to future work.
- It found that the Value Embeddings really like regularization and I wasn't applying any (oops).
- It found that my banded attention was too conservative (i forgot to tune it).
- It found that AdamW betas were all messed up.
- It tuned the weight decay schedule.
- It tuned the network initialization.
This is on top of all the tuning I've already done over a good amount of time. The exact commit is here, from this "round 1" of autoresearch. I am going to kick off "round 2", and in parallel I am looking at how multiple agents can collaborate to unlock parallelism.
https://t.co/WAz8aIztKT
All LLM frontier labs will do this. It's the final boss battle. It's a lot more complex at scale of course - you don't just have a single train. py file to tune. But doing it is "just engineering" and it's going to work. You spin up a swarm of agents, you have them collaborate to tune smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the edges.
And more generally, *any* metric you care about that is reasonably efficient to evaluate (or that has more efficient proxy metrics such as training a smaller network) can be autoresearched by an agent swarm. It's worth thinking about whether your problem falls into this bucket too.