This is genuinely wild.
We gave our AI a single challenge: build a fully working Spotify clone, wired into a live music catalogue.
No scaffolding. No boilerplate. Just Alo AI, operating end to end.
Copilots write code. Operators ship products.
Watch it build👇
5/
The gap generalizes across 4 models, 3 families.
It also appears before and after instruction tuning, suggesting its origin is in pretraining.
Read our paper:
Perfect Detection, Failed Control: The Geometry of Knowing vs. Steering in Language Models
https://t.co/ZsFT3oKta4
We just discovered something that changes how we work on AI hallucination.
We found that LLMs know when they are about to lie...
...and they lie anyway 🧵
ArXiv ⬇️
4/
Why?
The direction that detects fake entities is not the direction that makes the model refuse.
They sit at cosine 0.12. About 83° apart.
Perfect detection.
Failed control.
Knowing is not steering.
3/
For hallucination, detection was almost too good.
In LLMs, fake entities are perfectly linearly separable from layer 5: AUC = 1.000.
But when we used that same direction to steer the model toward refusal, it failed.
2/
A central hope in mechanistic interpretability is:
if we can find where a behavior is represented inside a model, we should be able to control it.
Find the direction.
Intervene on the direction.
Change the output.
But what if detection and control are not the same thing?
1/
New paper from AloLab on why detection and control are geometrically different problems.
A language model can contain an almost perfect signal that something is fake…
and that signal still may not be the steering wheel.
Spent years at the ECB and Bloomberg watching smart people reformat the same report every quarter. That's the first job any real AI agent should take, not writing marketing copy.
Most enterprise AI pilots die at the demo. Not because the model is weak, because nothing it does persists after the meeting ends. An agent that can't ship something usable isn't done, it's a proof of concept.
Our research found detection and steering point in different directions inside a model. You can catch a behavior without being able to control it. Most safety tooling assumes otherwise.
Most enterprise AI pilots die at the demo. Not because the model is bad, because nobody built the app around it that a team can actually use every day.
Alo is an AI that powers entire companies.
Very honoured to be building one the most extraordinary products in tech with the sharpest, kindest, most brilliant people I know.
Thanks to our great investors for supporting us and pushing the tech frontier with us.
🚀 Congratulations to Alomana on a €4M funding round to accelerate Alo, its AI operating layer that goes beyond simple assistance, focused on enterprise execution.
The round is led by CDP Venture Capital SGR (Corporate Partners I – ServiceTech), with participation from Italia Venture I – Fondo Imprese Sud, Kairos Ventures (ESG One), Gresilent Holdings, Italian Angels for Growth (IAG) and Club degli Investitori.
Alomana deploys a universal intelligence layer working across data, documents and applications, elevating AI from helpful chatbot to repeatable execution.
Over the past year, Alo has been implemented across finance, manufacturing and pharma, supporting workflows like risk and security, financial controls, operational automation and advanced analytics, delivering real operational gains.
Alomana joined the Founders Factory and Mediobanca MBSpeedUp Accelerator in 2024, and it’s been exciting to watch the team build toward production-grade enterprise deployments. Congrats to Giuseppe Ettorre, Daniele Ligorio and team.
#VentureCapital hashtag#EnterpriseAI hashtag#AIAgent
Deep in the AI cave✨
We visited @MicrosoftAI with @DataSparkApp to discuss exciting GenAI advancements. The team is building something truly special.
Huge thanks to the man @cedricvidal for showing us around and sharing his knowledge.
#microsoftai#GenerativeAI#ai