Get more out of GPT-6 Astra by revisiting your skills, AGENTS.md, and task prompts.
Make skill triggers specific, load guidance when it's relevant, and define what done looks like.
https://t.co/UGF0AC8Z5Y
Google DeepMind argues RAG is broken.
They published a paper that proved vectors databases are the dead end.
For the last three years, the default engineering response to any AI memory or data problem has been identical: "Just build a RAG pipeline."
Chunk the data, push it into a vector database, and let embeddings handle the rest.
Every company scaling enterprise AI assumes that if an embedding model fails, it's just a matter of time. Better training data, larger models, more parameters—throw compute at it, and the search gets smarter.
This paper proves that assumption is completely false.
They mathematically demonstrated that single-vector embeddings have a hard, uncrossable limit.
Here is the core flaw:
An embedding compresses an entire document or a complex query down into a single fixed-length vector of numbers.
When you run a search, the model takes the dot product of those vectors to measure similarity.
The math reveals a brutal constraint. The number of distinct document combinations a model can possibly retrieve for different queries is strictly bounded by the dimension of its embedding space.
It is a hard mathematical ceiling dictated by geometry and communication complexity.
No amount of data scaling can fix it. No amount of fine-tuning will punch through it.
Even if you give an embedding model infinite, unconstrained training freedom on the test set, it still hits the wall.
DeepMind built a stress-test dataset called LIMIT to prove it.
They threw state-of-the-art embedding models at it, models with thousands of dimensions.
The models completely failed. Even on simple, structured queries, the single-vector bottleneck forced the system to drop critical context and hallucinate irrelevant results.
Why? Because a single vector cannot capture complex, multi-faceted relationships between documents.
When you ask an AI to reason, follow complex instructions, or handle nuanced cross-document dependencies, the vector space simply runs out of room.
It collapses.
This changes everything for software architecture.
If your AI agent's memory relies on standard single-vector retrieval, it is structurally blind to complex logic. It is missing pieces of your data right now, and no prompt tweak can save it.
If we want AI that actually understands enterprise knowledge, we have to throw out the single vector.
And invent something entirely new.
Hello people of Sol! I've reset usage limits for all ChatGPT Work and Codex users. Together with that, a quick update on GPT-5.6 Sol usage limits.
Over the past few weeks, many of you have told us that Sol was using your Codex limits faster than expected. To be clear, we have not reduced usage on any subscription plans.
We’ve been digging into what was happening and have landed several improvements. As a result, we expect your usage to last around 18% longer during typical use of Sol. Some of you should already see significantly larger improvements from today. Tomorrow, we’ll also restore the five-hour limit that we temporarily paused while investigating.
Here’s what we found:
- GPT-5.6 Sol is much more willing to work for longer, make additional tool calls, and coordinate complex workflows across tools and subagents. That makes it better at solving hard problems, but some tasks were using far more than we intended.
- Sol also works harder at the same reasoning effort than previous models. High on Sol can use more tokens than High did on GPT-5.5.
- Programmatic tool calling, also referred to as code mode, gives Sol much more flexibility to run tool calls in parallel or continue working while waiting. But it also led to more responses per turn, more cached input tokens, and higher usage than expected.
- This was particularly noticeable when Sol was waiting for tool calls to finish or running many web searches. We’ve improved how we handle both cases and are continuing to make code mode more efficient.
- The impact was also very uneven. The median user actually found Sol quite token efficient, while some power users working on harder tasks saw their usage drain much faster. We were very focused on average and median usage before launch and missed some cases where the long tail could use significantly more usage.
Sol is a significant step forward in what Codex can do, but capability and efficiency do not always improve at the same pace, and some issues only become clear once people are using the model at real-world scale. We should have recognized this sooner and been more upfront about it.
You keep pushing the frontier and we’ll keep improving efficiency and sharing updates as we go.
My view of: Fable 5 vs GPT-5.6-Sol. They are not easy models to compare, these are my vibes - take them as you will.
My overall feel is that Fable is a 'wise owl' who is very thoughtful and very well spoken, GPT-5.6-Sol is like a rottweiler who will grab the problem by the throat and not let go until it is done.
In other words, Fable, is a fundamentally smarter model - even at low reasoning it can be very insightful and writes in a clear compelling way. GPT-5.6-Sol on the other hand is extremely diligent, I can give it a list of 8 things to do and you will be sure that they will be done.
Fable feels more arrogant to me, I was both to get it to build a new benchmark for me - 5.6 worked between 6 hours and 2 days (I tried several times) and it came up with very thoroughly tested, working benchmark. Fable came back within 40 minutes (twice) and the benchmark sounded smart, but was ultimately was 'vibe' based slop and since it was Fable's vibes that was doing the judging, it decided that it was good to go (it kept giving Fable 100% score btw).
Some thoughts by category:
UI & App building: Fable will still craft a better UI from scratch, the flow of the app would probably be a bit nicer. But I find that Fable often misses quite key things, which GPT-5.6-Sol doesn't. GPT's Frontend skills are big jump vs previous GPT models, but still not as great overall.
Writing: Fable is better hands down, Sol feels quite difficult to align to what I want to say or explain things to me simply. Though I think the 'Pro' model writes clearer.
Robustness & Reliability: This is where I think GPT-5.6-Sol wins for me hands down. Fable seems to do things of high quality, but I can never relax with it, it always misses something. With 5.6 this just almost never happens.
Other things where I liked GPT-5.6-Sol, but can't compare to Fable directly.
- Video editing is actually working now, it is not completely perfect, but with the right skill/guidance you can just give it 1h footage and it can give you a 5 min highlight clip no problem
- Computer use - getting really rather good, very usable
- Sub agents - it is very fluent at managing sub-agents and speaking to different threads, can help with some new workflows
- Adhering to existing code patterns - I love this, even without asking it would implement something in a way that aligns with you app - major problem for slop generation
- Research - I think it is getting quite a bit better, it still has some bad patterns (e.g being too tactical), but it feels like it is more steerable to be a good researcher
- Multi-day runs - the /goal feature is pretty insane with 5.6-Sol, you can run it for days if you wanted to and it does work. Useful to have another thread or /side to check up on it, but I have some great results with it
- Token efficiency - it is so much more token efficient and faster than 5.5, in reality it is now much faster than Fable too
On the downside, you can feel that Fable is naturally smarter, and I did have some baffling moments with 5.6 when I was getting it to make a fairly simple change in 8 turns - it seemed to get stuck in a dumb stream that was hard to get out of. So it is not AGI, don't get too carried away by the hype.
I have some phenomenal examples that I'm honestly blown away by that I'll share, but as a side anecdote, I have a kind of 'swear meter' which counts how often I'm rude to Codex. In GPT-5.5 era, the % was at around 4-5%, it dropped to 1-2% when I was testing GPT-5.6-Sol and it shot up to 7% when I went back to 5.5 - it was so shocking to go back to 5.5 and experience how much worse it was.
So is GPT-5.6-Sol better than Fable? On pure intelligence - no. But man, I missed it when I just wanted to get sh*t done. It is insanely capable workhorse that you can give any task to and just expect it to be done. No lectures or 'you are absolutely rightisms', nothing is beneath it, if it takes 2 days to do some dirty work, it will do it.
It feels like the first time in a while when we have quite different types of frontier intelligences that benchmark sort of similarly, but feel very different. If you can, you would be probably better off using both and iteratively finding what you'd use Fable or GPT-5.6-Sol for. Perhaps, something like - an architectural discussion with Fable, implementation with 5.6 and docs & comms with Fable.