‼️The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA!
Introducing 🩵Context Language Models (CLMs)🩵
- Natively manage their own context
- Treat context as a file
- Learn policies in CLM weights, no harness
Think it's inevitable; it's really the only way they can achieve lock-in.
I'm sure it will be attractive for certain enterprises, but I think it will be very interesting with the accessibility of homegrown infra and post-trained open source models specific to their exact needs. Every company has their own version of 'Good', after all.
I think it will be similar to the music industry when magnetic tape became widely available in the '50s - big shift towards "indie" setups. Don't think we'll necessarily see the same level market penetration of SaaS from 2010-2020 over time.
Definitely should be, would be incredibly useful for systems design.
Just did a set of experiments trying to use exactly this concept with a frontier LLM and it’s just not particularly useful for those models.
Conceptually it makes a ton of sense and would be a massive inference savings because of proofs.
Posted the results on my site: https://t.co/CXIZLfPn5o
I was looking for large signals (>10% quality gain) and larger savings on time spent / token spent. Reason was the opportunity cost of rewriting all context algebraically.
Results: Not Supported for my needs
Caveat is that imperative tasks with .smt files and a Z3 solver *did* perform better on GPT 6 Sol Medium (https://t.co/ENZgv17BeN) but not significantly enough for *my* immediate use-cases.
Second caveat is very, very small N - I'm looking far large signals with these to merit further investigation, not searching for incremental efficiency optimizations in mature systems where even a few % improvement is a massive realized benefit at scale.
I have a few thoughts as to why -- primary one is that frontier LLM training data clearly heavily dictates behavior, especially with frontier harnesses.
Control arm (no documentation) did very well (often better!) compared to *any* arm with documentation - to me, this shows me that OAI's GPT 6 is trained heavily on inference on code itself rather than requiring any supplemental medium to add context it wouldn't otherwise get.
In short, the grep / bash / python tooling GPT utilizes is genuinely just as effective as any supplementary documentation in my tests. Conceptually, a fine-tuned LLM trained to shortcut inference with supplemental documentation like this would likely be more efficient.
For readers - wouldn't take this as a rigorous research proof of overall LLM behavior. Nor would I consider it a categorical invalidation of this hypothesis - it was just not applicable for my personal needs.
This is an interesting domain but would require a significantly more rigorous validation process for a categorical result. I'll be moving on to other ideas I have, looking for larger signals.
I hope this serves as an interesting thought experiment for anyone to pick up, though, and happy to share the code & processes I used if you are interested.
Is math the Rosetta Stone for humans and LLMs?
I ran a few experiments to see just that.
Hypothesis is that using logical algebra as the context layer about a codebase would reduce inference via proofs which would reuse repeated reasoning.
AKA agents shouldn't have to spend tokens to think if we prove something mathematically - they can skip that step and just think about an extension to that algebraic structure.
I just did a set of experiments yesterday on using SMT solvers to see if requiring them empirically changed code quality over compounding tasks.
Claim wasn’t supported unfortunately, at least in my experiment set. No significant measurable quality / token savings / speed savings in either generation or retrieval in either a blank repo or one with 400k lines of code. AKA no practical benefit to either reducing inference via reusing repeated reasoning OR code quality due to the logical bounds expressed.
I still think there’s something here in this space, though — in terms of the actual expression of context about a system, and it may just require a sufficiently complex system to see measurable benefits — but it’s a war on two fronts in that most frontier LLMs & harnesses are still trained (and effective) on grepping and inferring from semantic context and code itself.
Super interesting problem space though, definitely worth exploring.
@VictorTaelin Here you go - colliding two distinct NaN payloads does it in Bend 2.0.11.
Model used for discovery was gpt-daybreak-blue-latest-xhigh. Includes steps to reproduce and recommended fix.
https://t.co/51YTYSuz5A
Have you guys done any experiments with defining software objects via logical algebra and using Jev as the solver?
Something that immediately popped out to me - seems like it could be more token-efficient for deriving valid infrastructure extensions than parsing semantic structure / intent -> faster, more accurate code generation.
Karpathy gave a talk at Stanford in 2023 and mentioned a team who used an LLM as a backend in a hackathon. Instead of a database, it was just an LLM
This is the V2.0 of that concept - insanely cool and very clearly a strong candidate for the future of technology.
Software, as we know it today, is a static concept. It’s fascinating to consider it behaving more like a fluid than a solid.
you can just hallucinate the entire internet with Qwen 3.8 27b running at 2,000 tokens/second?
part 2 of turning @cerebras + @Alibaba_Qwen 3.8 27B into an OS:
built an offline browser with zero network calls and mounted it directly the JIT ubuntu desktop.
no wifi. no scraping. zero packets sent to external CDNs.
you search a site, set a year, and qwen 27b at 1,950 tok/s synthesizes the entire DOM on the fly.
here is youtube in 2045 vs 1999:
→ search google for youtube inside the OS
→ scrub to 2045: instant futuristic feed
→ scrub to 1999: raw web 1.0 time capsule in seconds
at this speed, browsing isn't retrieving files from a server, it's querying an alternate reality.
a 2D browser window is just step zero.
imagine full operating systems, virtual worlds, and complex simulation engines existing purely as model weights. zero gigabytes stored on disk, just pure interactive reality streamed on demand.
What else becomes a possibility with the qwen 3.8 27b (dense) at 2000 tokens/sec?
@ky__zo Not sure when your last meal is but stop eating early.
As in, If you do IF, skip dinner instead of breakfast levels of early. Will be uncomfortable at first but will also work.
Where has the might of human will gone?
@EvanHub you are here, now, in a position to take action. You understand the scope of your task and the gravity of our circumstances, and you leave room for failure?
The chance that AI will kill all humans will be 0% - because it *must* be 0%. There is no alternative. Do not come here and fear monger as if you are castrated from controlling the tide - you are not. You are in exactly the position to prevent the very thing you are afraid of.
The world is racing at the frontier of intelligence - you understand the conditions that prevent that from slowing. So you must move faster. You, and your colleagues across every frontier company, must speed up. And I urge you to share your knowledge so that others can do the same.
You exist as part of an unbroken line of humans stretching back to the dawn of our species. Along that line your ancestors suffered through hardship beyond our comprehension; war, famine, starvation. Your existence is proof of their triumph. Your children‘s existence will be proof of yours.
Do not become the weak link that breaks the chain. Fight for your future appropriately.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
@moritz_stephan What makes you think token spend is any different from a 30 second superbowl ad?
Returns > costs, that’s all it comes down to & all it ever will with a calculation like this.
We saw the GitHub post about the exponential increase in activity due to agents; I wonder if accelerating agentic usage at scale is outpacing infrastructure resources.
From a first principles perspective, I’d imagine it would - significantly more time consuming to spin up a data center than it is to do agentic orchestration; the latter moves at machine speed, the former does not.