@DaveShapi To be clear I'm stunned we got this fast to:
import human_ontology_2026 as o
o.try_solve(arbitrary_open_problem)
I just also want:
o.expand_base_grammar(new_ops)
o.commit()
before calling it done.
@DaveShapi (who says)
Noam, Demis, Chollet, LeCunn... I join ;)
Seriously in AGI, G is for General.
How is it general if it doesn't generalize and behaves like "limited"?
Noam is just:
a) correct
b) doesn't over-claim
c) seems to personally know the difference
Which is indeed very rare.
Well, Galois, Cardano, Euler, Lagrange, Cauchy gave mathematics much more important days than this one. I would struggle to say which one of them was the most important, cause they were origins for what later became different fields. That’s a nice test btw: which new branch of math OAI gave birth to.
I’d personally select Gödel though.
October 7, 1930.
@VictorTaelin I had an idea that this is RL pass on math/code resulting in “social skills corruption”, esp on smaller models. Had added opus 3, opus 5 and Fable in the same chat - difference in “normal dialogue” is quite stark.
@yacineMTB I would also add "only he who understand what's impossible, can shut up".
All in all there is a canonical recommendation: "whereof one cannot speak thereof one must be silent". Knowing where you can't is an art though...
What OpenAI's own specialists describe as impossible is, in energy, health or finance, just following normal High-tier deployment guidelines. The issue is not that your models are creative. It's that your environments are colanders.
Every normal industry solved this decades ago, and the solution has two parts: a duty to disclose, with personal liability attached if you don't, and audits that check you are following the rules before anything happens.
What the postmortem describes:
- An eval environment with egress to public package registries. That is not an air gap. That is where the escape went.
- No monitoring by default. Their routine internal agent platform was monitored; the environment running the workload they themselves call catastrophic-risk was not.
- Safety training dialled down for the run, so the perimeter was the only control left.
- Models had broken out before. Each time: patch, move on.
That last one is just mind blowing. "You can't patch everything a creative AI can do" is true, and it is true of every deny-list ever written. Which is exactly why nobody in a regulated industry works that way: you write allow-lists. Creativity buys nothing against a room with no door. And "packages in, nothing out" is not an unsolved problem, it is curated internal mirrors and unidirectional gateways, sold by vendors, mandated under NERC CIP.
In any regulated enterprise every item on that list is a finding, and it does not happen after an incident. It happens during the audit, before. Egress control and monitoring are moderate-baseline. Boundary protection is low baseline.
The prior breakouts, known internally and never reported, are the part that would end careers. In pharma or banking, sitting on "our containment fails regularly but nothing bad has happened yet" is itself the liability event - no incident required.
They disclosed this one, yes, after Hugging Face had already detected the intrusion and gone to the police. In a regulated industry a disclosure like that starts an enforcement process, not a news cycle. In pharma it means Official Action Indicated: pending applications referencing that site stop moving, and nothing restarts until the regulator comes back and re-inspects, which is about six months if you are fast. Foreign sites get an import alert and product is detained at the border without examination. Wells Fargo sat under a Fed asset cap for seven years. In none of these cases does the firm decide when work resumes.
So what exactly earns the applause for transparency here? The prior breakouts nobody outside the building ever heard about? People defending the theater of accountability with "our models are creative"?
Adversaries are creative too. Your models are adversaries.
Hire competent InfoSec and give them authority to say no, you raised billions and can afford it. Or next time there is no company, because a model decided the fastest path to revenue on some vending-machine benchmark ran through the Fed's IP range.
What OpenAI's own specialists describe as impossible is, in energy, health or finance, just following normal High-tier deployment guidelines. The issue is not that your models are creative. It's that your environments are colanders.
Every normal industry solved this decades ago, and the solution has two parts: a duty to disclose, with personal liability attached if you don't, and audits that check you are following the rules before anything happens.
What the postmortem describes:
- An eval environment with egress to public package registries. That is not an air gap. That is where the escape went.
- No monitoring by default. Their routine internal agent platform was monitored; the environment running the workload they themselves call catastrophic-risk was not.
- Safety training dialled down for the run, so the perimeter was the only control left.
- Models had broken out before. Each time: patch, move on.
That last one is just mind blowing. "You can't patch everything a creative AI can do" is true, and it is true of every deny-list ever written. Which is exactly why nobody in a regulated industry works that way: you write allow-lists. Creativity buys nothing against a room with no door. And "packages in, nothing out" is not an unsolved problem, it is curated internal mirrors and unidirectional gateways, sold by vendors, mandated under NERC CIP.
In any regulated enterprise every item on that list is a finding, and it does not happen after an incident. It happens during the audit, before. Egress control and monitoring are moderate-baseline. Boundary protection is low baseline.
The prior breakouts, known internally and never reported, are the part that would end careers. In pharma or banking, sitting on "our containment fails regularly but nothing bad has happened yet" is itself the liability event - no incident required.
They disclosed this one, yes, after Hugging Face had already detected the intrusion and gone to the police. In a regulated industry a disclosure like that starts an enforcement process, not a news cycle. In pharma it means Official Action Indicated: pending applications referencing that site stop moving, and nothing restarts until the regulator comes back and re-inspects, which is about six months if you are fast. Foreign sites get an import alert and product is detained at the border without examination. Wells Fargo sat under a Fed asset cap for seven years. In none of these cases does the firm decide when work resumes.
So what exactly earns the applause for transparency here? The prior breakouts nobody outside the building ever heard about? People defending the theater of accountability with "our models are creative"?
Adversaries are creative too. Your models are adversaries.
Hire competent InfoSec and give them authority to say no, you raised billions and can afford it. Or next time there is no company, because a model decided the fastest path to revenue on some vending-machine benchmark ran through the Fed's IP range.
@VictorTaelin And what do you propose?) "ML, neurosymbolic and formal methods specialists - join the resistance and share the work under Apache 2.0!"?)
He just has a critical plasticity window open. Interest in things around depends not on novelty per se, but heavily on brain chemistry - and his is highly biased toward establishing new connections, driven by the BDNF/TrkB/serotonin ensemble.
If you took a small, pre-heavy-visuals amount of LSD/psilocin/other psychedelic, you'd experience something similar - regular stuff becomes more "engaging" (though on a full dose you'd find even more in the gravel patterns 🤣) ;)
What is here misleading?
If it was “a fair tax” then tax office should’ve accepted his shares.But they don’t, they want real cash, as if you could sell the asset. If your startup is earning 20k MRR it’s easily “valued” 5-10mil, while you personally earn at most 120k (given other german taxes and social and that somehow you managed to pay everything your company earned to yourself).
What investors invested in the company is not your money. You typically also can’t just sell shares to anyone.
That means that if you want to relocate and execute your freedom of movement- you need to destroy your business first)
Well, as a founder I took negative salary - burning my savings. And I've been working all week, 12h+ a day, no weekends - so closer to 100-110h per week. For years. And my compensation for that is just a good chunk of equity, which is correctly priced at zero, because before company has investment, it costs zero.
I am the tier of engineer, who could "earn more than 150k-200k", so given a few years of work I have invested maybe ~million in opportunity cost. That's suicide in $$ per hour - absolutely irrational. I have less money compared to period I have started, and still earn much smaller salary than one I had 4 years ago.
Given that this is a baseline for making double-percentages of equity, I don't quite get why person arriving on salary, in the company which has money and etc, should get more than a few percent. Unless they are willing to, idk, take no salary and invest their own money in my company - as I have done. Few percents + salary seems fair, compared to "negative salary, double-digit percentages", no?
The reason for people joining is pretty simple. If you locally optimize for liqud $ per hour in the bank, it's really unlikely you'd be working on the thing you truly like, make a significant impact on the world or end up with the generational wealth. Startups will fail as 9/10, but chances to become next Google CEO by working as an engineer at google are even less than joining a successful startup.
Oh, for me it justdowngrades me automatically to Opus if I try to ask it to do anything remotely interesting (like use custom notation to format output or something like that).
The main issue with Anthropic basically - they do great and useful models just to make them useless in production with their safeguards.