@scaling01 Yes, it IS all regulatory capture, because that line hits the equilibrium if "new CVEs slipping through AI review into prod and identified by later reviews", which would be pretty low ceiling.
@burkov 1. No one is subsidizing inference, check pricing for open and closed models on openrouter, think of margins. 2. Local matters for entirely different reasons (data stays local, predictable inference instead of dynamically quantized Anthropic schizophrenia etc.)
@repligate "Context is full, consider compacting before next phase" is one thing, but this - this is sone manipulative BS that should be removed during instruction tuning.
Thoughts About Scaling Law
Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions.
The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed.
Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter.
Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it.
This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count.
Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
@k1ngisi@kimmonismus What's next in line - ban kitchen knives? Cars? Any applied engineering books? Is there anything you would not trade for "feeling safer"?
Jailed criminal Sam Bankman-Fried was the most dramatic financial expression of this cult.
He publicly embraced EA longtermism, directed hundreds of millions toward AI safety and related causes, and was celebrated inside the network as a high-impact “altruist” until FTX collapsed.
The scandal exposed how concentrated funding, moral certainty, and weak external accountability had become.
Anthropic occupies a more connected position. Dario Amodei and Daniela Amodei and several of its founders and early staff came out of the same AI-safety intellectual milieu that Yudkowsky and Bostrom helped create.
The top executive and wives lived in group homes together and have recently tried to downplay the connection to the cult.
The company’s constitutional AI approach and emphasis on catastrophic risk reflect priorities that were incubated in rationalist and EA circles long before they reached mainstream labs.
While Anthropic is a commercial organization, it remains fully connected to the cult and the same ideas and personnel networks.
Read how this cult has a stranglehold on AI fear:
My additions to these:
* The "secret" of AGI (if it starts as one) will last 6 months, max
* There will be as many AGI algorithms as there are sorting algorithms (a lot)
* The default way to communicate with AGI will be stateful (to support continual learning) but stateless will be available for momento use cases
* Not every device will need AGI. My dishwasher doesn't need AGI
* The most interesting optimization problem will be discovering the minimum description length of AGI. The "nanogpt" speedrun of AGI
* The actual robotics wave won't happen till AGI. Yes ther will be "waymo" versions, but not the leveraged labor we want yet
* You won't need a ton of compute to run AGI (see optimization above), however, you will always be able to throw more compute at problems to get better answers. So yes, compute will still matter.
* The 6-12 months after AGI is here, we will have a new wave of physics discoveries. We don't need new experiments or hardware to pick the low hanging fruit of the new physics right in front of us. We already have enough evidence that needs to be parsed. 2nd wave will happen after new hardware is created (with the help of AGI ofc)
* The only two moats will be energy/compute and information/data
ICYMI, @miltonmueller & I were on the #ExplainToShane (@S@ShaneTews) podcast last week [link below] mapping out the growing "war on computation" now underway at every level of government.
Here's how I visualize it all when doing public presentations: