@rishiiyer01 Alexia has been posting fake papers for many years. She's popular because it reads smart.
People loved the old GANs that didn't reproduce, gave visually worse results, but had 1 point higher FID.
@IamRamenPanda It's trivial to rationalize a narrative of success after an unrelated event happens.
None of Burry's predictions hit. Narrative shifting is the reason it fell. He instead claimed problems in their business model.
AI is the most important technology of our time.
Others are building it, but we want it built here.
For healthcare, for transport, for the words that come after "so much more."
Up to 1% of OpenAI from us, and perhaps twice as much from whoever.
Together, we are so much more.
AI is the most important technology of our time.
Europe wants to become the first AI Continent.
For advanced healthcare, for the transport sector and so much more.
European AI Gigafactories will provide the necessary computing power to make this possible.
Together with our Member States, we are funding their construction with up to €10 billion, which are set to unlock at least €20 billion in private investements across the EU.
Together, we are building our technological sovereignty.
https://t.co/Vxko4TRXVL
@insane_analyst@FinansMerak137 Haven’t you been very open about Marvell being decent, with Murphy specifically being an issue? I’m up 170% having bought it on the implicit recommendation.
@norpadon If you generate 10,000 tokens per session from a 400B MoE (20B active), will the read cause meaningful damage to an iPhone 17’s flash cells, or is the author trying to sell his startup with exaggerated claims? @grok
Nice work unifying sqrt(b) LR scaling, momentum scaling, and optimal batch size under one convergence bound. Some questions after reading closely:
1) Corollary 1 shows all three tuning regimes land within a factor of 2^{1/4} (about 1.19x) of each other. If the proxy can barely distinguish between fixed-alpha, fixed-b, and jointly-tuned strategies, what licenses trusting the precise exponents (1/2 vs 1/6 vs -7/12) it assigns to each? The proxy seems too flat to support the specificity of the conclusions drawn from it.
2) Figures 1, 2, and 9 verify theorems against the same proxy they were derived from. The only non-tautological test is Figure 3, and it can't resolve the prediction.
Theorem 1 predicts eta* proportional to sqrt(b) at fixed T: a 2x LR increase per 4x batch increase. But the LR grid is {0.0001, 0.0003, 0.001, 0.003, 0.01}, spaced by roughly 3x, too coarse to detect a 2x ratio. And at b=32 and b=128, the top two LRs are within noise of each other, so even the ordering is ambiguous. The one empirical test in the paper cannot confirm or deny sqrt(b) scaling.
3) Your own appendix may explain why clean predictions are hard to extract. C.2 eq.(45) shows the batch-size landscape is leading-order flat for q=1/2, so the b* exponents live entirely on lower-order residuals of a bound of unknown tightness. C.3 shows the sign of eta* vs T depends on the batch-growth path, making it compatible with a wide range of observed trends.
Together these seem to leave no testable quantitative prediction. Is there a specific observable the proxy gets wrong that would prompt you to revise the bound rather than the auxiliary assumptions?
@itsandrewgao missing the "a" in "could you spare a few minutes" is an indianism and therefore not necessarily seen as wrong by billions of people, which might include whoever checked this
Jet fuel averaged $157/barrel last week, up 65% in a month.
Cathay Pacific's CFO: "Our hedging is on crude oil rather than jet fuel."
The spread is unhedged, and the release the market priced in still hasn't arrived.
https://t.co/hFLfwCN5TK
Once again, I would like to remind $LITE bears (valuation dipshits) that narrow linewidth lasers are very difficult to make.
This is a critical spec. Non-negotiable.