Remember the llm.c repro of the GPT-2 (124M) training run? It took 45 min on 8xH100. Since then, @kellerjordan0 (and by now many others) have iterated on that extensively in the new modded-nanogpt repo that achieves the same result, now in only 5 min!
Love this repo ๐ 600 LOC
LLMs have complex joint beliefs about all sorts of quantities. And my postdoc @jamesrequeima visualized them! In this thread we show LLM predictive distributions conditioned on data and free-form text.
LLMs pick up on all kinds of subtle and unusual structure: ๐งต
@beneater Long way to go for nuclear weapons. First, the retention bot will start using prompt injection attacks on the cancel bot and you'll end up upgrading your subscription.
@aureliengeron For the "easy" problem. In the very artificial scenario that you keep rejection sampling two whales until you get at least one male, then 1/3. In the real-life scenario where you only looked at one and formulated your question based on its observed sex, then 1/2.
@lawrennd We're doing pretty well here in South Australia, where Bayesian statistics is used for tracking. That's what we need, less flashy AI and more rigorous statistics.
@ducha_aiki@chriswolfvision@shortstein It's the editor's job to find qualified reviewers. If you get nominated for review then just assume you're qualified. You don't have to be qualified for everything in the manuscript, especially if it's cross-disciplinary. Just mention which parts you felt confident to review.
@nschawor Looks like the drift is shared by the conditions. Maybe fitting a Gaussian Processes to the mean would be an ideal form of de-trending. With two GPs you could also estimate the time-varying difference. Add a couple more GPs and heteroscedasticity could also be estimated.
@canyon289 We use "trace" often in physiology, and I think the name stuck from when a needle used to trace out a line on soot-covered paper in a kymograph https://t.co/hirwMiAP7E
@spaceLem@durand_sinclair@SwiftOnSecurity@vboykis I should've been more mathematically rigorous and like you said "nearly" all. I guess Fortran is written in assembly which doesn't really have arrays.