New work with @AlecRad and @DavidDuvenaud:
Have you ever dreamed of talking to someone from the past? Introducing talkie, a 13B model trained only on pre-1931 text.
Vintage models should help us to understand how LMs generalize (e.g., can we teach talkie to code?). Thread:
I still see a lot of people discussing LLMs as next-token predictors, which is by now quite a misunderstanding. A related opinion is that LLM progress will probably plateau. This post explains why I don't think the "plateau" argument holds up. https://t.co/fJPBoWs2aX
@chrisalbon@infinitehumanai Tell it to create reports with charts that let you explore inputs and outputs. Also to include code snippets etc. just as you would do with a direct report. The inspection layer does not need to be code.
I migrated cursor.com from a CMS to raw code and Markdown.
I had estimated it would take a few weeks, but was able to finish the migration in three days with $260 in tokens and hundreds of agents.
Here's how I did it + all my my usage stats.
https://t.co/QIAOmLsffx
Today, we at @OpenAI achieved a milestone that many considered years away: gold medal-level performance on the 2025 IMO with a general reasoning LLM—under the same time limits as humans, without tools. As remarkable as that sounds, it’s even more significant than the headline 🧵