This year, I read ten important historical novels: Jane Eyre, Middlemarch, To The Lighthouse, Bleak House, Portrait of a Lady, Anna Karenina, Life and Fate, Heart of Darkness, Madame Bovary, and The Magic Mountain.
Reflections:
• Four of these are more than 800 pages long. The Magic Mountain and Portrait of a Lady, while shorter, are not short. Of the ten, 5 are British, 2 are Russian, and there was one from each of France, Germany, and the US.
• For me the clear standouts are Middlemarch, Bleak House, Karenina, and Life and Fate. I would enthusiastically reread any of them. If I had to choose just one to go to again, I would probably select Middlemarch. There's something memorably compelling in Eliot's affection and empathy for almost all of her characters. If Succession is a show with no likable personalities, Middlemarch is the opposite. Bleak House is a close second. Life and Fate is quite different to the others: it’s not exactly entertaining (or even notably well-written), but it is true and profound. (Most works designated “important” are not, but Life and Fate surely merits that as well.) If kindness is one of the core adjurations of Life and Fate, Eliot is the author that most embodies it.
• I'd underestimated Dickens's lyricism. I had thought of him as a master of the plot (contra Nabokov), but he is just as accomplished in prose itself. “Chesney Wold is shut up, carpets are rolled into great scrolls in corners of comfortless rooms, bright damask does penance in brown holland, carving and gilding puts on mortification, and the Dedlock ancestors retire from the light of day again. Around and around the house the leaves fall thick, but never fast, for they come circling down with a dead lightness that is sombre and slow. Let the gardener sweep and sweep the turf as he will, and press the leaves into full barrows, and wheel them off, still they lie ankle-deep. Howls the shrill wind round Chesney Wold; the sharp rain beats, the windows rattle, and the chimneys growl. Mists hide in the avenues, veil the points of view, and move in funeral-wise across the rising grounds. On all the house there is a cold, blank smell like the smell of a little church, though something dryer, suggesting that the dead and buried Dedlocks walk there in the long nights and leave the flavour of their graves behind them.”
• Three of these four were written by authors in their 50s (Eliot, Tolstoy, and Grossman). Dickens was a mere 41 -- and this shows. The plot is very entertaining, and immensely intricate, but the characters are somehow flatter. So, maybe one lesson from the set is simply that wisdom is real, and that skill in the domain of fiction compounds for quite some time.
• Russian literature puzzles me. Why did it suddenly become so good in the 19th century, and why did it decline so much in the 20th? I don't think the latter answer can just be a story of oppression, since we got many great works during Stalin’s reign. But what's the best Russian novel since Master and Margarita? On the issue of the rise, I often encounter explanations claiming that it was related to Russian intellectuals being excluded from political influence and consequently retreating to the artistic domain -- but this feels obviously inadequate. Again, how does this explain the post-Bulgakov decline? And where are the great, say, Saudi works of the past 50 years?
• Whatever happened to the novel around the turn of the century (Conrad, Woolf, Mann in my reading) was not obviously salutary. All three are interesting works, and there is something very distinctly modern in Woolf's in particular, but they simply don't compel -- at least for this reader -- the way their predecessors do: maybe it's just the particular selection, but I was generally looking forward to finishing the early 20th century works, and a bit disappointed when completing those dating from before 1900. The dislocation that Blom describes in Vertigo Years is clearly manifest. Woolf's “Mr. Bennett and Mrs. Brown” essay, and her claim that “human character changed” in 1910, is consistent with the turn in the novels. She was speaking of different works, but her assessment rings true in a broader way: “Yet what odd books they are! Sometimes I wonder if we are right to call them books at all. For they leave one with so strange a feeling of incompleteness and dissatisfaction.”
• I should note that some of Woolf’s descriptions are great, even if her brooding interiority leaves me ultimately unenthralled. “The house was left; the house was deserted. It was left like a shell on a sandhill to fill with dry salt grains now that life had left it.” “He lay on his chair with his hands clasped above his paunch not reading, or sleeping, but basking like a creature gorged with existence.”
• Tom Wolfe attributes modern architecture and the international style to a post-1917 sympathy for the proletariat and a desire to strip indulgent bourgeois ornament from our construction, and, yes, Schoenberg explicitly motivated atonality in egalitarian ideals, but this set of novels makes me doubt the political explanations. You can clearly see the embrace of some kind of disharmony in the books, and I don’t think Conrad was trying to make any Marxist point. I still struggle to explain what happened, but I think I would reverse some of the standard causality, and it seems to me that the coopting of communist ideals is probably itself downstream of the broader social unease that also gave rise to these artistic tides. Blom’s description of the rise of various mental disorders -- neurasthenia and the like -- seems relevant. (All of this does make me want to better understand 1848.)
• The railway, and its attendant social upheaval, features repeatedly, and maybe most memorably in one chapter of Middlemarch. I hadn't appreciated just how significantly disruptive a force it was perceived as being even at the time. (Given the scale of the construction that was entailed, maybe this shouldn't be surprising.) More broadly, there is some sense of a society in transition through most of these works: these aren't neat and timeless tales. You have the rise of the bourgeois and broader urbanization in Bovary, the emerging social consciousness in Bleak House, the exposure of the shabbiness of Victorianism and its gender expectations in Lighthouse, and the postwar shell shock of Magic Mountain.
• The works written before 1900 are primarily about romance (Bleak House the exception, with romance only a subplot), and those written afterwards (Conrad, Mann, Grossman, Woolf) are emphatically not. I don’t know what to make of this. Perhaps just an accident of the selection.
• Ruxandra Teslo points out that there’s a moral gravity in the 19th century works that seems foreign today: people treat their own characters as important constructions in their own right. In a similar vein, I was struck by Grossman’s conception of freedom: he perceives it more as the right to self-define than a more typical liberty of action. Perhaps because actions were so circumscribed in Victorian societies (for women) and Soviet societies (for everyone), the seriousness of being weighed heavily.
* Money and its mechanics get extensive treatment in the pre-1900 novels. The details of Bovary’s debt were made famous by Piketty, but Eliot also spends time on Lydgate’s financial struggles, and Tolstoy on Levin’s agricultural economics. Pecuniary considerations are absent in the later works. Again, maybe just happenstance stemming from the particular selection, but I don’t get the feeling that it’s just that: I think something about authors’ attitudes to the topic changed.
• Today’s scientific papers are far harder to read, and jargon-replete, than those of 1960. However, the novels of the 19th century use significantly more sophisticated construction (and vocabulary) than those of today. What should we make of the countervailing trends? To me, both seem suboptimal.
• Pleasure aside, should one read these books? Does one derive moral betterment from doing so? I'm not sure. Probably not in any narrow sense. Ethicists are supposedly no more ethical than regular people -- if deliberate study doesn't help, what hope does mere fiction have? And, anecdotally, I don't consider the humanities majors to be the moral betters of the STEM students. I do think they've helped with my understanding of history, though. This year, I reflected on how the major historical moments that I've lived through -- the weeks after 9/11, the aftermath of Trump's 2016 election victory, March of 2020 -- cannot really be understood in terms of particular events, and must instead be apprehended through the vibes that prevailed. Rather than trying to assemble a logical causal chain, I think it's more helpfully explanatory to see many happenings as simply arising from a mood. History books struggle to capture such sentiments, and understandably so: the historian usually wasn't there; even if they were, vibes are ethereal things, and they feel out-of-place in a work that aspires to footnoted rigor and exactitude. As such, complements are required, and these novels have definitely helped me. This view also makes biography and autobiography seem of greater importance in developing such comprehension. The small details -- that Herbert Hoover's parents used to attend lectures and debates in a nearby town since that was the only entertainment available, that both died before age 35, or that Hoover himself once walked 80 miles in 3 days to join a geology class trip -- say a lot about a period, and are rarely captured in the grand sweep of events. I feel like I gained much more understanding of historical Vienna and of the emotions around WWI from reading Zweig's memoir than from any direct history of the period.
• Another argument made for reading these works is to simply better understand humanity and the human experience. There is almost certainly some extent to which this argument is valid, though I always wonder: do they help you better understand humanity, or better understand the kind of people who write books like these? Is Isabel Archer actually reflective of someone in that kind of position, or merely of the kind of hyper-intellectual James family? Karenina is ultimately a kind of demented obsessive (as was Tolstoy) – in learning about her, do we learn about love and its travails, or simply about unusually unstable personalities?
• There’s clearly some value in reading them for somewhat tautological reasons: they're worth reading because they are the books that we’ve decided are worth reading. They form part of our cultural context, and other works probably make somewhat more sense and are more memorable when interpreted through their lens. They are intellectual capital cities: you sorta have to go to Paris and New York in order to understand the rest of the world, and whether you “enjoy” them isn’t really the operative question.
• Ultimately, a utilitarian case for better understanding history or even humanity would not be my primary argument for why one might choose to read them, though. With self-consciousness about the platitude, they are simply some of the finest intellectual achievements of humanity, and worthy of engagement for that reason alone: a deeper appreciation for excellence is itself a valuable thing.
I can now no longer stumble across new ads without asking the magic 3 Q's (thanks to @harrydry)...
Can I visualize it?
Can I falsify it?
Can nobody else say it?
I got ahead of myself when I announced this project, and I am sorry. That was not my intention. I made a decision to ship this new approach based on the information that we had at the moment.
I know that many of you are excited about the potential for this and are now skeptical. Nobody is more excited about the potential for this approach than I am. For the moment, we have a team working tirelessly to understand what happened and will determine how to proceed once we get to the bottom of it. Once we have all of the facts, we will continue to be transparent with the community about what happened and next steps.
I'm excited to share that we've built the world's most capable AI software engineer, achieving 30.08% on SWE-Bench – ahead of Amazon and Cognition. This model is so much more than a benchmark score: it was trained from the start to think and behave like a human SWE.
LLaMA-3 is a prime example of why training a good LLM is almost entirely about data quality…
TL;DR. Meta released LLaMA-3-8B/70B today and 95% of the technical info we have so far is related to data quality:
- 15T tokens of pretraining data
- More code during pretraining (leads to better reasoning capabilities)
- More efficient tokenizer with larger vocabulary
- Super sophisticated (including LLM components) data quality filtering
- Extensive empirical analysis of data mixture
- Focus on quality filtering of post training data (for SFT/RLHF/DPO)
All of the cool stuff in this report is related to how to curate data effectively for pre/post-training! This really shows that data curation/filtering is the most difficult and impactful aspect of training foundation models.
(1) Model architecture: Only 5 sentences are provided about the model architecture, which simply state that LLaMa-3 uses a standard decoder-only architecture with grouped query attention to improve inference efficiency (and a longer 8K context). It’s pretty clear that model architectures are becoming standardized, and most of the research focus is going into constructing datasets. In fact, the main architecture modification made by LLaMA-3 is a more efficient tokenizer!
“Llama 3 uses a tokenizer with a vocabulary of 128K tokens that encodes language much more efficiently, which leads to substantially improved model performance.” - from LLaMA-3 blog
(2) Better tokenizer: LLaMA-3 comes with a custom tokenizer with a vocabulary of 128K tokens (LLaMA-2 had a vocabulary of 32K tokens). This tokenizer is more token efficient (i.e., fewer tokens are necessary to encode the same piece of text relative to LLaMA-2), which makes inference more efficient. Authors also note that the new tokenizer improves performance! In other words, making sure that we are encoding the model’s input data correctly is super important.
(3) Massive pretraining corpus: LLaMa-3 is pretrained over 15T tokens of text (5% non-English), which is a 7X improvement over LLaMA-2 and even larger than the 12T pretraining corpus of DBRX. The pretraining corpus also has 4X more code relative to LLaMA-2 (this was a big criticism of LLaMA-2). With this in mind, it’s not a surprise that LLaMA-3 has strong reasoning/code capabilities—several papers have correlated pretraining on code to better downstream reasoning in LLMs.
“We found that previous generations of Llama are surprisingly good at identifying high-quality data, hence we used Llama 2 to generate the training data for the text-quality classifiers that are powering Llama 3.” - from LLaMA-3 blog
(4) FIltering pretraining data: Few concrete details are provided on the filtering process for the pretraining corpus of LLaMA-3, but it’s clear that a lot of filtering is done. These filters include heuristic filters, NSFW filters, semantic deduplication, and text classifiers to predict data quality. Plus, authors note that LLaMA-2 is very good at detecting text quality, so they use these models in the filtering process (see above). Authors also mention that they do extensive empirical analysis to figure out the correct data mixture (DBRX also mentions this is hugely important).
(5) Overtraining: Chinchilla proposed the compute optimal training regime for LLMs, but recent work indicates that pretty much everyone overtrains their LLMs relative to the compute-optimal ratio. LLaMA-3 is pretrained on two orders of magnitude more data (for the 8B model) beyond the compute-optimal ratio, and we still see log-linear improvements. Sure, we could train a larger model on fewer tokens and achieve similar performance while spending less on training compute. But, this doesn’t consider inference costs! We almost always will pay for more training compute if it means we can deploy a smaller model with the same performance.
“The quality of the prompts that are used in SFT and the preference rankings that are used in PPO and DPO has an outsized influence on the performance of aligned models.” - from LLaMA-3 blog
(6) Post training data quality: Even beyond pretraining, data quality is pivotal for LLaMA-3! The model is aligned with a combination of SFT, rejection sampling, PPO, and DPO. During alignment, authors claim that the quality of supervised/preference data is super important. In fact, the biggest quality improvements in LLaMA-3 came from curating this data and performing multiple rounds of quality assurance on humans annotations!
Local Stable Cascade 1 Click Launcher
Just wrote a 1-click launcher for this Stable Cascade Gradio app. Works very well!
Works on all platforms: Windows, Mac, Linux
Ten months ago, we launched the Vesuvius Challenge to solve the ancient problem of the Herculaneum Papyri, a library of scrolls that were flash-fried by the eruption of Mount Vesuvius in 79 AD.
Today we are overjoyed to announce that our crazy project has succeeded. After 2000 years, we can finally read the scrolls:
This image was produced by @Youssef_M_Nader, @LukeFarritor, and @JuliSchillij, who have now won the Vesuvius Challenge Grand Prize of $700,000. Congratulations!!
These fifteen columns come from the very end of the first scroll we have been able to read and contain new text from the ancient world that has never been seen before. The author – probably Epicurean philosopher Philodemus – writes here about music, food, and how to enjoy life's pleasures. In the closing section, he throws shade at unnamed ideological adversaries – perhaps the stoics? – who "have nothing to say about pleasure, either in general or in particular."
This year, the Vesuvius Challenge continues. The text that we revealed so far represents just 5% of one scroll.
In 2024, our goal is to from reading a few passages of text to entire scrolls, and we're announcing a new $100,000 grand prize for the first team that is able to read at least 90% of all four scrolls that we have scanned.
The scrolls stored in Naples that remain to be read represent more than 16 megabytes of ancient text. But the villa where the scrolls were found was only partially excavated, and scholars tell us that there may be thousands more scrolls underground. Our hope is that the success of the Vesuvius Challenge catalyzes the excavation of the villa, that the main library is discovered, and that whatever we find there rewrites history and inspires all of us.
It's been a great joy to work on this strange and amazing project. Thanks to Brent Seales for laying the foundation for this work over so many years, thanks to the friends and Twitter users whose donations powered our effort, and thanks to the many contestants whose contributions have made the Vesuvius Challenge successful!
Read more in our announcement: https://t.co/rUlrdGXBMs
Thinking about making a list of some of the most impactful writings about tech and venture. Figure I'd let elon pay for the hosting costs.
Really interesting that Unicorn came from a data gathering project. Mostly descriptive, some great insights.
https://t.co/NA4uvyK0Ey
@lizo_mzimba So often articles have a fluffy ending, grounding the article in other recent events, or just generally wrapping things up. This is just such an amazing way to do it. Well done, this now lives rent free in my head
I am so excited to tell the world about what we've been working on for the last few months. World, say hello to @BesteverAI — GenAI tool for image & video ads.
The easiest way to generate creatives for campaigns is here.
You'll soon see lots of "Llama just dethroned ChatGPT" or "OpenAI is so done" posts on Twitter. Before your timeline gets flooded, I'll share my notes:
▸ Llama-2 likely costs $20M+ to train. Meta has done an incredible service to the community by releasing the model with a commercially-friendly license. AI researchers from big companies were wary of Llama-1 due to licensing issues, but now I think many of them will jump on the ship and contribute their firepower.
▸ Meta's team did a human study on 4K prompts to evaluate Llama-2's helpfulness. They use "win rate" as a metric to compare models, in similar spirit as the Vicuna benchmark. 70B model roughly ties with GPT-3.5-0301, and performs noticeably stronger than Falcon, MPT, and Vicuna.
I trust these real human ratings more than academic benchmarks, because they typically capture the "in-the-wild vibe" better.
▸ Llama-2 is NOT yet at GPT-3.5 level, mainly because of its weak coding abilities. On "HumanEval" (standard coding benchmark), it isn't nearly as good as StarCoder or many other models specifically designed for coding. That being said, I have little doubt that Llama-2 will improve significantly thanks to its open weights.
▸ Meta's team goes above and beyond on AI safety issues. In fact, almost half of the paper is talking about safety guardrails, red-teaming, and evaluations. A round of applause for such responsible efforts!
In prior works, there's a thorny tradeoff between helpfulness and safety. Meta mitigates this by training 2 separate reward models. They aren't open-source yet, but would be extremely valuable to the community.
▸ I think Llama-2 will dramatically boost multimodal AI and robotics research. These fields need more than just blackbox access to an API.
So far, we have to convert the complex sensory signals (video, audio, 3D perception) to text description and then feed to an LLM, which is awkward and leads to huge information loss. It'd be much more effective to graft sensory modules directly on a strong LLM backbone.
▸ The whitepaper itself is a masterpiece. Unlike GPT-4's paper that shared very little info, Llama-2 spelled out the entire recipe, including model details, training stages, hardware, data pipeline, and annotation process. For example, there's a systematic analysis on the effect of RLHF with nice visualizations.
Quote sec 5.1: "We posit that the superior writing abilities of LLMs, as manifested in surpassing human annotators in certain tasks, are fundamentally driven by RLHF."
Congrats to the team again 🥂! Today is another delightful day in OSS AI.
@jennybulstrode But that doesn't ring true. Heck, some of those innovations are still being used in Ukraine. Anyways, I'm sure you are actually an expert on Britain in the industrial revolution, so you probably don't care for the USSR, but thought I'd ask. Thanks so much for the great work!
@jennybulstrode Amazing work, thanks for this contribution to the history of innovation. Can I request a research paper? LOL. I'm interested to know what innovation looked like under the Soviet regime. Here in the US it's easy to wave it off as "oh they were just spying on us" or something silly