A 15-year-old dream has come true today. I started a PhD with the dream of creating a system that chants any Sanskrit shloka perfectly.
And here I am opening sourcing ๐๐๐ ๐๐ก๐๐ง๐ฎ - ๐ ๐ฏแน๐ญ๐ญ๐ (๐ฆ๐๐ญ๐๐ซ) ๐๐ฐ๐๐ซ๐ ล๐ฅ๐จ๐ค๐-๐ญ๐จ-๐๐ก๐๐ง๐ญ ๐ญ๐๐ฑ๐ญ-๐ญ๐จ-๐ฌ๐ฉ๐๐๐๐ก (TTS) ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ ๐๐จ๐ซ ๐๐๐ง๐ฌ๐ค๐ซ๐ข๐ญ. This is the world's first vrutta-aware, open-source TTS for Sanskrit Chanting.
@abhi_venigalla@julien_c Is the 1e15 the perf youre getting on H100s? That's 1 pflop. Would be interested in seeing it in action because that's way higher than what I expect from H100s, flashattn 2 and fp8 and all the other bells and whistles included.
@cis_female@tszzl When you're working at scale the allreduce alone kills a lot of perf no? I don't see anything other than TPUs and Nvidia GPUs perform as well when you're training on thousands of GPUs, even though raw perf should be comparable for many chips. Why is that?
Speculative execution for LLMs is an excellent inference-time optimization.
It hinges on the following unintuitive observation: forwarding an LLM on a single input token takes about as much time as forwarding an LLM on K input tokens in a batch (for larger K than you might think). This unintuitive fact is because sampling is heavily memory bound: most of the "work" is not doing compute, it is reading in the weights of the transformer from VRAM into on-chip cache for processing. So if you're going to do all that work of reading in all those weights, you might as well apply them to a whole batch of input vectors. I went into more detail in an earlier thread:
https://t.co/Lbtpq4VDeY
The reason we can't naively use this fact to sample in chunks of K tokens at a time is that every N-th token depends on what token we sample at time at step N-1. There is a serial dependency, so the baseline implementation just goes one by one left to right.
Now the clever idea is to use a small and cheap draft model to first generate a candidate sequence of K tokens - a "draft". Then we feed all of these together through the big model in a batch. This is almost as fast as feeding in just one token, per the above. Then we go from left to right over the logits predicted by the model and sample tokens. Any sample that agrees with the draft allows us to immediately skip forward to the next token. If there is a disagreement then we throw the draft away and eat the cost of doing some throwaway work (sampling the draft and the forward passing for all the later tokens).
The reason this works in practice is that most of the time the draft tokens get accepted, because they are easy, so even a much smaller draft model gets them. As these easy tokens get accepted, we skip through those parts in leaps. The hard tokens where the big model disagrees "fall back" to original speed, but actually a bit slower because of all the extra work.
So TLDR: this one weird trick works because LLMs are memory bound at inference time, in the "batch size 1" setting of sampling a single sequence of interest, that a large fraction of "local LLM" use cases fall into. And because most tokens are "easy".
References
https://t.co/sIBCSmsyKN
https://t.co/uSpmTzfWhR
https://t.co/7t7orHBybo
@abhi_venigalla@boborado It's crazy that the actual claims are that far off. 1pflop/s would still be 5x+ improvement over A100 which is big deal. But it's still 12.5% of what the actual claims are. What kind of nonsense math did nvidia even run to get that? I assume it involved reading nothing from SRAM.
@abhi_venigalla@boborado We're nowhere close to 50% MFU on H100s though are we? The claims on whitepaper are too insane for us to even get close. With fp8 and all the bells and whistles of sparsity nvidia claims 8pflops/GPU. Would be a surprise if we even get 1.
โHave you ever wondered what happened to the 56 men who signed the Declaration of Independence?
Five signers were captured by the British as traitors, and tortured before they died. Twelve had their homes ransacked and burned. Two lost their sons in the revolutionary army, another had two sons captured. Nine of the 56 fought and died from wounds or hardships of the revolutionary war.
They signed and they pledged their lives, their fortunes, and their sacred honor.
What kind of men were they? Twenty-four were lawyers and jurists. Eleven were merchants, nine were farmers and large plantation owners, men of means, well educated. But they signed the Declaration of Independence knowing full well that the penalty would be death if they were captured.
Carter Braxton of Virginia, a wealthy planter and trader, saw his ships swept from the seas by the British Navy. He sold his home and properties to pay his debts, and died in rags.
Thomas McKeam was so hounded by the British that he was forced to move his family almost constantly. He served in the Congress without pay, and his family was kept in hiding. His possessions were taken from him, and poverty was his reward.
Vandals or soldiers or both, looted the properties of Ellery, Clymer, Hall, Walton, Gwinnett, Heyward, Ruttledge, and Middleton.
At the battle of Yorktown, Thomas Nelson Jr., noted that the British General Cornwallis had taken over the Nelson home for his headquarters. The owner quietly urged General George Washington to open fire. The home was destroyed, and Nelson died bankrupt.
Francis Lewis had his home and properties destroyed. The enemy jailed his wife, and she died within a few months.
John Hart was driven from his wifeโs bedside as she was dying. Their 13 children fled for their lives. His fields and his gristmill were laid to waste. For more than a year he lived in forests and caves, returning home to find his wife dead and his children vanished. A few weeks later he died from exhaustion and a broken heart. Norris and Livingston suffered similar fates.
Such were the stories and sacrifices of the American Revolution. These were not wild eyed, rabble-rousing ruffians. They were soft-spoken men of means and education. They had security, but they valued liberty more. Standing tall, straight, and unwavering, they pledged: โFor the support of this declaration, with firm reliance on the protection of the divine providence, we mutually pledge to each other, our lives, our fortunes, and our sacred honor.โโ
Michael W Smith
I've engineered a hand-crafted Transformer to do long-hand addition. All of the weights were hand-chosen by me. It is perfectly accurate.
How? Fall down the rabbit hole here: https://t.co/bg2bZ3Al3r
Short on time? Here's a summary ๐
1/
@the_yanco@ESYudkowsky@PradyuPrasad That's no different from pretraining where you calculate a "loss" corresponding to how "wrong" the model's output is compared to what it's supposed to be. As far as this analogy is concerned RLHF and pretraining are the same thing.
@ESYudkowsky@PradyuPrasad Why would an LM feel as if taking a training step is like dropping bombs on it whenever it's output doesn't match a desired target?
@JosephJacks_ Why does it matter that they're unreadable? Plus with advancements in interpretability that won't even be the case for long, not that it should matter in the first place.
Is it just my lack of experience with good infrastructure on GPUs, or are TPUs really that much more stable? This is me training a model on a v4-1024 slice, which is basically equivalent to 512 A100s and I've had no incident in almost 5 days! ECC would've shown on GPUs by now.
@abacaj You're never gonna train a 65B param anywhere close to convergence with 2 GPUs, even if it's 2 H100s. The comparison doesn't make sense because it's targeting companies who can pretrain extremely large models. The "smaller toes" will be comfortably far behind, dw.
Visualize transformer attention!
AttentionViz, created by Catherine Yeh and expanded by Yida Chen, helps you explore transformer self-attention by visualizing query and key vectors in a joint embedding.
Paper: https://t.co/GKom1BN5Zl
Website: https://t.co/rCeA2e5xKv
@norabelrose It's almost certainly an MoE with 1.6T params and trained for like 10T tokens. Have been hearing that for a year now. Trillion param dense model training is still infeasible for everyone.