One of my personal favorite features announced at WWDC will I suspect be a sleeper hit: container machines, allowing your Mac to run a lightweight, persistent Linux environment with your home directory and repos automatically mounted: https://t.co/dOBdfOOVxC
@elonmusk@pbeisel I guess everyone here is ignoring the key word โtraining dataโ. FSD driven miles are the real world test data. I think Teslaโs dataset has far surpassed 10B+ already.
@ylecun@_arohan_@grok How does lejeppa solve the need of teacher ? Donโt we need a teacher to give predictions in tokken space while pretraining the encoder ?
Help me out here ! If they are compressing the keys and values from dxt dimensions to dxc dimensions using a mlp, doesnโt than mean the memory requirement has gone up ? What am I missing here ?
๐ Introducing NSA: A Hardware-Aligned and Natively Trainable Sparse Attention mechanism for ultra-fast long-context training & inference!
Core components of NSA:
โข Dynamic hierarchical sparse strategy
โข Coarse-grained token compression
โข Fine-grained token selection
๐ก With optimized design for modern hardware, NSA speeds up inference while reducing pre-training costsโwithout compromising performance. It matches or outperforms Full Attention models on general benchmarks, long-context tasks, and instruction-based reasoning.
๐ For more details, check out our paper here: https://t.co/HJiqzwnUV7
Why I think deep seek or any non-Indian Ilm is a problem of India.
1) During the printing revolution people got their ideas and knowlege for m books that were printed. All the ideas that couldnโt take the form of books vanished.
3)Now, whoever has the best Ilm will control the narrative. These models can be tunned to act in a certain way to specific responses. Hence if someone wants create chaos in a region they can fine tune it to create uncertainty in minds of people.
@___Harald___ Truly amazing just a quick question. I have been trying to train transformers on time series it fails compared to lstm etc but why does it works so well here im assuming actions are a time series ?
ChatGPT is so last month ๐ด
๐ฅ Stanford/Google researchers just dropped some mindblowing new research on generative agents, and it's like they brought Westworld to life. ๐ค
Here's what you should knowโคต๏ธโคต๏ธโคต๏ธ
Using a simulation video game they created, researchers made 25 characters that could:
Communicate with others and their environment ๐ฌ
Memorize and recall what they did and observed ๐ง
Reflect on those observations ๐ค
Form plans for each day ๐
Then, they gave them some memories:
An identity (name, occupation, priorities) ๐
Information about/relationships with other characters ๐
Some intention about how to spend their day ๐ญ
Then, they pressed play. โถ๏ธ
With just this information alone, the characters acted much like humans do:
They shared information with each other ๐ฃ๏ธ
Example: Isabella starts the day with a plan to host a Valentine's Day party. She spreads the word, and by the end of the simulation, 12 characters know about the party ๐ฅณ
Much like humans, 7 of them flaked - 3 of them had "other plans" and the other 4 just didn't show. ๐
They form new relationships and remember them ๐
Example: Sam and Latoya don't know each other at the start. They meet at a park, and Latoya says she's working on a photography project... ๐ธ When Sam and Latoya meet again later, Sam says: "Hi, Latoya. How is your project going?" ๐ท
They coordinate with each other ๐ค
Example: Researchers gave Isabella (the v-day party host) and Maria two pieces of info:
Isabella: You will throw a party ๐
Maria: You have a crush on Klaus ๐
Without any further instruction, Isabella invites people to the party, decorates the venue, and asks Maria for help. Meanwhile, Maria jumps at the opportunity to get closer to Klaus by inviting him along as well. ๐
This is fascinating new research ๐
We've officially moved past 'AI models can write blog posts for me' and into "How much can AI models act like humans?" territory. ๐คโก๏ธ๐ซ
We're moving fast ๐จ.
Paper - https://t.co/mdbGAJJABC
Credit - @nonmayorpete