Everything you love about generative models — now powered by real physics!
Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications.
Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: https://t.co/bEkIlCKqdf).
The Genesis physics engine and simulation platform is fully open source at https://t.co/DhBv7NdyqH. We'll gradually roll out access to our generative framework in the near future.
Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism.
We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications.
Open Source Code: https://t.co/DhBv7NdyqH
Project webpage: https://t.co/SBNyhFB0yn
Documentation: https://t.co/3yuBoaealV
1/n
The #NobelPrizeinPhysics2024 for Hopfield & Hinton rewards plagiarism and incorrect attribution in computer science. It's mostly about Amari's "Hopfield network" and the "Boltzmann Machine."
1. The Lenz-Ising recurrent architecture with neuron-like elements was published in 1925 [L20][I24][I25]. In 1972, Shun-Ichi Amari made it adaptive such that it could learn to associate input patterns with output patterns by changing its connection weights [AMH1]. However, Amari is only briefly cited in the "Scientific Background to the Nobel Prize in Physics 2024." Unfortunately, Amari's net was later called the "Hopfield network." Hopfield republished it 10 years later [AMH2], without citing Amari, not even in later papers.
2. The related Boltzmann Machine paper by Ackley, Hinton, and Sejnowski (1985) [BM] was about learning internal representations in hidden units of neural networks (NNs) [S20]. It didn't cite the first working algorithm for deep learning of internal representations by Ivakhnenko & Lapa (Ukraine, 1965)[DEEP1-2][HIN]. It didn't cite Amari's separate work (1967-68)[GD1-2] on learning internal representations in deep NNs end-to-end through stochastic gradient descent (SGD). Not even the later surveys by the authors [S20][DL3][DLP] nor the "Scientific Background to the Nobel Prize in Physics 2024" mention these origins of deep learning. ([BM] also did not cite relevant prior work by Sherrington & Kirkpatrick [SK75] & Glauber [G63].)
3. The Nobel Committee also lauds Hinton et al.'s 2006 method for layer-wise pretraining of deep NNs (2006) [UN4]. However, this work neither cited the original layer-wise training of deep NNs by Ivakhnenko & Lapa (1965)[DEEP1-2] nor the original work on unsupervised pretraining of deep NNs (1991) [UN0-1][DLP].
4. The "Popular information" says: “At the end of the 1960s, some discouraging theoretical results caused many researchers to suspect that these neural networks would never be of any real use." However, deep learning research was obviously alive and kicking in the 1960s-70s, especially outside of the Anglosphere [DEEP1-2][GD1-3][CNN1][DL1-2][DLP][DLH].
5. Many additional cases of plagiarism and incorrect attribution can be found in the following reference [DLP], which also contains the other references above. One can start with Sec. 3:
[DLP] J. Schmidhuber (2023). How 3 Turing awardees republished key methods and ideas whose creators they failed to credit. Technical Report IDSIA-23-23, Swiss AI Lab IDSIA, 14 Dec 2023. https://t.co/Nz0fjc6kyx
See also the following reference [DLH] for a history of the field:
[DLH] J. Schmidhuber (2022). Annotated History of Modern AI and Deep Learning. Technical Report IDSIA-22-22, IDSIA, Lugano, Switzerland, 2022. Preprint arXiv:2212.11279. https://t.co/Ys0dw5hkF4 (This extends the 2015 award-winning survey https://t.co/7goTtI5Uwv)
Instead of another LLM comparison/test, this time I've tested and compared something very different:
🐺🐦⬛ LLM Prompt Format Comparison/Test: Mixtral 8x7B Instruct with **17** different instruct templates — https://t.co/7FP00aB4f2
Open-Source LLMs vs. ChatGPT:
1. General Capabilities: Llama-2-chat-70B variant exhibits enhanced capabilities in general conversational tasks, surpassing the performance of GPT-3.5-turbo; UltraLlama matches GPT-3.5-turbo’s performance in its proposed benchmark.
2. Agent Capabilities (using tools, self-debugging, following natural language feedback, exploring environment): Lemur-70B-chat surpasses the performance of GPT-3.5-turbo when exploring the environment or following natural language feedback on coding tasks. AgentLlama-70B achieves comparable performance to GPT-3.5-turbo on unseen agent tasks. Gorilla outperforms GPT-4 on writing API calls.
3. Logical Reasoning Capabilities: fine-tuned models (e.g., WizardCoder, WizardMath) and pre-training on higher quality data models (e.g., Lemur-70B-chat, Phi-1, Phi-1.5) show stronger performance than GPT-3.5-turbo.
4. Modeling Long-Context Capabilities: Llama-2-long-chat-70B outperforms GPT-3.5-turbo-16k on ZeroSCROLLS.
5. Application-specific Capabilities:
- query-focused summarization (fine-tuning on training data is better)
- open-ended QA (InstructRetro shows improvement over GPT3)
- medical (MentalLlama-chat-13 and Radiology-Llama-2 outperform ChatGPT)
- generate structured responses (Struc-Bench outperforms ChatGPT)
- generate critiques (Shepherd is almost on-par with ChatGPT)
6. Trust-worthy AI:
- hallucination: during finetuning - improving data quality during fine-tuning; during inference - specific decoding strategies, external knowledge augmentation (Chain-of-Knowledge, LLM-AUGMENTER, Knowledge Solver, CRITIC, Prametric Knowlege Guiding), and multi-agent dialogue.
- safety: GPT-3.5-turbo and GPT-4 models remain at the top for safety evaluations. This is largely attributed to Reinforcement Learning with Human Feedback (RLHF). RL from AI Feedback (RLAIF) could help reduce costs for RLHF.
🔗https://t.co/AQ7UpsL2ev
Thanks to the authors for the great paper! @CaimingXiong@HailinChen3@FangkaiJiao@qcwntu@XingxuanLi@RuochenZhao3@MatRavox@JotyShafiq
Excited to announce DPO has gone multi-modal! New paper out on RLHF for text-to-image diffusion models! We obtain large-scale state of the art results with 70% win rates against Stable Diffusion XL on human evals! Deep dive below 🧵
With some exceptions, the biggest impacts in AI come from people who are experts at both software and machine learning.
Though most people expect the opposite, it’s generally much faster to learn ML than software.
So great software engineers tend to have outsize potential in AI
Check out this blog post by Luis Serrano on sentence similarity! He dives into the importance of embeddings in large language models & explains dot product and cosine similarity. Add some excitement to your tech reads!🚀💻 Read more:
https://t.co/kKYs4lEZOa
Generative Fill, a new superpower integrated throughout Photoshop, launching in beta today.
Powered by Firefly, our generative AI family of models, Photoshop now let’s you summon new objects and augment creations layer by layer. Saves time, increases possibility, and pretty 🤯
AI continues to revolutionize industries, accelerating the proliferation of software in various sectors. Jay Alammar explores the value game in AI, generative AI tech stacks, and where the competitive moats might be. Learn more:
https://t.co/oCxyZbI9vd
Superpower analysis with ChatGPT Code Interpreter.
Here is a CSV with data on superheros. Give me interesting graphs. Then perform a network analysis of the various powers, and provide measures of centrality.
Most interesting: reach some conclusions about what this all means.
Only 18 hours since Midjourney 5.1 is out...
and people are already generating insane images.
My 10 favorite examples so far...
(P.S. I stole this from @_Borriss_. Whatever you do, don't follow him!!!)
I used GPT-4-32K (+ other models) to analyze hundreds of files and explain how @Twitter's open-source algorithm works.
Now, I'm sharing the code I used, so you can do this on ANY Github repo!
Here's the AI's explanation, my approach, and the code for your own use:
Introducing CogLayer: A self-structured, semi-autonomous thinking system that figures out what the human wants to think about then uses each GPT call as a thought in an interconnected thought process. Here’s a demo from @replit demo night. Web app launching soon! (built w replit)
Curious how the RedPajama effort by @togethercompute is progressing and where it stacks up? We evaluated the 7B model they just released 2h ago! Here is how it looks 800B tokens in. (Eval took 16 minutes on 32 A100s.)
Started a list of open-source LLMs with commercial licenses so you can fine-tune your own applications. Contributions welcome! 🙏
https://t.co/olH351Yoir
The original license for MosaicML mpt-7b story writer (with full commercial rights) is *void*, and was officially replaced by CC non-commercial.
It's the right move, credits where it's due! 🤝
See how fast companies move when they see legal problems?
https://t.co/dkz9gFNcqh
Our team at @MosaicML has been working on releasing something special:
We're proud to announce that we are OPEN SOURCING a 7B LLM trained to 1T tokens
The MPT model outperforms ALL other open source models!
Code: https://t.co/sNEKG7gzYY
Blog: https://t.co/dNpfg5gLtR
🧵