Pop quiz: Why must we reduce learning rate as we scale transformers?
@jmgilmer has the answer - and the solution!
Attention logits can grow uncontrollably, destabilizing training. The fix: Layer norm on queries and keys!
(Follow Justin and 👀 for upcoming paper on this topic!)
We've just released the first version of our Deep Learning Tuning Playbook! This is our attempt to distill our process for actually getting good results with deep learning. We emphasize hyperparameter tuning since it has been a large pain point. https://t.co/PjeJVWeOzS
I tried HSBC chat. Here's what they said:
I’ll transfer you to an agent now. It could take up to 6 hours to get connected because our agents are helping other customers like you. Please feel free to log out of Online Banking, and check back later for our response. Thank you.
AI is eating software. I’ve said this for a while. Why? Traditional software never improves. AI enables ‘smart’ software to learn & improve constantly. The world runs on software and AI is changing everything. Yet few people see or understand the massive AI wave on the horizon.
In the NYT today, Cade Metz implies that I left Google so that I could criticize Google. Actually, I left so that I could talk about the dangers of AI without considering how this impacts Google. Google has acted very responsibly.
Dishonest CBC headline:
"Canada's AI pioneer Geoffrey Hinton says AI could wipe out humans. In the meantime, there's money to be made".
The second sentence was said by a journalist, not me, but you wouldn't know that.
US sanctions are driving Chinese firms, like Huawei, Alibaba, Baidu, to seek ways to advance AI without cutting-edge chips. We dug through open-source research papers and spoke with employees & analysts to understand what they're doing. w/ @raffaelehuang https://t.co/361AOgauKH
Two years ago I discovered scholarship arguing that AI is creating a new colonial world order. The idea haunted me & pushed me to investigate. Today I begin to release my findings: a culmination of reporting across five continents. Here’s my introduction. https://t.co/NlPAQDVpuR
@zSafwan@omarsar0 Saved! Here's the compiled thread: https://t.co/ebQ8hKbOjI
🪄 AI-generated summary:
"This thread is about a comprehensive prompt engineering guide used by thousands of AI developers and researchers working with LLMs. It has been translated into Chinese and...
Added a few more translations. Spanish, Turkish, and Italian. 7 total languages are now available for the Prompt Engineering Guide.
We also added a more detailed explanation to ReAct prompting.
More coming soon
https://t.co/Dbm9hWMix3
https://t.co/24k6YQrMcz
Prompt Engineering Guide (20K⭐️)
We started with basic prompt examples and have expanded to a comprehensive prompt engineering guide used by thousands of AI developers and researchers working with LLMs.
- Now over 200K+ learners
- Chinese & Japanese translations are now available
- GPT-4 & ChatGPT guides and notebooks
- Collection of all the latest tools and papers on prompt engineering
- Added LLM collection
- Papers explanations in progress
... and much more
We aim to build the ultimate resource to learn how to work and build with LLMs. A lot more to come.
Support and contributions are welcome!
web: https://t.co/o4KzoHf52W
repo: https://t.co/S5YqgigNlb
DeepSpeed Chat
Impressive open-source effort by Microsoft!
DeepSpeed Chat offers an end-to-end RLHF pipeline to train ChatGPT-like models. This is the missing piece from other efforts like Alpaca and Vicuna. The RLHF pipeline is replicated from the InstructGPT paper.
The other challenges are cost and efficiency. DeepSpeed Chat aims to make this process more accessible and affordable through a unified hybrid engine (DeepSpeed-HE) for RLHF. For example, DeepSpeed-HE can train an OPT-66B model in 2.1 days for $1620.
With further scalability (e.g., using multi-node multi-GPU systems), DeepSpeed-HE can train an OPT-13B model in 1.25 hours for $320 and an OPT-175B model in under a day for $5120. That's a big deal!
More info here: https://t.co/rmlM3dtu5F
TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion Synthesis
extensive experiments on the KIT-ML and HumanML3D datasets show that TMR outperforms the prior work by a significant margin, for example reducing the median rank from 54 to 19
abs: https://t.co/B4QbGLlmnH
project page: https://t.co/KH4w0442YP
@Gradio demo: https://t.co/QHyF5Knxqm
China wants peace and the US wants wars — this is the situation now due to structural characteristics of the two economic systems.
China wants peace because its economy is based on production, infrastructure and trade.
Peace means Chinese trains and ships can move around freely importing and exporting goods.
Peace means Chinese engineers can build bridges and dams.
Peace means Alibaba can get customers from Syria and Iraq.
On the other hand, the US engine is running on dollar vaporware, which can be sustained only by a Ponzi scheme, which needs a continuous supply of buyers.
Peaceful and prosperous countries are not going to buy US debt. They are going to trade in local currencies. They are going to create trade blocs like BRICS.
A peaceful and prosperous Middle East will start selling oil for Yuan. Saudis will use oil profits to build futuristic cities (Neom) rather than buy US weapons.
Peaceful and prosperous countries will invest in colleges, R&D, manufacturing and technology. Soon they will be competing with US companies. Look at China over the last 40 years.
What happens next in the world will depend on the intelligence of human beings. Unfortunately, maintaining peace is 100x more difficult than starting wars.