Meet Watchman - your fully autonomous 24/7 watcher companion. It is a fully offline system that works on hardware with min 4GB VRAM.
@cognition#AppBuildersPH
If quantum entanglement could arise in neurons, the phenomenon could explain how different parts of the brain work together. Now calculations show how nerve fibres can produce entangled particles https://t.co/VfuE9cgFLZ
With today’s launch of our Llama 3.1 collection of models we’re making history with the largest and most capable open source AI model ever released. 128K context length, multilingual support, and new safety tools. Download 405B and our improved 8B & 70B here. https://t.co/F8cI1bUL8h
Cosine-Similarity of Embeddings may not be always about Similarity? 🤔
📌 This Netflix paper cautions against blindly using cosine similarity and proposes alternatives such as training the model directly with cosine similarity, projecting embeddings back to the original space before applying cosine similarity, or applying normalization/popularity bias reduction before or during training.
It concludes that cosine similarity can be arbitrary and meaningless depending on the regularization used during training.
📌 Experiments on simulated data with known ground-truth item clusters illustrate the large variability in item-item cosine similarities for the first model under different choices of D, compared to the unique solution of the second model.
Grounding is absolutely essential for GenAI applications. Today, we just added new search grounding to the Reader. Now you can simply write a query as 𝗵𝘁𝘁𝗽𝘀://𝘀.𝗷𝗶𝗻𝗮.𝗮𝗶/𝗪𝗵𝗲𝗻+𝘄𝗶𝗹𝗹+𝘁𝗵𝗲+𝗻𝗲𝘅𝘁+𝗦𝗽𝗮𝗰𝗲𝗫+𝗹𝗮𝘂𝗻𝗰𝗵+𝗯𝗲 and it will return you the top-5 search results from the web, each with LLM-friendly text and a URL pointed to the source. With this, devs can easily incorporate latest world knowledge into their LLMs, which is one step closer to improving the factuality of LLMs, making responses more trustworthy and helpful. 🧵
today i learned something very interesting at ICLR, a term called order of evidence.
During RAG, we send the top retrieved documents as evidences and prompting them to LLMs. Normally prompting is based on cosine similarity: top match always being prompted first.
However, different LLMs might have difference "preference" on the order of evidence:
1. ChatGPT prefers top evidence, which is good.
2. GPT4 has no preference on evidence order, which means the similarity score is not being considered, you only need to decide top-K, and all evidences will be equally treated.
3. Surprisingly, Llama2 and PaLM prefers last order evidence, so you need to reverse the rank list then prompt to the LLM :)
https://t.co/fHRR0FGO26
Here’s an early preview of ElevenLabs Music.
All of the songs in this thread were generated from a single text prompt with no edits.
Title: It Started to Sing
Style: “Pop pop-rock, country, top charts song.”
New research from FAIR: Better & Faster Large Language Models via Multi-token Prediction
Research paper ➡️ https://t.co/Q36b6FUjDj
We show that replacing next token prediction tasks with multiple token prediction can result in substantially better code generation performance with the exact same training budget and data — while also increasing inference performance by 3x.
While similar approaches have previously been used in fine-tuning to improve inference speed, this research expands to pre-training for large models, showing notable behaviors and results at these scales.
Early fusion multi-modal models will bring next-step architecture changes. Text scaling-based architectures are suboptimal. Here's a peek at a de-risked transformer arch/recipe that, if scaled right in MM, can achieve 4x efficiency. Paper coming soon (~mid-June).
til, Ilya sutskever gave john carmack this reading list of approx 30 research papers and said, ‘If you really learn all of these, you’ll know 90% of what matters today.’
https://t.co/6eNmrgyq7k
We’re going back 2 back! 🔥 Introducing the first 1M context window @AIatMeta Llama-3 70B to pair with the our Llama-3 8B model that we launched last week on @huggingface. Our 1M context window 70B model landed a perfect score on NIAH and we’re excited about the results that we’re seeing on OpenLLM!
As we continue to give back to the community, we’re actively working with thought leaders like @winglian on evals. As always, thank you to our friends at @CrusoeEnergy for the compute and if you’re interested in joining us on this journey to ensure quality models, send us a DM.
🔗 https://t.co/mKT5K7ZhPm
# CUDA/C++ origins of Deep Learning
Fun fact many people might have heard about the ImageNet / AlexNet moment of 2012, and the deep learning revolution it started.
https://t.co/2xjLWODMOf
What's maybe a bit less known is that the code backing this winning submission to the contest was written from scratch, manually in CUDA/C++ by Alex Krizhevsky. The repo was called cuda-convnet and it was here on Google Code:
https://t.co/ch137VSYZ4
I think Google Code was shut down (?), but I found some forks of it on GitHub now, e.g.:
https://t.co/zYhzdUxoEN
This was among the first high-profile applications of CUDA for Deep Learning, and it is the scale that doing so afforded that allowed this network to get such a strong performance in the ImageNet benchmark. Actually this was a fairly sophisticated multi-GPU application too, and e.g. included model-parallelism, where the two parallel convolution streams were split across two GPUs.
You have to also appreciate that at this time in 2012 (~12 years ago), the majority of deep learning was done in Matlab, on CPU, in toy settings, iterating on all kinds of learning algorithms, architectures and optimization ideas. So it was quite novel and unexpected to see Alex, Ilya and Geoff say: forget all the algorithms work, just take a fairly standard ConvNet, make it very big, train it on a big dataset (ImageNet), and just implement the whole thing in CUDA/C++. And it's in this way that deep learning as a field got a big spark. I recall reading through cuda-convnet around that time like... what is this :S
Now of course, there were already hints of a shift in direction towards scaling, e.g. Matlab had its initial support for GPUs, and much of the work in Andrew Ng's lab at Stanford around this time (where I rotated as a 1st year PhD student) was moving in the direction of GPUs for deep learning at scale, among a number of parallel efforts.
But I just thought it was amusing, while writing all this C/C++ code and CUDA kernels, that it feels a bit like coming back around to that moment, to something that looks a bit like cuda-convnet.
I’m excited to share that I’m working on a new book about building applications with foundation models! AI Engineering builds upon Machine Learning Systems Design, but with a focus on large scale, ready made models.
The book covers:
- The new AI stack (e.g. how it differs from traditional ML engineering)
- Different approaches to evaluate open-ended systems
- Dataset engineering
- Prompt engineering, RAG, agents
- Finetuning
- Compute infrastructure, including how to mitigate latency and cost
AI Engineering is scheduled for late 2024. The first 3 chapters are available on the O'Reilly platform: https://t.co/8YrSmUH9qw
I’ve learned a lot during the research and writing process for this book. I hope you’ll find the learnings useful. Feedback is much appreciated!