"Unintelligent people are more easily misled by other people, intelligent people are more easily misled by themselves. They’re better at convincing themselves of things they want to believe rather than things that are actually true."
Great article! https://t.co/jca5HOwV2C
Today, we announced that we’ve gotten dictionary learning working on Sonnet, extracting millions of features from one of the best models in the world.
This is the first time this has been successfully done on a frontier model.
I wanted to share some highlights 🧵
Every saturday morning for the last 6 years, I watch this obscure video.
It's Jeff Bezos talking about leadership, but really, it's the most succinct blueprint for how to achieve greatness I have ever found.
You should watch it often, but I summarized it here: (thread, sry)
Enshitification of platforms:
Here is how platforms die: first, they are good to their users; then they abuse their users to make things better for their business customers; finally, they abuse those business customers to claw back all the value for them. https://t.co/bOjMFKMvxm
In case I need to discuss that again in the future, leaving it here for further reference. 😄
Death By a Thousand Microservices:
The software industry is learning once again that complexity kills https://t.co/WBGMAGFBny
Given the popularity of retrieval augmented generation (RAG) for LLMs, one question I’m constantly asked is: What model should I use to embed my data for RAG? This question has a simple answer that I use for (almost) all applications…
TL;DR: Sentence BERT (sBERT) is an encoder-only transformer, based upon the BERT model, that has been optimized for vector/semantic search. sBERT models (and tons of variants) are openly available via the Sentence Transformers library and are easy/cheap to finetune over your own data to get the best possible results in any search application (including RAG!).
What is BERT? To understand sBERT, we first need to understand BERT. Put simply, BERT is an encoder-only language model that is pretrained using an infilling/Cloze objective (as opposed to next token prediction for generative LLMs). This model, which is small and efficient compared to most generative LLMs (110-340 million parameters), can be cheaply fine-tuned to solve many different token/sequence-level classification tasks with a small number of training examples.
BERT for search. When provided two pieces of text as input (i.e., cross-encoder setup), BERT can be fine-tuned to accurately predict their level of similarity. However, performing search in this manner would require exhaustively computing similarity for every document we want to consider in our search application, which is incredibly inefficient. Instead, we need a bi-encoder that can separately embed pieces of text, allowing us to find similar documents using efficient vector search (i.e., approximate nearest neighbors) over a large number of documents! Unfortunately, vanilla BERT functions poorly in this setup—its textual embeddings are not semantically meaningful.
Introducing sBERT. Sentence BERT (sBERT) takes a pretrained BERT model and further trains it to produce semantically meaningful embeddings. To do this, we can use a siamese/triplet network structure, where we separately embed two (or three) pieces of text with BERT, producing an embedding for each piece of text. Then, given pairs (or triplets) of text that are either similar or not similar, we can train sBERT to accurately predict textual similarity from these embeddings using an added regression or classification head. This approach is cheap/efficient, drastically improves the semantic meaningfulness of BERT embeddings, and produces a high-performing (BERT-based) bi-encoder that we can use for vector search.
RAG with sBERT. A variety of open-source sBERT models are available via the Sentence Transformers library. To use these models for RAG, we simply need to split our documents into chunks, embed each chunk with sBERT, then index all of these chunks for search in a vector database. From here, we can perform efficient vector search (ideally, hybrid search that also includes lexical components) over these chunks to retrieve relevant data at inference time for RAG.
Continual improvement. To make this approach more effective, we should capture good/bad examples of data retrieved for RAG within our application. E.g., we can collect pairs of user messages with chunks of text that should or should not have been retrieved. Given this data, we can finetune sBERT to improve its effectiveness for our application. By finetuning sBERT over the data that we collect, the model that we use for RAG will continue to get better over time.
@cschleiden@lufthansa Couldn't agree more. I regularly think that their apps and website were quite good 10+ years ago (in our Consulting days) and are getting worse and worse - even though the use case hasn't changed much over the years.
Tony Fadell (co-creator of the iPod and iPhone) on opinion-based decisions
“When you make the first version of anything—something revolutionary—there are a lot of opinion-based decisions… And when you have those opinions, and you’re trying to work with a team to implement those decisions, you have to really tell the ‘why’ of those decisions. That way everyone can feel like they’re a part of those decisions and understand the tradeoffs... A lot of times, people want a data-driven decision, but with v1s you don’t have data... If you look at most companies that are paralyzed and cannot make new innovations and new products, it’s because they’re trying to turn opinion-based decisions into data-driven decisions so that they don’t lose their jobs”
— Tony Fadell
Testing out @HeyGen_Official translation on French and German. I don’t speak either language so let me know if it sounds natural if you do.
I hope if you pay you can turn off the color correction.
It didn’t work on my phone so I had to upload on my pc.
https://t.co/FMJp9sJEBI