1/ Can Large Language Models (LLMs) truly reason? Or are they just sophisticated pattern matchers? In our latest preprint, we explore this key question through a large-scale study of both open-source like Llama, Phi, Gemma, and Mistral and leading closed models, including the recent OpenAI GPT-4o and o1-series.
https://t.co/2tv8Pp9MSz
Work done with @i_mirzadeh, @KeivanAlizadeh2, Hooman Shahrokhi, Samy Bengio, @OncelTuzel.
#LLM #Reasoning #Mathematics #AGI #Research #Apple
I built a free Cursor course (using Cursor!)
It’s my birthday today but I hate gifts so here’s one for you instead.
Learn how to get started coding with Cursor even if you’ve never coded in your life.
Check link in the replies and let me know what you wanna learn!
Here’s a list of my favorite blogs to stay updated on everything around building things with LLMs and the generative AI space in general:
(No specific order)
The two hottest topics in AI right now are RAG and Agents.
Every top company in the AI community is laser-focused on making AI Agents a reality.
I'm working on a couple of YouTube videos to teach you how Agents work and how to build one. In the meantime, let me show you this:
@langflow_ai is an open-source visual framework for building RAG applications and agentic workflows.
It's super-popular, with over 15k stars on GitHub:
https://t.co/F91niosr8x
What's cool about LangFlow is that anyone can use it to build langchain-based applications:
• Visual interface
• Drag-and-drop
• Reusable components
You can write all the code you want, but a visual interface with predefined components will save you a ton of time.
Something cool:
@DataStax just acquired LangFlow. They'll be backing the tool financially to speed up innovation.
By the end of the year, developing AI applications will be 100x better than today.
Uploaded quantized 4bit models (4GB size) for @MistralAI's 32K new model courtesy of alpindale: https://t.co/OSSwAe94iG
Plus get 2x faster 70% less VRAM QLoRA finetuning with @UnslothAI! & Colab for Mistral v2: https://t.co/X7BN2P3STB helpful for the @cerebral_valley hackathon!
🚀 🌐 Build your own video generation model like #Sora! Experience the power of replication without the price tag! Open-Sora delivers a low-cost implementation of Sora, cutting costs by a staggering 46%. Expand your sequences to nearly a million with this innovative open-source project. #OpenSora #ColossalAI 🌟 Learn more here: https://t.co/12Fl4ZysIG
I just finished reading
"Fuck You, Show Me The Prompt."
by @HamelHusain
This is a M-U-S-T.
Every single word is pure gold.
Most LLM frameworks are not only unnecessary but are actively crippling you and slowing you.
Their only value is their prompts
https://t.co/nKpVkppaYi
Thank you for already over 12K views in the first 3 days of publishing my latest article on DSPy! 🎉
In case you haven’t read it yet, here’s what you’ve missed:
DSPy is a TRENDING framework developed by Stanford NLP researchers to help you build LLM-based applications. It aims to tackle the fragility problem of developing LLM-based applications by prioritizing programming over prompting.
Below, you can see the core concepts DSPy introduces:
🧩 Hand-written prompts-> replaced by signatures
🧩 Prompting techniques -> replaced by modules
🧩 Manual prompt engineering -> automated with teleprompters & DSPy Compiler
Read more on @TDataScience: https://t.co/AnrpyPaPku
Instead of reading people complain on twitter about models being "over RLHFd", "hand tuned refusing" and "intentionally woke-ized", go watch this talk from @johnschulman2 on the actual challenges of doing this https://t.co/VS9CmPyo5c
Build a full-stack RAG-powered Restaurant Menu 🧑🍳🥗
Excited to feature a @weights_biases blog post showing you how to build a full-stack chatbot that can answer any questions about a restaurant menu, with @llama_index + in-built logging using @weights_biases Weave for app usage.
Check it out: https://t.co/o7Z8D10Mua
In this new 90 minute lecture, I show how to pretrain a 3B LLM from scratch. No edits. No detail skipped.
Companies want you to believe pretraining models is super hard and costly. With the right tools, it's not.
- We start by tuning the model on a cheap A10G.
- Then we scale to 4 A10G to speed up training by 2x.
- We handle OOM errors and tune hyperparameters.
- We end with scaling to 8 H100s (1 machine) to speed up by another 4x.
- Tune again to scale from 1B to 3B params.
- Then scale to 16 H100s (multi-node).
At the end of this video you'll develop a good development workflow for pretraining LLMs.
https://t.co/XXhUHhp3bZ
Awesome new MLX package - mlx_embedding_models.
Generate sentence embeddings easy and fast on your laptop.
pip install mlx-embedding-models
Code: https://t.co/e4MUedYalt
Quick-start:
Advanced QA over a lot of Tabular Data (combine text-to-SQL with RAG) 📊🪄
Our brand-new mini course 🧑🏫 is a comprehensive overview of how you can build simple-to-advanced query pipelines from scratch, by composing components into complex DAGs. Presenting this in three levels:
1️⃣ Basic text-to-pandas / SQL
2️⃣ Query-time table retrieval in text-to-SQL prompt
3️⃣ Query-time row retrieval in text-to-SQL prompt
Steps 2 and 3 introduce RAG concepts by vectorizing the tables and rows for few-shot example selection.
Adding on these layers ensures that your pipeline can scale to more tables (table retrieval), and that your queries are less prone to failure with the right examples.
Uses WikiTableQuestions as a dataset (@IcePasupat et al.)
Logged with @ArizePhoenix tracing (works with any of our observability partners).
Video (part 2 of our advanced RAG orchestration series): https://t.co/7MxTuPc3dY
Colab: https://t.co/DcC3fNjO0V
Source docs:
SQL: https://t.co/blb1Xfs0wN
Pandas: https://t.co/us6PEUfk49
Introducing a Short Course Series on Advanced RAG Orchestration 🪄🤖
As an AI engineer, it can be daunting to dive into how to build high-quality, advanced RAG yourself - there’s literally hundreds of options at every stage of the pipeline.
Easily stitch together custom modules into DAGs over your data, with observability baked in (here we show @ArizePhoenix 🔬).
Check out our first course in the series on query pipelines. We show you how to compose basic workflows like prompt chaining, output parsing and streaming to advanced RAG with query rewriting, retrieval, and more.
YouTube: https://t.co/6hjuoADLcv
Reranking with RAGatouille library using ColBERT - quick nice use case 💡
📌 Re-rank documents retrieved by another retriever, such as your existing RAG pipeline
📌 The example in the code will use the `RAGPretrainedModel` wrapper class which Allows you to load a pretrained model from disk or from the hub, build or query an index.
------
📌 In the last line of the attached code(in the image) after running the `rerank()` method - The relevant extract is now all the way at the top of the results, ready to be passed to the rest of your pipeline!
So why not just use `rerank()` on the whole index if it's so good? Because that is NOT efficient.
📌 ColBERT is an extremely fast querier, but it needs to have an index built to do so. When you're using ColBERT to rerank documents, it's doing it index-free, which means it needs to encode all your documents and queries, and perform the comparison on the fly. This is fine for a handful of document on CPU or a few hundreds on GPU, but it's going to take exponentially longer as you add more documents!
📌 Re-ranking the results of another retrieval method is a good compromise: it allows you to leverage ColBERT's power without having to modify the rest of your pipeline, just increase the `k` value of your retriever and let ColBERT rescore them!
Ten months ago, we launched the Vesuvius Challenge to solve the ancient problem of the Herculaneum Papyri, a library of scrolls that were flash-fried by the eruption of Mount Vesuvius in 79 AD.
Today we are overjoyed to announce that our crazy project has succeeded. After 2000 years, we can finally read the scrolls:
This image was produced by @Youssef_M_Nader, @LukeFarritor, and @JuliSchillij, who have now won the Vesuvius Challenge Grand Prize of $700,000. Congratulations!!
These fifteen columns come from the very end of the first scroll we have been able to read and contain new text from the ancient world that has never been seen before. The author – probably Epicurean philosopher Philodemus – writes here about music, food, and how to enjoy life's pleasures. In the closing section, he throws shade at unnamed ideological adversaries – perhaps the stoics? – who "have nothing to say about pleasure, either in general or in particular."
This year, the Vesuvius Challenge continues. The text that we revealed so far represents just 5% of one scroll.
In 2024, our goal is to from reading a few passages of text to entire scrolls, and we're announcing a new $100,000 grand prize for the first team that is able to read at least 90% of all four scrolls that we have scanned.
The scrolls stored in Naples that remain to be read represent more than 16 megabytes of ancient text. But the villa where the scrolls were found was only partially excavated, and scholars tell us that there may be thousands more scrolls underground. Our hope is that the success of the Vesuvius Challenge catalyzes the excavation of the villa, that the main library is discovered, and that whatever we find there rewrites history and inspires all of us.
It's been a great joy to work on this strange and amazing project. Thanks to Brent Seales for laying the foundation for this work over so many years, thanks to the friends and Twitter users whose donations powered our effort, and thanks to the many contestants whose contributions have made the Vesuvius Challenge successful!
Read more in our announcement: https://t.co/rUlrdGXBMs
Anyone got any leads on good LLM prompts for RAG with citations?
I have a bunch of embedded content, I want a prompt I can use with it to show an answer to a question that includes little [1] citation links that link to the content that was used to answer a question
I used to find writing CUDA code rather terrifying. But then I discovered a couple of tricks that actually make it quite accessible.
In this video I introduce CUDA in a way that will be accessible to Python folks, & I even show how to do it all in Colab!
https://t.co/WGXXctbalv
There's a new promising method for finetuning LLMs without modifying their weights called
proxy-tuning (by Liu et al. https://t.co/3PjF0NtlOM).
How does it work? It's a simple decoding-time method where you modify the logits of the target LLM. In particular, you compute the logits' difference between a smaller base and finetuning model, then apply the difference to the target model's logits.
More concretely, suppose the goal is to improve a large target model (M1).
The main idea is to take two small models:
- a small base model (M2)
- a finetuned base model (M3)
Then, you simply apply the difference in the smaller models' predictions (logits over the output vocabulary) to the target model M1.
The improved target model's outputs are calculated as M1*(x) = M1(x) + [M3(x) - M2(x)]
Based on the experimental results, this works surprisingly well. The authors tested this on
A. instruction-tuning
B. domain adaptation
C. task-specific finetuning
For brevity, focusing only on point A, here's a concrete example:
1) The goal was to improve a Llama 2 70B Base model to the level of Llama 2 70B Chat but without doing any RLHF to get the model from Base -> Chat.
2) They took a 10x smaller Llama 2 7B model and instruction-finetuned it.
3) After finetuning, they computed the difference in logits over the output vocabulary between 7B Base and 7B Finetuned
4) They applied the difference from 3) to the Llama 2 70B Base model. This pushed the 70B Base model's performance pretty close to 70B Chat.
The only caveat of this method is, of course, that your smaller models have to be trained on the same vocabulary as the larger model. Theoretically, if one knew the GPT-4 vocabulary and had access to its logit outputs, one could create new specialized GPT-4 models with this approach.
hi! Open Interpreter 0.2.0—The New Computer Update—is out today.
everything's new.
- OS Mode lets vision models operate your computer
- We included a new model for precise GUI control
- We're launching a Computer API for LLMs
↓