Introducing Sine your research copilot
Input your topic, and sine will cover different aspects of the topic in-depth.
Join waitlist here https://t.co/njrAsshvpw
@_buildspace
https://t.co/zmFYzR7BKr
@hex4_makes@_buildspace Thanks for your question ! I plan to use open-source model, e.g. llama3 or free apis first, e.g. groq. It definitely supports search engine, and it is a very good idea to support uploaded files !
i want to understand and learning many things deeply and scientifically, and before that is not easy even if we have search engine. By using LLMs, we could create the best content for what you are interested in.
Are you interested ?
@_buildspace
๐ We're thrilled to have successfully hosted the Local-First Shanghai Meetup! A huge thanks to everyone who attended and contributed. Here are some highlights from our amazing speakers. Check out the photos below! ๐ธ
@cwolferesearch Thanks for the thread ! Especially about `invest into a really good evaluation framework`, could not agree more, could you share more practices on this? thanks :)
The memory in Transformers grows linearly with the sequence length at inference time.
In SSMs it is constant, but often at the expense of performance.
We introduce Dynamic Memory Compression (DMC) where we retrofit LLMs to compress their KV cache while preserving performance and vastly surpassing GQA!
The throughput of Llama 2 7B/13B/70B increases by up to 370% on a H100 GPU.
Paper: https://t.co/uUnh4g92VX
Code and models are coming soon!
@AdrianLancucki@PontiEdoardo@nvidia@EdinburghNLP
very interesting blog over long llm and rag. I am wondering if the long and shared computed kv cache worth offloading to RAM/disk for later reloading ? e.g. learning specific domain knowledge, the knowledge is a very long text and has been used computed, it is a waste to discard
Towards Long Context RAG ๐ฎ
Gemini 1.5 Pro is impressive. Naturally this begs this question of what RAG will look like in a long-context LLM future - which techniques will disappear and which will remain?
We did a deep dive into Gemini, and consolidated our thinking about long-context LLM benefits, challenges, and new architectures ๐ก
โ Long-context LLMs will help alleviate the need to do precise chunking and retrieval, and RAG over small sets of documents
โ๏ธ Long-context LLMs still donโt resolve the issue of RAG over big knowledge bases (present in most organizations/enterprises)
We propose some new RAG architectures that combine the capabilities of long-context LLMs with even bigger external knowledge bases
1๏ธโฃ Small-to-Big Retrieval over Documents
2๏ธโฃ Intelligent Routing for Latency/Cost Tradeoffs
3๏ธโฃ Retrieval-Augmented KV Caching
A big shoutout to @Francis_YAO_ for his feedback and thoughts on the current state-of-the-art long-context research.
Our core mission in @llama_index is to enable developers to build LLM-powered apps over their data, whether thatโs using RAG or any new framework that comes along. Weโre excited to continue innovating on the techniques that will help you take advantage of these long-context models.
Blog: https://t.co/FRR45Pxh7D
the story of how my newsletter with 100 subscribers ended up bringing me from a small town, to sf.
+ landed me my first $25,000 to start this co.
and along the way, introduced me to @FurqanR + @ShaanVP -- two people that changed my life and trajectory.
how to ask for help:
feeling sluggish or unmotivated to finish the tasks and end day w/ regrets? we have all been there, and i find a helpful method and make an app:
find a DuDu Buddy who is someone just like you to take care each other's daily tasks progress under some rules
cc @_buildspace