I got to go back to my college and just talk with current students. I feel like I was just there, but I did graduate 7 years ago 👴🏿 And then I signed some of their iPads 😅
Oh, and it was recorded if you want to watch: https://t.co/ZXCekfg98N
Our work was just accepted at NeurIPS 26! As the paper promises, the benchmark has grown to 2622 problems now! Would love it if more teams work on this hill-climb.
https://t.co/108vKgpRJk
Just found this benchmark from Balaji Rao et al: s2n-bignum-bench, it tasks agents with writing proofs in HOL Light for crypto assembly code from Amazon's s2n-bignum repo which they use in production.
Codex 5.3 xhigh gets just 6.3% so this could be good to work on.
Just found this benchmark from Balaji Rao et al: s2n-bignum-bench, it tasks agents with writing proofs in HOL Light for crypto assembly code from Amazon's s2n-bignum repo which they use in production.
Codex 5.3 xhigh gets just 6.3% so this could be good to work on.
It’s year 2024, and n-gram LMs are making a comeback!!
We develop infini-gram, an engine that efficiently processes n-gram queries with unbounded n and trillion-token corpora. It takes merely 20 milliseconds to count the frequency of an arbitrarily long n-gram in RedPajama (1.4T tokens), while also retrieving all of its occurrence positions in the corpus. That’s right, your query can be arbitrarily long and appear arbitrarily frequently, and it takes 0.02 seconds to process!
We created an infini-gram index on RedPajama, which makes it the biggest n-gram LM ever built to date.
Using larger n improves the predictive power of n-gram LMs. Existing n-gram LMs are mostly built with n<=5, which limits their capability to predict the next token (shown in image below). If we increase to n=16, we still get a non-zero n-gram count in RedPajama, but now the model makes a correct prediction. This is why we generalize n-gram LMs to unbounded n, i.e., an ∞-gram LM. ∞-gram uses the longest context possible (right before the count becomes zero), and our infini-gram engine perfectly supports this.
Paper: https://t.co/nWIEi9Brsb
@huggingface 🤗 Demo: https://t.co/B7JeEsqCdQ
Code & pre-built indexes: https://t.co/JeISXCcTID (coming soon)
🧵(1/n)
Evaluation Metrics for RAG Reranking
How do we measure the performance of our post-retrieval rerankers?
Use two metrics:
• MRR
• nDCG
MRR
Mean Reciprocal Rank measures how well the reranking model finds first relevant document.
MRR is useful when the most important factor is finding the most relevant document quickly.
It's like asking: "How quickly did we hit gold?"
nDCG
Normalized Discounted Cumulative Gain examines the whole list of ranked documents.
It considers both the relevance and the order of each document.
nDCG is useful for understanding the overall quality of the reranked list.
It's like asking: "How rich is the entire gold mine?"
Friends don't let friends make bad charts!
Chenxin Li, pulled together a lot of great advice for data visualization, with clear "do this, not that" examples for each item.
Here are a few of my favorites, see the link below for more.