@arxiv Hi @arxiv,
Paper been on hold for 2 weeks🥹
1. https://t.co/po1VWrmYbN
2. Ye Liu
3. Procedural Memory Distillation: Online Reflection for Self-Improving Language Models
Thanks for the great of help!
📢 New Preprint from @raghavlite on Multimodal Contrastive Learning: Breaking the Batch Barrier (B3) 📢
TL;DR: Smart batch mining based on community detection achieves state of the art on the MMEB benchmark.
Preprint: https://t.co/BU85rYwDBp
Code: https://t.co/bo2xLRBG7Q
@HongLiu9903 Hi @HongLiu9903, thanks for interest in our model. The performance mismatch was due to the <eos> token not being added. We have fixed this by setting add_eos_token to True in the tokenizer_config.json. You can quickly verify that by checking the model performance on Apps.
🚨🚨🚨Just released!🚨🚨🚨
🚀Introducing the Salesforce Code Embedding Model Family (SFR-Embedding-Code), ranked #1 on CoIR Benchmark! 🚀
Available in 2 sizes: 2B, 400M. Key Highlights:
1️⃣ 2B Model: Achieves #1 on CoIR.
2️⃣400M Model: Best-performing model under 0.5B parameters.
3️⃣ Multi-lingual, multi-task unified training framework for code retrieval
4️⃣ Supports 12 programming languages, including Python, Java, C++, JavaScript, C#, and more!
🧑💻✨Empower your next AI Coding Agent with the best code embedding models! 🧑💻✨
Join us in advancing #AccurateAI:
📎Paper: https://t.co/jbQpo8tesX
🤗400M Model: https://t.co/as8duCE6w2
🤗2B Model: https://t.co/tOIDOz7kLZ
#CodeAI #MLResearch #SOTA #OpenScience @Salesforce
Big thanks to our research team for SFR-Embedding Code:
Ye Liu @YeLiu918
Rui Meng @RuiMeng_
Shafiq Joty @JotyShafiq
Silvio Savarese @silviocinguetta
Yingbo Zhou @yingbozhou_ai
Caiming Xiong @CaimingXiong
Semih Yavuz @semih__yavuz
🎆I am pleased to announce the release of the latest version of the Salesforce Embedding Model (SFR-embedding-v2), which has reclaimed the top-1 position on the MTEB benchmark.
✨ Key Highlights:
🥇 Achieved the distinction of being the second model to surpass a 70+ performance score on MTEB.
🔧 New multi-stage training recipe to enhance multitasking capabilities.
📊Significant improvements in classification and clustering tasks, while maintaining strong performance in retrieval and other areas. 💪
https://t.co/6Xoe88zCF2
🚀 Introducing our latest breakthrough: the SFR-embedding model, a new champion on the MTEB benchmark! 🥇
But why does it excel?🤔 Read here: https://t.co/UVOGKbDwDc
Experience the future of AI with SFR-Embedding-Mistral!⭐
Introducing 🔥SFR-Embedding-Mistral🔥 has clinched the #1 spot on the MTEB leaderboard!🥇
Key highlights:
Retrieval and Reranking: New SoTA.
Retrieval Score: a massive leap from 56.9 to 59
Clustering Tasks: Achieved a +1.4 absolute improvement
https://t.co/KoWdFvXKl4
We introduce 🔥XGen-7B 🔥, a new 7B LLM trained on up to 8K sequence length for 1.5T tokens. Achieves better or comparable results with MPT, Falcon, LLaMA, Redpajama, and OpenLLaMA in the text and code tasks.
🔗Blog: https://t.co/IrZjQvHPuw
🔗Code: https://t.co/QfsDxCZYDj