Checkout our new Muse Spark 1.1 (available with an API)!
We have made huge leaps in its agentic coding capabilities.
Excited to see how the world builds on it!
1/ muse spark 1.1 is an industry-competitive agentic and coding model. across many agentic evals it rivals gpt-5.5 and opus-4.8.
available now through the new meta model api and in meta ai. 🧵
Another release this week🔥
Muse Spark 1.1 is a considerable upgrade in capability, especially for agentic tasks. Plus… It’s available via API!
We see this model as a workhorse: fast and very reasonable priced, with strong all-round capability.
(1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
1/ muse spark 1.1 is an industry-competitive agentic and coding model. across many agentic evals it rivals gpt-5.5 and opus-4.8.
available now through the new meta model api and in meta ai. 🧵
1/ releasing muse image today — the first image generation model from MSL. it's agentic: pairs with muse spark to reason through your prompt, search the web, and plan before it generates. people get what they meant on the first try. live now in the Meta AI app.
Excited to release PrefEval (ICLR '25 Oral), a benchmark for evaluating LLMs’ ability to infer, memorize, and adhere to user preferences in long-context conversations!
⚠️We find that cutting-edge LLMs struggle to follow user preferences—even in short contexts. This isn't just about long-context ability—but also a lack of proactiveness in following preferences.
💭 LLMs struggle with implicit preferences revealed through conversation, requiring more reasoning.
🚀 SFT on our benchmark greatly improves performance w/ enhanced attention to preference regions!
🧩 What's in PrefEval?
- 3,000 user preference & query pairs
- 3 Implicit & explicit preference forms
- 20 conversational topics
- Evaluated up to 100K context w/ various baselines
- Flexible task forms: MCQ & generation
- Extensive Error Type Analysis
🔗 Learn more:
Website: https://t.co/nwSEbefqFf
Code & Data: https://t.co/5h1SphruHN
Paper: https://t.co/J3bejwsIpq
Huge thanks to my mentors during my internship at Amazon @linkaixi @Mingyi552237, Yang Liu, @hdevamanyu!
If you're interested in discussing how retrieval-augmented generation and in-context learning can be used to create safer chatbots in #NLProc, come see our poster at #EMNLP2023 tomorrow (Dec. 9) from 11:00 AM - 12:30 PM in the East Foyer!
📣📣 #EMNLP2023@Tahaksu1 presents his joint work with @AmazonScience today! They introduce a framework bridging the compositinal task gap between closed & open source language models. Find him at the venue in the poster session PS3, 16:00 today! #NLP#LanguageModels
Exciting workshop coming your way @ SIGDIAL/INLG (https://t.co/DVWDXAGuUA) ... on the intersection of controllability and LLMs.
Submit papers by July 7th!
#TamingLLM#INLG2023@sigdial@inlgmeeting
We have extended our deadline to July 7, if you missed (or will miss) the EMNLP deadline, welcome to submit to the 1st Workshop on Taming Large Language Models!
Can retrieval-augmented generation reduce toxicity in #NLProc? We find retrieving even from a pool of 10 demonstrations can reduce toxicity by ~14% and when using a larger pool of demonstrations, we outperform BlenderBot3 in human evaluation!
📎: https://t.co/Ar6s3euLXS
[1/7]
📢 Call for Student Volunteers 📢
#ACL2023NLP is looking for student volunteers to help us with conference activities (both online and in-person).
Checkout the call https://t.co/7Om9e92zg5 for more details.
#ACL2023Toronto#NLProc
📢 Call for Student Volunteers 📢
#ACL2023NLP is looking for student volunteers to help us with conference activities (both online and in-person).
Checkout the call https://t.co/7Om9e92zg5 for more details.
#ACL2023Toronto#NLProc
Our KILM paper was accepted at ACL 2023! In this paper with @Tea_XuYan, @hdevamanyu, Di Jin, @AishwaryaPadma4, Yang Liu, and @dilekhakkanitur we show how atomic pieces of knowledge could be injected into a pre-trained language model. https://t.co/PBzz8v33dc @AmazonScience
The #ACL2023NLP committee is hard at work to finalize the submission decisions, which will be tentatively out by the end of May 1st (Anywhere on earth)....19 hours to go... Fingers crossed!
#ACL2023NLP#ACL2023Toronto