🌟New Benchmark! 🌟
Do you work on RAG? Are you interested in Multi-Turn conversations? Very excited to share the new MTRAG benchmark we have released!
Data: https://t.co/KtJQgtB5Uj
Paper: https://t.co/QNRccmrEV5
@danish_c@kpfadnis@chulaka_g@vpshah95@lucian_popa_us
🎉Today, we're pleased to announce the release of the Granite 3.0 model family, the latest open-licensed, general purpose LLMs from @IBM 🎉
These have been a labor of love for my team at @IBMResearch, working closely with a host of collaborators across the company. We're excited to share them with you.
Some key details:
📏 available in 4 sizes, ranging from 400m to 8B active parameters, with both dense transformer (2B and 8B) and mixture-of-experts (3B-A800M and 1B-A400M) versions.
💪 strong performance across a range of tasks (see figure below), standing toe-to-toe with the best open models in their weight class.
⚡️fast and light, designed to run anywhere, from GPU servers to CPUs and edge devices. We have also released a companion speculator model—Granite Accelerator—which can more than double the speed of the 8B model.
🏦 built with business in mind, with a permissive Apache 2.0 open source license, and special focus on enterprise tasks. Multlingual, coding task-enabled, and with built-in function-calling support for agentic patterns
🔎 comprehensive transparency—our technical report carefully details the data sources that went into our models, along with details on how we trained them.
🏰 Companion Granite Guardian models are available in 2B and 8B sizes that achieve state-of-the-art input and output guardrail performance
🌱 Built using 100% renewable energy
IBM presents Granite 3.0, lightweight foundation models ranging from 400 million to 8B parameters.
Supports coding, RAG, reasoning, and function calling, focusing on enterprise use cases, including on-premise and on-device settings.
The technical report contains detailed discussions on how they collect synthetic datasets for code, reasoning, RAG, tool use, and more.
"Granite 3.0 language models demonstrate strong performance across a battery of academic benchmarks for language understanding, reasoning, coding, function calling, and safety"
The release includes pre-trained and post-trained versions of all Granite 3.0 models under a permissive Apache 2.0 license.
Very strong release by the Granite Team at IBM.
Granite 3.0 is our latest update for the IBM foundation models. The 8B and 2B models outperform strong competitors with similar sizes. The 1B and 3B MoE use only 400M and 800M active parameters to target the on-device use cases. Our technical report provides all the details you need to train a state-of-the-art 8B model from scratch!
https://t.co/Q58KSz3UaL
Can you make an existing pretrained LM behave like a BIGGER pretrained LM?
Judging by our new paper on Auto-Contrastive Decoding, in some ways you can! 🤯
https://t.co/PvzSZO9qyE
@IBMResearch#acl2023#NLProc
Label Sleuth is coming to @emnlpmeeting on Friday!🤖
after lawyers fell in love in it at #ELTACON and data scientists at @PyDataGlobal.
Easily label and build text classifiers in a few hours with no need for AI nor programming skills!
#opensource#NLProc
https://t.co/cdDlmLT284
@hemantbuch I think we all misunderstood what Khaled Mahmud said. He meant to say that they have world class "no-ball"ers and SL has none. Remember how many no-balls they bowled in the SL game.
Want to build a text classifier in a few hours?
Even if you don’t have any:
labeled data
#machineLearning knowledge
programing skills
Label Sleuth https://t.co/ViNd3sQNkT a new open-source no-code system for annotations 🧵 @IBMResearch@NotreDame@StanfordHCI UT Dallas #NLProc
Welcome PrimeQA at #NAACL2022! Replicate the state-of-the-art on multilingual open QA quickly! Here’s a new open-source repo in collab with with @stanfordnlp, @huggingface, @Uni_Stuttgart @NLPIllinois1. Link: https://t.co/R6GbTT3mKq Talk to me or read:
https://t.co/GblOy35JGK 🧵
Sri Lanka is in a death spiral. Today, I measure LKA's inflation at 122%/yr. Things are so bad even IMF refuses to offer Sri Lanka a bailout loan. SPOILER ALERT: Sri Lanka has had 16 IMF programs. None have worked.
https://t.co/OT20JfwvzY
I did interview Gota. The Rajapaksas realized it was a mistake that they did declined to comment for our Hambantota port story & requested a meeting before elex
He did not say much interesting so we didn’t use it. It wasnt 3 hours long. I was heavily pregnant &not drinking coffee
Outstanding body of work from IBM Research featured at NeurIPS2021 -- nearly 90 accepted papers, workshops, workshop papers, demos and datasets. Check out https://t.co/L37tLrW1A3 and follow us at @IBMResearch to explore the latest. #ibmresearch#NeurIPS2021