🚨New paper: Reward Models (RMs) are used to align LLMs, but can they be steered toward user-specific value/style preferences?
With EVALUESTEER, we find even the best RMs we tested exhibit their own value/style biases, and are unable to align with a user >25% of the time. 🧵
Cannot attend #ICLR2025 in person (will be NAACL and Stanford soon!), but do check out 👇
▪️Apr 27: "Exploring the Pre-conditions for Memory-Learning Agents" led by @viishruth and Vishwa Shah, at SSI-FM workshop
▪️Apr 28: our @DL4Code workshop with a fantastic line of works & speakers!
▪️Apr 28: @dan_fried's talk at the DL4C about "Inducing Functions to Improve LLM Agents"
Another banger from my group that tbh raises more questions about creating data agents then answers.
Here's the core issue:
When creating an agent to query structured & unstructured data for business insights, how do you describe these data tools?
Let me elaborate 🧵👇
🧵 1/ Recent hype suggests long-context LLMs remove the need for retrieval in RAG pipelines—"just put all your docs in context." We tested this theory rigorously in finance, focusing on SEC filings.
Spoiler: Retrieval & chunking strategies still dominate.
Snowflake Cortex Agents, now in public preview!
Cortex Agents orchestrates across structured and unstructured data for accurate AI-driven decisions from within the secure Snowflake perimeter; Cortex Agents use Cortex Analyst (now in GA) and Cortex Search as tools.
@AnthropicAI's Claude 3.5 Sonnet is used by Cortex Agents to deliver accurate, efficient and governed data insights at scale. All the details: https://t.co/jSGMFmmD3N
How far are we from having competent AI co-workers that can perform tasks as varied as software development, project management, administration, and data science?
In our new paper, we introduce TheAgentCompany, a benchmark for AI agents on consequential real-world tasks.
Excited to share TimeSeriesExam for systematic evaluation of time series reasoning capabilities of LLMs. Think your LLM can reason on time series concepts? Take it for a spin on the TimeSeriesExam! Now publicly available on HuggingFace :)
🚀 I am thrilled to introduce @SnowflakeDB 's Arctic Embed 2.0 embedding models! 2.0 offers high-quality multilingual performance with all the greatness of our prior embedding models (MRL, Apache-2 license, great English retrieval, inference efficiency) https://t.co/hEcd0niVyr🌍
We are excited to share SwiftKV, our recent work at @SnowflakeDB AI Research! SwiftKV reduces the pre-fill compute for enterprise LLM inference by up to 2x, resulting in higher serving throughput for input-heavy workloads. 🧵
Tired: Bringing up politics at Thanksgiving
Wired: Bringing up @datologyai’s new text curation results at Thanksgiving
That’s right, we applied our data curation pipeline to text pretraining data and the results are hot enough to roast a 🦃
🧵
🧵We’ve spent the last few months at @datologyai building a state-of-the-art data curation pipeline and I’m SO excited to share our first results: we curated image-text pretraining data and massively improved CLIP model quality, training speed, and inference efficiency 🔥🔥🔥
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
by @s_waghjale, @viishruth, @ZhiruoW, and @dan_fried
Session: NLP Applications 1, Session 02, 11:00-12:30
https://t.co/YpAEBhpUNf
As we prepare for EMNLP 2024 (@emnlpmeeting), we're thrilled to congratulate all of the LTI researchers with accepted papers at this year's conference. In all, 32 papers with LTI authors were accepted! Read about them all here:
https://t.co/o77FizuTE4
Sad to miss #EMNLP2024 but do check out our paper "ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?" presented by @viishruth and Siddhant
Tuesday 11-12:30 at Poster Session 02‼️
I’m attending #EMNLP2024 in Miami from 11-16th Nov to present ECCO on Tuesday 🏖️
Looking forward to meeting folks and chatting more about code generation and LLM agents!
Can current code LMs generate sufficiently efficient programs? 🤔
More importantly, Can these LMs improve code efficiency without sacrificing correctness?
Check out ECCO, our code-gen benchmark for correctness-preserving program optimizations!
🧵 1/n