At Sepal, we create RL environments and tools for Frontier Labs to evaluate models & build real-world capabilities.
This is made possible by working with the world’s leading experts who contribute their skills and domain knowledge to test and evaluate models. Over the last week alone, hundreds of new experts have joined Sepal’s mission: to extend humanity’s knowledge and capabilities through the development of safe AI.
If you have years of professional expertise in a domain we’re recruiting for, you want to contribute and learn more about the AI industry while making money working remotely on your own time. Join us!
Apply Here: https://t.co/4KMunYcQrH
Our team at Sepal AI is at NeurIPS! If you're interested in applying for a role or just want to talk, come say hi!
We're hiring across dozens of industries to find the brightest in their fields to contribute to our mission. If you're ready to help develop frontier AI, join us!
Apply Here: https://t.co/yhtw4nocYi
The world runs on spreadsheets. Can AI agents handle real-world workflows?
We've worked with @sepal_ai on SheetBench-50, a financial analyst benchmark verified by experts from PwC, Cisco, Charles Schwab, and Fannie Mae
For now, Sonnet 4.5 is the top scorer, beating OpenAI CUA👀
This enables scaled recruitment for agentic and reasoning data pipelines for models and agents, making it easier than ever to hire the top experts across 12+ domains from Sepal's Global Talent Network.
Introducing the new Sepal AI Global Talent Search Engine - a powerful solution that we use to expedite the recruitment of experts for training, evaluations, and model testing. This platform is one of the things that enables our customers to maintain a leading edge in scaled human data operations.
If you are interested in systematically studying risk, building a custom benchmark, or accessing the Sepal AI talent grid, please contact us 🤝
https://t.co/h2ZQQpcaLS
🚨 @AnthropicAI’s Claude 4 breached the ASL-3 risk threshold 🚨, delivering a 2.53× boost in participant capabilities on key tasks in a set of Uplift Trials supported by Sepal AI. Here’s why it matters 🧵
This type of pre-launch testing, coupled with a thoughtful risk framework like Anthropics ASL or Open AI’s preparedness framework, is extremely important for individuals building AGI to consider investing in as model performance continues to improve 📈
Congratulations to @AnthropicAI for the launch of Claude 4! Sepal is honored to be featured on the model card for being partners in supporting Anthropic’s important work leading in AI safety.