Been looking forward to this release, and it’s here
Super excited about the
decrease in base input tokens on a default-agent turn by 65%! 👀
Will be trying this out today
Most AI audio models have never heard a maqam.
Team Motif fine-tuned Stable Audio 3.0 on Arabic maqam, built an Ableton plugin for microtonal style transfer, and won our Stable Audio 3.0 Challenge at Music Hackspace running locally on device.
Watch Jad Al Masri break it down 👇
Nice analysis from Peter here about one of the puzzles in AI - given all the changes AI is causing within the economy, why hasn't it had measurable impacts on unemployment?
Opus 5 is available today on all paid plans and the Claude API, priced the same as Opus 4.8. It’s the default model on Claude Max, and the strongest on Claude Pro. It’s also offered in Fast mode, which runs around 2.5× the default speed.
Read more: https://t.co/YMw7gLMT7e
Voice mode now runs on Claude's more capable models and reaches the tools you've connected mid-conversation.
Talk through the hard problems out loud, in many more languages.
a website where i can follow my friends and see what they are up to. who is building this. novel idea. this could get so much in funding. then to maintain their valuation all they'd have to do is slowly replace all my friends with a global feed of generic slop. oh
2002-2004: Software guy thinks he can build rockets and cars
2016: Car guy thinks he can do brain surgery
2018: American guy thinks he can build factory in China
2022: Hardware guy thinks he can do software
2026: “matters outside his sphere of expertise”
March 31st is the last day to submit proposals for projects with measurable results in reducing indoor airborne pathogens, like breath-based multi-pathogen detection systems, continuous HVAC compliance verification systems, and indoor air infrastructure deployment.
New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was built with @Cerebras and taught by @zhennydez, @duerr_seb, and @MilksandMatcha.
When a model generates text, much of the time is spent moving its weights out of memory and into the compute units. Inference-optimized hardware minimizes that movement, making token generation several times faster than on a typical GPU setup. In this course, the hardware you'll use is Cerebras' Wafer-Scale Engine, which is designed for fast inference by keeping the model's weights close to the compute units.
Fast inference makes lengthy agentic workflows go faster, and also unlocks latency-sensitive, real-time applications like live translation and voice agents.
Skills you'll gain:
- Compare how GPUs, TPUs, and Cerebras' Wafer-Scale Engine each handle the memory-to-compute bottleneck
- Build real-time applications powered by fast inference, including personalizing a webpage and running a multi-step workflow to analyze market signals
- Adopt concrete habits for agentic coding with fast inference, keeping your sessions focused and steering the model more effectively
My teams use Cerebras for several applications that are latency sensitive. Join and build LLM applications that respond quickly:
https://t.co/P8vchGAr22
Meta AI apps got upgraded with Scheduled actions and Artefacts on mobile!
> Do more with Meta Al - Set up daily briefings, get help with your calendar, and create interactive artifacts.
Meta is expanding 👀