LLMs aren't just for GPUs! @cerebras releases a family of models up to 13B parameters on @huggingface to promote open research into scaling laws and demonstrate the capability of CS-2 hardware
Why do this?
1/4
My lobbyists are very nervous about me posting this, but over-regulation is working against us all. The costs are astronomical to us all, but hidden.
So, I'm taking a risk, and sharing my stories from Charm and Revoy:
https://t.co/UsaltST5gJ
Apologies for the delay on getting MoE 101's last episode out!
Originally I planned to cover inference arithmetics only, but we turned it into the MoE inference 101!
I know you enjoyed the MoE training perf modeling, this one is for inference. On both gpus and cerebras.
https://t.co/zrtN8DgA2n
🧵1/n
This might be the most information dense blog I've ever written. Added "show me the math" section into MoE 101 p4 episode. We believe it fully models MoE training perf on both gpu and cerebras wse devices.
https://t.co/uW6H78ZE56
🧵1/n
Enterprise companies are hoping GenAI agents change their business. But security and IP is a non-negotiable.
Wrote a blog on how @googlecloud ensures security while bringing “chat with data” agents to to customers
Thrilled to officially announce what I've been working on for the last year: https://t.co/pWT8MNdLMI!
At Strella, we believe that the customer’s needs should be a company’s North Star. Using Strella’s AI, we enable companies to make informed decisions in hours, not weeks ⭐️🌟🚀
Man so excited we could finally unveil this. This is THE applied AI project. Google walked away from it. We embraced it. The world is a different place because of it. It provides so many foundational learnings that we are now applying to the commercial world via AIP. The origin story of Silicon Valley innovation reincarnated.
Crazy speed from the team at @cerebras! Unlocks lots of interesting use cases across fast agent tool calling, multi-agent systems, self-consistency, and more!
Verified by @ArtificialAnlys, Cerebras Inference achieves 1,850 tokens/sec on Llama 3.1 8B and 450 tokens/sec on Llama 3.1 70B!
By dramatically reducing processing time, we're enabling more complex AI workflows and enhancing real-time LLM intelligence. This includes a new class of intelligent agents that can “think faster” than ever before.
Cerebras Inference will power a new era of Instant AI.
👉Try it today: https://t.co/jREGhLI2nj
👉Read our blog: https://t.co/1rYd32ELnf
👉Check out Artificial Analysis for more data: https://t.co/71Br7To8Qf
I read @leopoldasch's essay on the future of AI research and geopolitical competition. It's well-researched, well-presented, and passionate. However, Leopold advocates for an unreasonably strict and exclusionary future for AI development—a view that's gaining traction. (1/9)
🥳 CerebrasGPT proved to the world in March how effectively you can train LLMs on @cerebras hardware. Now BTLM surpasses 1M downloads in ~3 weeks on @huggingface! 🚀
Cerebras BTLM-3B-8K model crosses 1M downloads🤯
It's the #1 ranked 3B language model on @huggingface!
A big thanks to all the devs out there building on top of open source models 🙌
The Cerebras team has had a great time sharing our work at #ICML23. Below is a summary of the posters we presented, let us know if you are interested in discussing any of them further!