Biotech R&D is generating more scientific AI models than ever, from protein structure prediction to molecular docking to sequence analysis. But the infrastructure to run them hasn't kept up.
Today we're announcing Benchling Inference, powered by Baseten. Together with @benchling, we're delivering on-demand GPU capacity built for the bursty, high-stakes demands of scientific workloads. With Benchling Inference, scientists can:
→ Deploy models in seconds, not weeks
→ Keep proprietary models inside their VPC if needed
→ Benefit from economics that work even at small and mid-size biotech scale
Benchling and Baseten decided to team up because we believe that research teams shouldn't have to manage HPC queues, negotiate cloud contracts, or become GPU experts to run frontier models on their own data.
Six years of inference expertise are now available where science happens.
Read more here: https://t.co/vqmtnXnAT1
We painted San Francisco green and pink, and the message is clear — you need to own your inference.
If you spot us around the city, share a picture with us. We’ll send you something!
Generational AI companies are powered by Baseten.
Why? We obsess over the milliseconds, so they can ship the future.
Focus on what actually differentiates you. Leave the inference to us.
The biggest hurdle to widespread AI adoption isn't just model capability, it's the cost and speed of inference. At Baseten, our mission is to make the world’s best models run at peak efficiency.
Our engineering team just reached a significant milestone. By implementing a hybrid speculation engine (combining Multi-Token Prediction with a Suffix Automaton), we’ve achieved up to 40% boost in throughput for production coding workloads.
https://t.co/xylJcoDPqJ
We boosted acceptance rate by up to 40% with the Baseten Speculation Engine.
How? By combining Multi-Token Prediction (MTP) with Suffix Automaton (SA) decoding.
This hybrid approach crushes production coding workloads, delivering 30%+ longer acceptance lengths on code editing tasks with zero added overhead. An open source version for TensorRT-LLM is now available to the community.
Read the full engineering deep dive: https://t.co/Z1kS863vOC
Baseten’s day 0 bet was that inference was the technology that would enable the best user experiences AI could deliver–fast, smart, reliable, secure. And that those experiences would rely not only on a handful of giant general intelligence models, but millions of specialized models built by companies for their specific customers and use cases.
Whether you’re a doctor, developer, lawyer, mechanic, researcher, construction worker, marketer, etc, you’re accelerated by specialized tools worthy of your craft. To me, this is one of the most meaningful promises AI can deliver on.
We’re starting to see it now. Many of the main-character AI companies on the application layer are built on highly-specialized models for highly-specialized workflows–Abridge, Clay, Cursor, OpenEvidence, Hebbia, Mercor, Notion–these businesses are booming because customers love specialized tools.
There are probably hundreds of custom models in production today. Soon, there will be thousands and then millions. All enabled by a high-performing inference layer.
Inference has emerged as one of the hardest problems in modern AI systems. Delivering reliable, low-latency experiences requires deep coordination across distributed infrastructure, kernel-level performance, and software ergonomics—even world-class teams struggle to do this well. As a result, as consumers and developers, we’ve grown to accept sluggish performance, frequent downtime, and inconsistent quality across both application companies and model providers.
Meanwhile, the demands on inference are accelerating: AI adoption is trending towards ubiquity with reasoning models that are orders of magnitude more compute-intensive. This will only increase as more companies catch on to the virtues of owning their end-to-end IP rather than relying on black-box model APIs on shared infrastructure. Whether we can realize the impact of this generational shift will depend on our ability to serve these models reliably at scale.
We knew we could make the technology work, but the biggest delight of it all has been seeing what our customers do with it. The (many-model) future is bright.