It's important to support newly released open-weight models on day 1. But it's not noteworthy. What's noteworthy is to have the inference optimization muscle to immediately blow the competition out of water on latency and throughput.
As measured by OpenRouter:
It's important to support newly released open-weight models on day 1. But it's not noteworthy. What's noteworthy is to have the inference optimization muscle to immediately blow the competition out of water on latency and throughput.
As measured by OpenRouter:
๐ Our "technical" marketer might not be looped in, but today is our biggest launch day yet.
We're introducing two new products to serve the inference lifecycle: Model APIs and Training.
Model APIs are frontier models running on the Baseten Inference Stack, purpose-built for production. Baseten Training (Beta) provides infra and tooling without limitations for AI models destined for production.
Huge shoutout to the many partners and customers we've worked with as we built these two new productsโmore details below.
2025 is the year of inference.
We're thrilled to announce our $75m Series C co-led by @IVP and @sparkcapital with participation from @GreylockVC, @conviction, @basecasevc, @spc and @lachy. We're also excited to add Dick Costolo and Adam Bain from @01Advisors as new investors.
Check out our CEO Tuhin's blog to learn more.
It's time to build!
Baseten launches Mistral 7B API with leading performance ๐
@baseten has entered the arena with their first serverless LLM offering of Mistral 7B Instruct. Artificial Analysis are measuring 170 tokens per second, the fastest of any provider, and 0.1s latency (TTFT) in-line with the next fastest provider.
Mistral 7B Instruct from @Mistralai has become the leading 7B-class model for use-cases from RAG chatbots to data extraction. With pricing more than two orders of magnitude lower than large models like GPT-4 and Claude 3 Opus, models like Mistral 7B Instruct enable use-cases that wouldnโt be possible with the economics of larger models.
We're excited to announce that we've raised a $40M Series B to help power the next generation of AI-native products with performant, reliable and scalable inference infrastructure.
https://t.co/NAn8LduZ6I
@tylerangert yeesh where are you seeing $1k to tune llama? my take is cost doesn't seem to be the issue (see https://t.co/xBJAcW7Ryk its free to start) it's the data set. curating and and getting it right just takes time