Doing a private conversation with @bryan_johnson on longevity + tech on Aug 26 (with bioage testing on site)
Thanks @baseten for hosting.
Limited spots: https://t.co/Oyy0o7Jx8T
Last week, the entire Baseten team came together in Chicago for our biannual company offsite.
Since the beginning, we've come together twice a year in different cities to connect and build together: talent shows, hackathons, customer conversations, and late-night karaoke sessions.
At one of our early offsites in 2022, there were 19 people. It's hard to imagine that now, with nearly 400 people (and counting), this is the last time we'll be this "small."
We're grateful to all the wonderful people who make Baseten's culture what it is today: fast, quality-obsessed, and deeply committed to our customers.
If you want to join us for the next one, we're hiring.
Calling it now, in 6 months most companies will just be self hosting OpenWeight models at 90% cheaper price point and if you don't hop on this right now you will be playing catch up.
To pursue self hosting, we simulated @cline live production traffic onto 8-GPU B200 nodes, measuring TTFT, per-stream speed, and cache hit rates and now we are sharing everything about self hosting so that you can get all the open weight benefits like protecting your data, not being blocked by model nerfings of frontier labs and SAVING a lot of money.
If you are an AI engineer, a tech lead or in leadership at any AI company and the AI spending is causing you stress, this blog will answer all of your questions about how to manage your money, engineering requirements and upholding quality. Most of you think that your only saving grace is using a cheaper model and begging providers for discounts but self hosting is the much better option for cheaper inference and data autonomy.
You don't have to do everything yourself, you can get really really far by just understanding how your own AI traffic is working and then getting any inference provider to work with you using the advice posted in the blog.
Special thanks to @baseten (The OG @philipkiely ) and @CoreWeave team for helping us understand inference and walking us through how to scale and try different traffic cases.
We encourage people to read through the blog and get the possible inference deals from any provider that works best for their specific use case.
Introducing ask-web: Rox’s in-house web search agent.
ask-web sits on the cost-per-accuracy pareto frontier of the hyper-parameter grid when compared to frontier labs and commercial search agent providers.
The agent delivers 91.3% accuracy at 1.03 cents per query on real production prompts.
It has been running in production for more than 6 months with continuous evals.
Inference partners: @togethercompute, @baseten, @modal
Commercial Search vendors benchmarked: @perplexity_ai, @ExaAILabs, @p0.
Frontier Search vendors benchmarked: @OpenAI, @AnthropicAI
Exa, OpenAI and Anthropic excel on accuracy. Parallel and Perplexity are cost-efficient.
Here’s the breakdown:
Been working on this for a while - Built on Baseten is happening.
August 4th, SF. Four early-stage founders getting on stage to share what they built and the infra decisions behind it
More about them below 👇
You can now access our GLM-5.2 API through the Merge Gateway!
GLM-5.2 matches frontier model intelligence while running 4x+ faster and at 1/5th the cost.
Try it out: https://t.co/OuguJONiFK
We've been happy to partner with @baseten as a model vendor on Merge Gateway for a while now, and GLM-5.2 is now available through them too.
For those considering open-source models for coding, agentic or other use cases, GLM 5.2 has comparable intelligence to Opus 4.8 at ~1/5th the cost and 4x performance (280+ TPS) on Baseten.
Check it out here: https://t.co/4jUpjeUq3h
Was a lot of fun telling @cursor_ai how we manage 128+ agents and making some predictions on the future of language models. Plus hear about how Harry names his agents after mathematicians and I name mine after NBA players, which probably unnecessarily hobbles mine in comparison but I love it too much to stop
Good take
My guess is
- demand for intelligence is near infinite
- but 80% of workloads will be running on 99% cheaper models within 12-18 months
- 20% of workloads will still run on latest gen models where IQ maxing is important (scientific breakthroughs, higher level ochestrator agents?)
- rough analogy might be what % of macbooks or gaming PCs sold have the maxed out specs for CPU/GPU, prices are falling much faster than Moore's law here though
- this leads me to think the limiting factor will be energy and compute, not better models
At Coinbase we're working hard on routing prompts to cheaper models where appropriate, and in some cases have been able to keep costs roughly flat, while token usage continues to grow exponentially.
10M developers use @opencode every month. This means the experience has to feel the same every hour of every day; slow or inconsistent inference breaks productivity and flow.
Enter Baseten. With Baseten's Model APIs, OpenCode achieves 5x faster TPS, sub-second TTFT, and 33% blended cost savings passed directly to users through cache token pricing.
Philip worked with @rimelabs to create an enterprise-grade voice clone for an ambitious project: an AI-narrated audiobook that provides an enjoyable listening experience.
Now you can "read" Inference Engineering while running your agents.
Model labs should spend their time pushing the frontier, not thinking about API keys, rate limits, metering, and billing.
Today, we're launching Baseten Frontier Gateway: the fastest path from trained weights to a production, white-labeled API. https://t.co/1tmF8Xq9OE