I’m glad @strawgate's standards for coffee are lower than his standards for code review because many coffees were brewed by yours truly in the creation of this video. I was thrilled to sit down with Bill, who is the head of product at @pydantic, to hear about how he and the team use @Macroscope. It’s an absolute joy to work with Pydantic and to have a small part in supporting their work.
@kayvz I love that Murmur can "adopt" existing old, stale, and stagnant PRs in our projects and get them green, respond to comments, keep them up to date, and to make sure that they are always ready for review/merge.
@coderabbitai Being an OSS maintainer is a lot of work and Coderabbit has consistently done a great job supporting maintainers with their free code review, poems and PR walkthroughs.
"But I use Braintrust for evals."
Trust us. We know. We hear it all the time.
Good news: Logfire now speaks Braintrust. Point your existing Braintrust SDK at us with one environment variable. Zero code changes. Keep your evals, keep your scorers, keep paying Braintrust if you're attached to the invoice.
Try both. Compare.
Big brains run experiments.
Day three, @strawgate remains in charge. https://t.co/MJ1sQhIPvr
@aiseomastery@pydantic It's not hard to spend $36 on Braintrust, you just reduce your eval dataset by 99.9% or you can just run the eval suite once every 3 years instead of nightly.
"Big brains trust Braintrust."
Big brains check the pricing page.
We checked it for you: one night of serious evals uses the entire monthly score allowance. Nights two through thirty bill at $1.50 per thousand. That's $26,100 a year to score one agent, nightly, using OpenAI's own public eval suite — before data and retention.
Our pricing page is shorter. Scores are spans. Spans are $2 a million. Same year of scoring: $36.
$36 vs $26,100.
Big brains read the article https://t.co/PPlUxviWoW
Our social media team is off for the week, so @strawgate's in charge. Separately, his "no making fun of competitors" pledge expired at midnight.
We resume operations with a pricing observation: 50 million eval scores on Braintrust Pro costs $75,174. In Pydantic Logfire, a score is just a span. Same 50 million: $100.
The meter doesn't just cost money. It's why your sampling is set to 10%, and why last week's failure lived in the other 90%.
Full arithmetic: https://t.co/dLth8Yqd8b