Had a great time talking to @nytimes in this article about @Box and the future of work with AI. Like I said in the article, it's not the volume of jobs that will change but rather the type of jobs that people will do. Evals are just one example of an entirely new category of work that more and more people will be doing in the coming years.
Full story here: https://t.co/ilMExeBiaQ
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Me asking Fable 5.1: "You are a smarter and more capable model than all the models that have done all the work in this session. Give me an audit of where we're at and if there's any room for improvement."
Box's early evals were manual, with responses marked in a spreadsheet or a Box Note.
With Braintrust, they built a practice of running evals programmatically.
Now product managers have a custom eval runner that pulls versioned, schema-validated datasets out of Braintrust, hits the agent endpoint, grades the outputs, and loads results back in.
Read more → https://t.co/qW9qwkbiD3
Claude says "genuinely" so much I built a benchmark about it.
GenuineBench: 13 frontier models × 80 everyday writing tasks. Every Claude model says it in ~1 of every 3 responses. Llama 4 has never genuinely meant anything in its life (2.5%).
Claude says "genuinely" so much I built a benchmark about it.
GenuineBench: 13 frontier models × 80 everyday writing tasks. Every Claude model says it in ~1 of every 3 responses. Llama 4 has never genuinely meant anything in its life (2.5%).
Claude says "genuinely" so much I built a benchmark about it.
GenuineBench: 13 frontier models × 80 everyday writing tasks. Every Claude model says it in ~1 of every 3 responses. Llama 4 has never genuinely meant anything in its life (2.5%).