@arizeai OSS team member @ehutt_ put PXI, Claude Code, and Codex through the same tasks with Harbor + Phoenix.
Even if the agents succeeded, vtheir traces told very different stories: different tool calls, latency, and efficiency.
Read more: https://t.co/mDRST1dfyS
Arize Phoenix continues as an open-source project. OpenInference continues as an open, OpenTelemetry-compatible project.
If you're already building with Arize Phoenix, keep building.
@HamelHusain The bot replies are often completely nonsensical too. Do they just like to hear themselves “speak”? It’s such a bummer. Makes this platform very unpleasant to spend any time on
Started rewatching the X-Files. I got a kick out of the episode about the killer AI made by “Eurisko” (totally not Cisco). If you’re looking for a good show to get into the mood for spooky season, I highly recommend it. The truth is out there! 👽🛸👻
Benchmarking AI agents & tool use with Harbor + Arize Phoenix.
Join us Oct 8 at 11am PT / 2pm ET to learn how to compare agents or models over the same task set, separate behavioral scores from infrastructure failures, and inspect ATIF traces.
https://t.co/wPqe8nR2kG
@henrytdowling Try Arize Phoenix! It’s free and open source, we built instrumentation for all kinds of coding agents and frameworks, and it’s designed specifically for the workflow you’re describing. Disclaimer: I work on it.
@seldo compared Jev, Claude Opus 5, and GPT-5.6 Terra with 23,325 judgments on accuracy, cost, latency, and calibration.
We found that Jev matched Claude Opus 5 at 87% hallucination-detection accuracy while running 23x faster and at roughly 1/300 the cost.
https://t.co/oOKt8kSJiR
i've been surprised at the response to Jev, but it makes sense in retrospect. sure it's just a classifier but it's a zero shot classifier with frontier-ish intelligence. i'm surprised someone hadn't built it before. i wonder what other old ML ideas are also worth rescuing
Benchmarking agents and the tools they use (mcp, cli, skills, etc.) is HARD. I'm going to show you how we use @harborframework (plus the new Phoenix plugin for Harbor!) to evaluate all of our Phoenix interfaces (our PXI agent, the MCP server, the px CLI, and all our skills).
Webinar coming up: benchmarking AI agents & tool use with Harbor and Arize Phoenix.
Learn how to compare agents and models on the same tasks, diagnose regressions, and inspect scores, errors, and traces.
Oct 8 · 11am PT / 2pm ET
Save your spot:
https://t.co/wPqe8nR2kG
TypeSafe came out of stealth this week with Jev, the first System One Model: a frontier model that never generates prose. Unstructured state in, typed probabilistic decisions out.
The power of LLMs comes largely from their ability to 1) accept natural language inputs and 2) generalize across tasks. But ever since they became widespread, we have been trying to wrangle those outputs to be more usable and trustworthy.
Jev from @typesafeai is one of those obvious-in-hindsight ideas that makes me wish I’d thought of it myself.
We’ve long treated LLM evals as a classification problem, borrowing from classical ML to develop and validate our judges. Jev takes that framing to its logical conclusion: strip away the generation we don’t need, and get the classification at a fraction of the cost.
I’ve been banging my head against the wall all week with Astra and thought I was doing something wrong. It was starting to drive me a little insane. The tasks were complicated but well specified, and astra went in truly bizarre directions. Every round of review and feedback led to more issues. I’ll try going back to Sol.
a portion of our team has gone back to Sol
astra is good and can do some novel things but it has some downsides
and so far our effective spend looks doubled so tough to justify
PSA: Did you know that when you call your bank/doctor/airline whatever and they put you on hold, that hold is recorded? Elevator music and all, both sides are recorded.
How do I know this? Well, I used to work on AI for contact centers and have listened to some of those recordings... it's very creepy hearing people talk when they don't think anyone is listening.
If your agent needs 85 MCP turns to answer a SQL-shaped question, the problem may be the interface you gave it.
Retrieval is great when a coding agent needs to inspect a few traces. It gets clumsy when the task is really about counting, filtering, joining, or aggregating across lots of telemetry.
Arize Phoenix now gives agents another option: https://t.co/DURe4UJ8uF