Ran some evals on Jev vs GLM-4.6 on Healthbench and got some crazy results.
jev in terms of -
accuracy : statistically same
latency : 100x faster
cost: 100x cheaper
parse failures: 0 on both models
@typesafeai@CompleteSkeptic
If you haven't setup evals for your product or are looking to explore evals, you can try out @evaris_ai . It helps you setup evals using inspect-ai and evaris platform and takes hardly 1 min.
You can quickly setup evals in your CI using our Agent skill.
Its simple. Try it.
🚀 introducing etrace — open-source AI agent tracing library
one primitive: trace
auto-instruments LLM calls
tracks costs across 1700+ models
16 semantic span kinds built for AI workflows
-
📊 etrace studio — visualize, inspect, and debug every trace locally