There are so many useful signals to collect from your traces in @langfuse like topic, frustration, signs of success...
love @typesafeai making those checks cheaper to run together at scale with Jev!!
Jev-powered evals are now available in Langfuse
👉️ Pay 40-400x less than with frontier models
👉️ Run Jev (by @typesafeai) on all production traces without sampling
👉️ Check multiple criteria at minimal extra cost
https://t.co/1RKrhsfYnW
Jev-powered evals are now available in Langfuse
👉️ Pay 40-400x less than with frontier models
👉️ Run Jev (by @typesafeai) on all production traces without sampling
👉️ Check multiple criteria at minimal extra cost
https://t.co/1RKrhsfYnW
Added cost waterfall to @langfuse cost tracking breakdowns. You can now immediately spot what drives your generation cost (and whether caching kicked in or not)
the trace timeline now fits any trace on one screen. twelve spans or twelve hundred.
zoom like a map: scroll to pan, pinch or cmd-scroll to zoom both axes, double-click an observation to fly to it. zoomed out, color carries the observation type.
https://t.co/Y0DMvCgKsj
Introducing out new evaluator template gallery!
The right evaluator setup depends on the specific application context and the failure modes observed. But some general evaluation and classification approaches provide a great starting point early on.
Use the templates to:
✔️ Classify common topics of user requests
✔️ Check if retrieved context was accurately used
✔️ Catch user interactions where users disagree
A good eval set ≠ a large eval set. Read @lotte_verheyden's latest article on picking the right metrics for your AI application. Now also available on https://t.co/4bH9oOYfwc
Before you write good evals, you have to pick the right things to evaluate. Main take-aways:
- source metrics from failures you see in traces
- keep the set small and maintainable
- only keep a metric if you'd take action when it moves https://t.co/G2FZjfidrb