Trying to build LLM as a judge harness for my agent for evaluation on custom rules. Let’s see how it goes!
And ofc, im building it on Emergent :)
The UI and UX on mweb feels so cool! Dopamine from building nice-feeling apps is next level!
Thinking about and planning for different evals is so fun!
How should the agent be behaving, catching wrong/right/golden behaviour etc so many other dimensions!
Anyone here figure out decent strategies for automating evals??