The frontier moved again. Happy to keep up.
In all honesty, prompt screening with a nano model and regex was not the most effective idea to begin with.
Most SMBs have small data sources, it just doesn't make sense to train a classifier at their scale.
Tried out Jev in my system today for two purposes: screening user queries and judging answer fidelity.
It was awesome at the first and was pretty bad at the second. Am I doing something wrong?
10 minutes on X and you'll see 2 camps:
1. Put MCP over your systems, plug in Claude, ship.
2. Custom harness
1 is cheaper. We picked 2.
Why?
You trade time/money for better long-horizon performance, domain fit, and control. I think that's a good trade. What do you think?
@suraj_sharma14@BainandCompany I feel like it's too hopeful to expect them to reach these numbers (as of now at least), so the real question is what happens when the bill comes due?
What if Jev was used for UI/UX use cases?
In the chat assistant I'm building for a company, it's very useful in deciding which artifacts to render into a chat panel.
Like do I need to show sources in this answer or not? Do I need to render a clarification card or not?