OPUS DOESN'T NEED TO MAKE EVERY DECISION IN YOUR AGENT
Jev found 51 of 54 instructions buried in 30,143 lines of a call transcript for $0.32.
I published the tests and a prompt to audit your agent. https://t.co/cjQUagj9ue
I put Jev in front of an AI writer to save tokens.
The writing got worse
Then I gave it the small decision *before* the writing. If you're building your first agent, this is the job I'd test first https://t.co/ZbU4ioAxro
@higgsfield_ai@gregisenberg This is amazing! Can we also get “same building all the way through” button? I make videos of real places and that's still the thing I fight most 😅
@seroundtable Imagine spending months deciding exactly what goes on your hotel's website, then Google puts an AI Overview beside your name and writes the answer for you
@karrisaarinen@linear Can it switch models mid-task? A lot of my work starts as a messy research question and ends in code. Sometimes the model I’d pick at the start isn’t always the one I want finishing it
Testing Jev tonight on 936 multilingual intent-routing examples.
It got 65% right, while assigning its top answer 80% probability on average.
It’s fast as hell. That gap is what I’d want to understand before letting it make decisions inside my agents.
@marclou Very true, was about to buy some hyped jev domain last night. l’m looking at Jev right now but want to find one place in the agents I already run where it actually replaces a full LLM call. Seems like a better starting point than building a new wrapper
@lilyraynyc I can see that someone landed on a page from ChatGPT. I still can’t see what they asked or what answer made them click. That’s the report I’d actually want!
@muratcan Did Opus see ground-truth frames of the location, or only the prompt and output? I’m making videos of real hotels; a shot can look physically plausible and still be the wrong building!
@tiangolo I posted about this today. The v1 can be impressive and still be useless for the final user. I’m much prouder of v13, when it finally survives the actual constraints.
Can someone explain me what is this "one shotted" or "one prompt" obsession?
So proud of my "_v13" iteration that is actually working and perfect for my prod environment or for my clients
Still can't believe I left my 9-5 very well paid job 3 months ago to found a startup, and I am still with no salary, spending my savings and working nonstop.
back to X and want to start seriously building in public
Opened my Codex stats: 35.1B tokens.
I can get AI to execute almost anything when I already know what I want. But I sat down to write a video and got stuck on the part before the prompt: what the hell is it actually about?