AI/ML Research Engineer building sensor-related inference pipelines, air-gapped DoD LLMs, & defensive algorithms for drone swarms. MS student at U of M.
We gave Jev 2,029 real phone calls.
No transcripts or audio; it never heard a word. Our AI receptionist's calls were reduced to pure structure, meaning turns, tool calls, workflow stages and timing.
During the calls, Jev made 38,012 turn-level forecasts at 118 ms median latency, reviewed every call with five typed questions and produced 10,145 answers in 26 seconds with 256 requests in flight.
The experiment was zero-shot, with no fine-tuning or examples from our data. We compared Jev's forecasts with what actually happened in the EHR.
By the halfway point, Jev could meaningfully separate calls that would book from those that wouldn't (AUC 0.78), and near the end it ranked them correctly 94% of the time.
Even though Jev over-focused on visible errors our agent usually overcomes, it's still pretty incredible that it analyzed thousands of real calls in seconds for only $3.
Basically a Jev configuration, which is neat to see. Low cost model too.
Still prefer working on OpenJev and mini-jev, but it's nice to see them moving quickly on new tech.
Decisions API, for lightning fast constrained decision making powered by Luna. Supports visual inputs, and tuned to be able to make decisions in less than a few hundreds of milliseconds end to end.
Crazy news day. @AnthropicAI filed for an IPO, and @OpenAI is both scrapping model releases and releasing a bunch of other stuff for tomorrow's dev day.
I see lots of excitement over usage limits with Opus 5.5, but I don't get it. I'm burning my weekly limit 2.7x faster than normal.
Usually hit about 15% per day. I'm at 52% just 18 hours after my reset. At this rate I'll get 3 days.
Make it make sense @AnthropicAI
@letstalkksports Yes. Bought League Pass after she joined the Fever. Stopped buying it this year due to the politics. Only watch Fever games now.
She has flaws, but it's their hate of her that makes me want to watch her more. To see her persevere against it.
@buger@OpenAI@thsottiaux Based on what I've seen, it waits a bit until it thinks they are done, but in my most recent case, they were all waiting on (multiple) command approvals so they could finish.
Anyone else finding that @OpenAI models in the Codex app are opening new threads for every subagent instead of running them in the background?
Maybe a bug @thsottiaux
Seems like @typesafeai is botching the invite flow. Get on the waitlist. Get an invite to create account. Signup flow says it's full.
Will forget about it and not go back.
Not great business.