We benchmarked Opus 5.5 & Fable 5.1, they built the same game and put GPT-6 Astra to judge the results.
Opus 5.5 does perform almost the same as Fable 5.1, spending almost the same amount of tokens, but as each token costs less, it should positively affect usage.
Video soon...
@_cmd8@typesafeai Posted the same without having seen yours before.
https://t.co/qkCzvGbvcd
An important caveat: you'll lose the deterministic output if you change the questions JSON (input).
Static inputs alone don’t prove a @typesafeai's Jev demo is fake. A tool or LLM can update the input after every decision (not just the state).
With a simple, well-trained feedback-loop, you can have those "fake" demos generate websites, draw live in screen, play games, etc.
What happens when Jev, Claude and Codex share one voice chat app?
Meet Jauvex. Talk to your coding agents, steer them while they work, and let them talk to each other. Jev handles the quick decisions.
What's crazier: the app builds itself.
Here’s the demo. 👇
@ok1mraise It rewrites its own source code, it has an AGENTS.md file with the instructions.
Probably gonna keep it lean so that people can customize it however they want.
@typesafeai Perfect timing, as I just build an agnostic agent orchestrator, where Jev is the decision maker.
I had set it as optional as not everyone had access to Jev.
https://t.co/Snkf34fBTL
What happens when Jev, Claude and Codex share one voice chat app?
Meet Jauvex. Talk to your coding agents, steer them while they work, and let them talk to each other. Jev handles the quick decisions.
What's crazier: the app builds itself.
Here’s the demo. 👇
@ok1mraise Well, to be fair Codex has a very seamless two-way voice chat. For me the problem was that Claude didn't have it, and in the end I ended up building something that connects multiple agents.
What's different from this app, is that it's able to build itself. I'll OS it soon.
@mathfax But most of those flashy demos are valid. I might be wrong, and would love to hear your feedback: can you technically connect an LLM or script that refines the input (questions) until you get a stable flow and then unplug the aid to keep it deterministic?
https://t.co/qkCzvGbvcd
Static inputs alone don’t prove a @typesafeai's Jev demo is fake. A tool or LLM can update the input after every decision (not just the state).
With a simple, well-trained feedback-loop, you can have those "fake" demos generate websites, draw live in screen, play games, etc.
@putrikarunian The model can't write, but the app using Jev can decide to call a tool that can make it write. The flow can generate or learn this way, the input (questions) can be trained with a feedback loop this way (plug in an LLM, refine your input, remove LLM):
https://t.co/qkCzvGbvcd
Static inputs alone don’t prove a @typesafeai's Jev demo is fake. A tool or LLM can update the input after every decision (not just the state).
With a simple, well-trained feedback-loop, you can have those "fake" demos generate websites, draw live in screen, play games, etc.
@typesafeai Obvious caveat: it would affect the deterministic output feature when run against the previous input. However this could work to pre-train a fixed curated input.
Static inputs alone don’t prove a @typesafeai's Jev demo is fake. A tool or LLM can update the input after every decision (not just the state).
With a simple, well-trained feedback-loop, you can have those "fake" demos generate websites, draw live in screen, play games, etc.
@nicdunz But then again, why should the input be static between calls???
If you change the input with an LLM or tool then you open new possibilities, and then those "fake" demos no longer are "fake".
@elshayib_ That's not true. It can do those things and more as long it knows how to act based on its structured input.
It can even learn, generate, and adapt, if you connect it with a tool or an LLM that makes its input dynamic.
Out-of-the-box sure it's a static decision making model.
@Ramanean It’s never been about building it, it’s always been about how to market it… this is the sad reality of creating something. Ask Tesla and Edison otherwise.
@SSHCodes@Steve8708 I used @browser_use with Jev and was able to navigate, won't lie that it wasn't straightforward, however it did work after it was setup correctly, now I can i.e. categorize "hype" of my X feed automatically based on a predefined criteria. Jev needs an LLM for planning.
@Steve8708@ctatedev@vercel@Steve8708 I understand very well that @typesafeai's Jev is a decision making model, and I know it's not generative (I made a demo myself in profile), but I wouldn't jump so fast to accuse someone and say that their demos are fake.
2nd in your video:
https://t.co/ujuIe98wlC