Streaming app UI challenge, with complex animation in the star button.
The artifacts in the video because of the recording, the actual app is very smooth.
https://t.co/aohUv2hNHn
created with @FlutterDev
@evilsocket Probably because you had too much context and got a cache miss. In Claude the cache duration is 1h and I assume they are similar.
And since cache hits are 10x cheaper, one miss on a long session really hurts.
@pkyanam@neural_avb@liviusa@OpenAIDevs@thsottiaux For codex cache IS shared access accounts. You can personally test it. It's hard to believe but it's true.
It's not a security issue because you can't do anything with it.
@jezell I guess it depends on the app, so the apps I use every day, I guess I will not mind even if they are 500 megabytes, because I have 100 GB of free space, which is most new smartphones, and nowadays. But for commodity apps, I would choose the smaller one if all else is equal.
@ThomasBurkhartB I think you still need design taste, this is because, like with software development, AI can do pretty much anything, but if you don't know what's right and what's not, the entire package will be just slop. Perhaps this will change soon, but I don't think we are there yet.
@badlogicgames I don't think this is happening today, but I'll concede this point to you.
I actually do 80% of my research using my coding agent and give it an MCP server to scrape websites.
@badlogicgames It has the added benefit of doing this in parallel and without needing to have an entire browser open just for this.
I think the actual useful use case will be filling forms, doing research on sites that need authentication, etc.
@badlogicgames Extracting text out of articles has been a solved problem for years now. That's why reading mode on almost all browsers works fine, and I think I can trust OpenAI to get this right.
The major issue for me, at least, is the LLM itself hallucinating.
@badlogicgames >With e.g. ChatGPT research mode, you have no idea what's being put into the context.
I can see what sources they are referencing, and I can check myself. That's something I already do to confirm the main parts if I care about the output. How does this extension do better here?
@vitbramm@OfficialLoganK@GoogleAIStudio For me, the only thing that solved it for me is to limit the output size, and have a post-processing step to retry when I get these cases, which is somewhat easy to catch with regex.
@KhalidWarsa Not as expensive as you might think, I would guess between 50 and 100 USD, which is all considering a good deal.
Most likely it will be less than this.
@onlinedopamine@edrick_dch The only scenario I can see where it turns out good for everyone is if the next generation models are really good and small and the hw is extra efficient, so we can squeeze 2 OOM. So, an opus performance in 30b model next year, which doesn't sound that impossible.
@onlinedopamine@edrick_dch People are consuming 200$/day using Claude code. No way this is profitable for them. The heaviest users can go above 10k/m. There are very few people who are willing or even able to pay this. So even with 1k/m they will still be losing money.
@JustJake 3.7 is very good at green field projects or even screens; it wants to generate stuff.
3.5 is lazy, so it just usually does what is asked for and doesn't try to touch unrelated stuff, so it's very good at editing existing code.