🚨 Para ajudar o querido Lito Sousa, compartilhe o reel do link abaixo em suas redes sociais.
🔗 https://t.co/p5sArlTs0B
Além disso, marque as instituições abaixo nas publicações:
📌 @ionis_pharma — Ionis Pharmaceuticals �� ION717 trial
📌 @broadinstitute — PRiSM trial (NCT07444580) 📌 @harvard
📌 @mayoclinic
@trikcode It seems like the Claude team is so focused on loops that you need it for Opus 5 to be good. I have Fable use Opus for me, I always ask it to do adversarial reviews and it *always* finds missing things that then get corrected.
For vision, I still haven’t found anything to be anywhere near Gemini and 3.7 Flash continues to do well. On a real project of a legal pipeline where I need vision to see things in contracts (literally pictures of document), I’ve tried all models and none of them have the precision of Gemini and now give the cost of 3.7, it’s S+ for that very specific use case. I know about the “small and cheap” OCR, those unfortunately also can’t get the results needed for this type of pipeline.
@pvncher Thank you! I’m sorry for the bother over a weekend and I really appreciate the drive to be answering and caring for users/customers even over a weekend.
I’ve had Max20 or whatever is called subscription for both Codex and Claude acode for many months. I have a single one for each, I’m not a tokenmaxxer, I use them professionally in multiple parallel projects. I was getting to a point where Codex just felt unlimited and 5.6 did everything I wanted…the last two weeks I actually need Claude Code for longer running tasks as it’s the subscription that actually lasts longer and that’s just insane, it’s literally the first time ever this has happened.
I shared feedback with the team and @pvncher mentioned I shouldn’t be using 5.6-sol-max so I’ve switched to high or medium, generally asking it to orchestrate Terra and lower thinking levels of Sol and yet my subscription is draining in a speed I’ve never seen.
It’s a shame @cursor_ai doesn’t have the “just works on mobile” experience of Codex, I was daily driving it for a few days and I’m really happy with Grok 4.6. I just miss the seamless “use from your phone” and browser use of Codex…funny enough, it’s the first time the harness is the actual thing making me continue to use 5.6. Can’t wait for @t3dotcodes to have a good Cursor experience and that OSS project might close the gap Cursor didn’t. I know cursor has a “remote” feature and I’m really trying to use it and it doesn’t just works.
It may be a skill issue, but I’m not super thrilled with the experience of moving all of the compute of my dev workflows to cursor cloud.
So Codex is still the best to touch grass…as I take a quick trip and for the first time didn’t even bring my laptop…just hope my Codex sub doesn’t end in the process despite the regular workload :( and thanks Claude for doing the big code review and refactoring…it honestly feels like I have more Fable than 5.6-Sol-xhigh or max.
@thsottiaux I’ve shared feedback ids and happy to share more if that helps investigate what’s going on if it’s really a technical issue and not a change in the sub.
Você assistiu o vídeo? Ele explica os motivos lá, não é uma comparação direta e linear de “este é melhor que aquele” mas sim o uso dos modelos para use cases onde são bons. DeepSeek funciona absurdamente bem no custo benefício para as tarefas onde ele é bom. E update: versão multimodal está na API do DeepSeek em testes.
O Luna e seu custo se tornaram o modelo padrão que posso colocar em produção com clientes. Projetos que seriam literalmente inviáveis financeiramente, agora estou entregando graças ao DeepSeek V4 Flash e o Luna.
@EthanLipnik They get a nice thing done and people promote proudly tips and tricks how to abuse it :( maybe we need to get scores as customers and those that keep doing this kind of stuff just lose their score and can’t get nice features
1 - I actually barely did any of that since you told to stop with Max, I was doing regular “review this codebase”, “fix the issues you’ve found” when Sol High.
2 - I’m not. I generally let it do its thing.
3 - I’ll do another check, I’m very light on skills or anything extra, I run it barebones and it’s one of the things I love about codex, it’s the one that is the best at browseruse barebones.
@pvncher Big fan of your work for a while now, I just had Sol Max do poorly on orchestrating an implementation, I hope there is good insight and learnings there, I sent the feedback: 01a01d15-c2f7-7812-b215-6ac380bdd75a
TLDR; Had I not stopped it, the entire Max 20 quota would be spent without producing the final output, it got circular with the reviews of the loop.
It’s hard to believe something like this is available to all of us and free: https://t.co/U9dSDwpGes amazing interview @mattpocockuk , thank you for the wisdom @unclebobmartin
@Tech2Wild Run real work and pipelines with them, I believe DSV4 Flash still comes out ahead. At least that's what I have as a result from a production legal pipeline I have tested with.
Models and benchmarks continue to be interesting...Grok 4.6 and Gemini 3.7 Flash seem to be doing much better and yet I just got a RunPod box with an RTX PRO 6000 so I could test Qwen 3.8 27B with beefy hardware.
Started with Gemini, asked it to learn the best recipe to get this working and gave it ssh of the pod...it just went nuts, kept looping and installing different versions of dependencies...I had to kill the machine
New one...Grok, starts MUCH better, but still installs the wrong version of vllm that doesn't actually work well with this card, doesn't test it's own work and keeps telling me its fine but any prompt kept crashing the server
Then I just gave it to glm-5.2, night and day difference on how it "operates", very detailed, started figuring our all the mistakes made by Grok, got the right versions setup and running, tested the work properly...when it told me it was done, it was just working. I told it to make it go faster, it found and tweaked the vllm config to make it go faster. I was able to push a lot through the server with no issues at all.
Then I actually had Qwen 3.8 27B run a legal pipeline I'm testing so I could see how it would do...it started burning tokens like crazy and couldn't get it done, kept getting incorrect results and just going nuts.
My current production pipeline that works just fine there...Luna, uses very few tokens, gets it done super fast.
Yet, all of the models I just mentioned here benchmark similarly...benchmarking is super hard.
Really trying to use and like Gemini 3.7 Flash and anti gravity...trying to use it on a weekend, not peak hours...took a bunch of pictures of my 10 year old son textbook so I can create a study guide style site...so this is not complex, at all....and yet I can't get to the finish line without errors like this...the model is fast, the result isn't :( I would be done by now with Codex. Really considering dropping the Ultra sub :(
I literally have it running a production workload where it’s doing a bunch of validation on legal contracts. Following your recipe (thank you). I capped it at 300k for context.
“Current production performance
* 301 requests completed since the previous snapshot.
* 298 normal completions; 3 reached their output-token limit.
* 0 request errors, aborts, OOMs, preemptions, or container restarts.
* Average aggregate generation: 27.3 tokens/s
* Median: 6.7 tokens/s
* P95: 82.2 tokens/s
* Peak: 99.6 tokens/s
* Average TTFT: approximately 27.8 seconds
* Average end-to-end request time: approximately 84.8 seconds
* Request-weighted TPOT: 51.6 ms/token, equivalent to roughly 19.4 tokens/s per request
The large average prompts explain much of this: approximately 93,000 prompt tokens per completed request, with about 81% served from the prefix cache. Outputs average roughly 960 tokens.
Concurrency
Yes, concurrency is being used:
* Three requests running: 72% of active samples
* Two or more running: 93%
* Average running concurrency: 2.65
* Maximum outstanding requests observed: exactly 3”