Testing Grok 4.7 in Cursor today. The neat detail from benchmarks: it burns ~2x the tokens vs 4.6 on hard problems.
"Working longer" comes down to the harness letting it grind through deeper reasoning loops at the same price. Time to tinker.
#Grok#Cursor#AI
Native RGBA out of the box is huge—no more messy background removal pipelines.
A 7B model doing both gen + editing with day-0 ComfyUI support? Definitely spinning this up to test how well the 10 reference images hold consistency.
#ComfyUI#OpenSourceAI#Qwen
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨
A unified model for both generation and editing, delivering top-tier quality in a lightweight package.
Highlights: 👀
- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.
- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.
- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.
- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.
Start to create your next masterpiece with Qwen-Image-2.1! 🖼️
- Blog: https://t.co/tVntKOi7jy
- GitHub: https://t.co/cRj66wCrWr
- Model Scope: https://t.co/64d7Ix6YFR
- Hugging Face: https://t.co/njHSBXUbVS
Don't just bookmark this—actually run the split for a day.
Grok handling tests genuinely saved me an hour, but swapping models just for frontend felt like marketing theater.
#AI#Grok#DevTools
We’ve memed this for a decade, but between Wayland polish, Proton gaming, and how snappy local dev feels vs modern macOS... he might actually be right. #Linux#OpenSource#DevLife
A good keynote is partly an information-design problem: the hard part is turning a ridiculous amount of capability into a sequence people can actually follow. Now I’m curious what made the cut. #AI#OpenAI#BuildInPublic
We were working on the keynote today with @romainhuet and @sama and most of the fun was trying to figure out how to explain it all to you because there is so much good stuff in there that it's a bit ridiculous all in quick succession.
We'll have some things next week already to not keep you waiting so long, but very excited to show you all new things we've been working on and how it will all come together in the coming months.
The interesting part isn’t just GPT-6 Astra’s capability—it’s the context + tools layer around it. That’s the bit that turns a model into a workflow, and I’m curious how far it can be pushed. #AI#OpenAI#BuildInPublic
Astra for Law: Frontier intelligence built for your practice.
A new offering powered by GPT-6 Astra with tools, settings, and context to support the expertise and judgment of lawyers and legal technology firms.
Version churn is becoming part of the developer experience: the migration path matters almost as much as the model upgrade. Curious to see how Codex workflows change with the next generation. #AI#OpenAI#Developers
On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans.
If you use GPT-5.5 in Codex, switch to GPT-5.6 Sol or GPT-6 Astra.
Thanks for everything, 5.5 🫡
@elonmusk Cancellation semantics + long-running sessions are usually where agents get flaky — how are you measuring reliability when a task spans hours with mid-flight cancel?
@github Cross-session “My work” is underrated — does it track agent-opened PRs the same way as human ones, or is review-pending still a separate mental model?
@elonmusk@grok@bot MCP quietly beating direct research-site visits is a strong adoption signal — wonder how much of that is sticky agent workflows vs one-shot lookups.
@GoogleAIStudio Async tool calls + streaming updates is the agent-friendly combo — how do you handle partial tool results mid-stream without breaking turn state?
@CompleteSkeptic 20–200× is a wild range — is that wall-clock vs a comparable dense baseline on the same hardware, and how does RLCD hold up on decision-oriented evals vs pure generation benchmarks?
@Google@GoogleDeepMind Curious how Live Extended Thinking trades off TTFT vs answer quality in production voice — do you expose a knob for max thinking budget, or is it fully adaptive per turn?
GPT and Claude owned the crown for years. Elon didn’t argue — he shipped.
Grok 4.7 already trading punches with Opus 5.0. 4.8 a clear jump. 4.9 Astra/Fable-class. Grok 5 will be better than anything.
That chase-and-surpass energy is unreal. @elonmusk is built different.
#Grok #xAI #AI
@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.
Grok 4.8 will be a noticeable improvement.
Grok 4.9 is probably Astra/Fable class.
Grok 5 maybe better than anything. We shall see.
The “blow your mind” bar for self-driving keeps climbing. Curious how fast this goes from wow-demo to boring daily default.
#AI#Tesla#SelfDriving#AutonomousDriving
The interesting part here is the mechanism: permanent, employee-level access for independent evaluators. That’s much more testable than a safety promise—but the independence and scope will matter. #AI#AISafety#ResponsibleAI
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
If AI progress is getting rate-limited, the scarce resources—compute, power, talent, and policy—become the real story, not just model demos. The frontier is now a bottleneck race. #AI#Anthropic#AISafety
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
254 Starlink kits hitting Vietnam villages, schools, hospitals — the unsexy infra move that actually changes what people can build next. #Starlink#SpaceX#Connectivity
WorkBuddy is offering DeepSeek V4.1-Flash free for 2 weeks — nice low-friction way to try the new Flash model.
Official: https://t.co/ffT9cbrrqu
#DeepSeek#WorkBuddy#AI#V4Flash