This post made me curious: what if “time to input box” was separate from app startup?
I built a small PTY prototype that shows an editable input shell immediately while Claude Code continues initializing behind it.
Across 200 interleaved launches, median time to editable input improved by ~99%: from 812ms to 5ms.
For that brief transition, it’s plain text. / commands, @, and ! gain their special behavior once the full UI loads. Submission still waits for initialization to finish, but you can use that half-second to start typing.
🚨 Cursor just got a big win from Anthropic. 🤩⚡
OpenAI is pulling future models from Cursor, while Anthropic is moving in the opposite direction.
Anthropic says it plans to "increase compute for Claude models on Cursor" 😂
That means more Claude capacity right as OpenAI is stepping back.
And with Anthropic’s big plans for next week…
The timing couldn’t be better. 👀
South Korea is committing KRW 2.3T ($1.66B) to build a sovereign, full-stack humanoid robotics industry by 2030.
The state-led strategy to challenge the US & China:
✦︎ Boost core component localization from 45% to 80%
✦︎ Pivot legacy auto parts suppliers to precision robotics
✦︎ Direct state purchase of 1,080 humanoids & 5,000 autonomous units
✦︎ Build unified "cleanroom" data libraries for training
✦︎ Deliver a sovereign Robot OS by 2031
Full breakdown on how Seoul plans to scale physical AI:
https://t.co/qPc7PppQSL
Who’s buying made in China cars?
Pretty much people in every time zone. More than 100 different countries.
This year European demand has exploded higher.
Open weights do not make models honest, but they give developers more control over versions, fine-tuning, deployment and verification safeguards. 🔓🧩
#ArtificialIntelligence#Technology#AINews#AI
https://t.co/dbaLX1oPg3
this is a massive blow to Cursor
opus 5 is trash; fable 5 is solid but will bankrupt you, and using the same claude model (with the same reasoning) is night-and-day better in claude code than cursor
also makes you wonder if openai now sees SpaceXAI as a real threat
One plausible future is not enough
Video generators often look convincing in a single rollout, but PAWBench reveals they fail to reproduce the correct distribution of possible outcomes across repeated samples.
Cohere Parse 5 scores 79.2 on ParseBench — ahead of Mistral OCR 4, Databricks AI Parse, and Azure Document Intelligence.
Best price-performance on the market for high-volume document parsing.
The enterprise data pipeline just got a new default.
Most AI benchmarks test whether a model can give the right answer. CommerceAgentBench asks whether an agent can actually finish the work.
Accio has open-sourced 107 e-commerce tasks across procurement, product listings, operations, fulfillment and after-sales. Agents work across browsers, email, calendars, documents, APIs and files.
Crucially, they are not graded on what they claim to have done. The benchmark verifies what they actually changed, saved or submitted.
Take the Gmail procurement case. The agent must search roughly 300 messy emails, identify the real suppliers, reconstruct the latest quotes, compare six Incoterms and four currencies, calculate landed costs and detect payment fraud.
Then it must choose a supplier, label the relevant emails, save a reply draft and create a kickoff event.
This is the kind of benchmark I find genuinely useful. It measures agents more like workers than chatbots. And the results show why human oversight still matters: the best observed run completed only 66 of 107 tasks, a 61.7% pass rate. And since 2026 is literally the year of agents, this is more important than ever.
Accio says the tasks draw on 10M SMB users, 1.6M conversations, 200K agent trajectories and Alibaba’s 27 years of e-commerce experience.
The project and task specifications are open source:
https://t.co/Tq4UKfJDBb
🚨Hy4 Preview looks absolutely cracked
It only activates 49B params, and is basically GLM 5.3 / Kimi K3 level on benchmarks, maybe ahead slightly
And this is still only the Preview model too
I genuinely didn’t expect Hy4 to be this competitive with the frontier models already tbf
being a taxi driver is peak rn
*630am uber to airport*
- dude helps put bags in the back
- autopilot drives us there
- dude helps me get bags out
golden era for taxi drivers lol
1/ Self-improving agents refine answers, not the process that produces them. Meta^n keeps one meta-operation Ω and recurses on its input: each layer writes code that rewires the layer below. Depth emerges, and so do roles.
📄 https://t.co/J0QEBcrCUy
💻 https://t.co/f1EJnCIe7a
🚨GLM 5.3 goes open weights tomorrow
Zai said the full model weights would drop two weeks after launch once safety testing was finished
This is the proper GLM 5.3 too, not Flash
Open models are about to get ridiculously interesting over the next couple months
VoiceMem
ByteDance's new memory system for real-time voice AI. Uses a dual-brain architecture: left brain for facts, right brain for emotions. Hits 91.2% on LoCoMo with just 5 memories, and retrieves in 134ms.
Its official now: Ox Alpha is GLM(-5.3 Flash) by zAI. Offical release tonight via zAI
Via Bloomberg
"The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News."
Ox Alpha is now so popular it beats Claude Opus 5's monthly usage in just four days with double the usage on openrouter. crazy part is it’s still free to use for everyone
OpenAI Catches Kimi K3 in Pricing
Model | Input 1M Tokens | Output 1M Tokens
Terra-5.6 | $2 | $12
Kimi K3 | $3 | $15
Sol-5.6 Standard | $4 | $20
Sol-5.6 Batch | $2 | $10
Sol-5.6 Flex | $2 | $10
ChatGPT Terra-5.6 has similar performance to Kimi K3 with Kimi K3 winning on some benchmarks while Terra-5.6 winning on others. For example, on DeepSWE 1.1 Terra-5.6 is at 70.0% while Kimi K3 is at 69.0%.
Standard Sol-5.6 is compatible in price to Kimi K3. For slightly slower latencies, Sol-5.6 is cheaper than Kimi K3! Batch mode has deterministically slower latency while flex may be faster or slower than batch depending on compute availability.
🚨Qwen might have two unreleased models in Arena
A model codenamed "Paloma" looks like it could be a Qwen model
Another unreleased model codenamed "Korrine" is likely a Qwen model too
Paloma looks better than Korrine from first impressions
GPT-5.6 Sol Max was already the better deal on DeepSWE v1.1 - The recent price cut widens the gap even further. 🔥
Sol scores 72.7% at $6.47/task, compared with Fable 5 Max at 69.7% and $21.63/task.
DeepSWE tests coding agents on 113 original, long-horizon engineering tasks.