@teortaxesTex I don't know but it seems to be Alibaba Cloud stress testing their newly delivered Atlas 950 SuperPoDs, the 100t is likely calculated on compute bases and not memory
The model is likely 80b8a or 122b10a or something similar
Ox Alpha is likely a Qwen MoE model and will debut next week, 100T daily token capacity goes with the model being MoE and backed by large infra (Alibaba Cloud)
Ox Alpha (stealth model) is free for the next week
- 1M Context
- Multi-modal
- Zero Data Retention
Generous rate limits, near unlimited usage
We have capacity for 100T tokens per day, lets see what you can do
@bnjmn_marie I would say purpose of benchmark is to also measure adaptability, so 76.2% would be actual score, but in a real-world task the user would likely install Numba to help it more forward like you did.
DeepSeek v4 Flash 0731
- Scores 52 in AA
- 13b active params
Qwen 3.8 27b
- Scores 52 in AA
- 27b active params
Does that mean future Qwen 27b still has more to give probably even Fable 5 level score ...
@scaling01 here Qwen 3.8 27b zero shot the same prompt, and that with UD-Q4_K_XL quant not even full precision, it didn't work for you maybe because of the harness you used
literally just tried to let qwen3.8 27b one shot a fluid simulation webapp
it thought for 40k tokens, took around 1 hour to generate and it's just black and doesn't work
meanwhile Opus 4.5 just oneshots it in a minute and works wonderfully
exact same prompt
A month ago, we asked for prompts that models still struggled with. We received many thoughtful examples.
Today, GLM-5.3 reaches 60 on the Artificial Analysis Intelligence Index, with 743B base.
Thank you to everyone who contributed.
defenders can see the future, and have a narrow window to uplevel their cybersecurity practices now.
key is to uplevel fundamentals and apply the best AI tools.
what weโre doing at OpenAI, and where other organizations can start: https://t.co/P3IMZkV234
Stripe has finalized an agreement to acquire OpenRouter for more than $7 billion, according to a report from Bloomberg. This is more than five times the $1.3 billion valuation from OpenRouter's funding round just 82 days ago.
@addyosmani Like Patrick, I too don't get the terminal-based coding harnesses, I quite liked Roo code but that is only single session, so I built https://t.co/fcNfCFwxOn , it is multi session, multi project, multi repo, I find it help with focus on task instead of waiting for LLM response.
Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows.
Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category, and is designed to run entirely on consumer hardware like a Mac or PCs with performant GPUs.
In keeping with our long tradition of sharing fundamental AI research, weโre releasing model weights under a permissive Apache 2.0 license.
๐งต๐