Strange quirk ive noticed about chinese LLMs: They dont detect sarcasm like the american LLMs do.
Can’t say stuff like “why not just delete my whole codebase while you’re at it” to Chinese LLMs. You have to be very very literal with models like Deepseek and GLM.
Not sure if thats a technical difference in how the models are trained or a cultural difference, but its very noticeable IMO.
Grok is by far the funniest AI model IMO. Good banter. Witty responses. Will keep the joke going.
I'm starting the conversion work from new DeepSeek v4 Flash checkpoint to GGUF. If the model is as good as it looks, I'll probably remove the GGLM 5.2 support from the system, since now we have a model that is best suited for local inference that is smaller. Feedbacks?
If they do this it’ll go down as one of the stupidest moves in history
You don’t restrict it, you loosen restrictions in our models to actually Be useful against those and encourage more and faster innovation for frontier models here
Here’s another example: Hugging Face tried using American frontier models to analyze an AI-powered cyber attack. But the guardrails blocked requests containing real exploit payloads so they switched to GLM 5.2 running locally. The guardrails actually impaired defensive security.
Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails.” There’s no reason to limit American models on tasks that Chinese models handle without issue. We’re only making ourselves less competitive.
Yang Zhilin, the founder of Moonshot AI, an AI startup that created Kimi K3 to overthrow Anthropic and OpenAI, recently gave a 40-minute masterclass on their progress.
This is the clearest explanation yet of the engineering behind China’s cost-efficient yet highly competitive AI models.
Coupled with the recent Kimi K3 launch (their new 2.8-trillion-parameter model rivaling top closed AI systems), you’ll never look at “Chinese cheap models” the same way again.
Instead of watching another series on Netflix today, watch this talk.
@InsiderPhD Wrote about the open weight problem as well a few weeks back. IMO, it’s the biggest issue that we have with agent security/trust with no obvious mitigation.
https://t.co/ERSH8NshPj
China really thought they were just going to monopolize open-source AI and it was going to be a cake-walk🤣.
Mira Murati is literally the one who built ChatGPT 4 from the ground-up before they even knew how to spell "LLM".
American open-source AI is SO. FUCKING. BACK.
Today, we are introducing Inkling.
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
https://t.co/Ghebq5mG30
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Thinking Machines has released Inkling, the new leading U.S. open weights model, debuting at 41 on the Artificial Analysis Intelligence Index
@thinkymachines has previously released research previews of models and this is their first production language model release. The model is 975B total parameters, has 41B active parameters, and accepts text, image, and audio input modalities. The model is accessible via Thinking Machines’ Tinker platform API (256K context window) and weights are available on HuggingFace (1M context window).
Key results:
➤ Inkling debuts at 41 on the Artificial Analysis Intelligence Index, making it the leading open weights release from a U.S. lab. Inkling scores 3 points higher on the Intelligence Index (41) than the previous leading U.S. open weights model, Nemotron 3 Ultra (38), and also beats Gemma 4 31B (29) and gpt-oss-120b (24)
➤ Inkling stands out on agentic performance. It scores higher than both Kimi K2.6 and DeepSeek v4 Flash on both GDPval-AA v2 and 𝜏³-Banking: Inkling scores an Elo of 1238 on GDPval-AA v2, higher than Kimi K2.6 (1190) and DeepSeek v4 Flash max (1189) and scores 24% on 𝜏³-Banking, higher than Kimi K2.6 (21%) and just above DeepSeek v4 Flash max (23%)
➤ Inkling is token efficient compared to open weights leaders. Inkling averages 25K output tokens per Intelligence Index task compared to 43K, 38K and 37K by GLM-5.2 (max), Kimi K2.6 and DeepSeek v4 Pro (max) respectively
➤ Inkling natively supports image and audio multimodal inputs, a key differentiator among open weights models. Inkling accepts text, image, and audio input modalities. Images and videos are encoded via a hierarchical patch encoder and audio via discrete token encoding, with all modalities projected into a shared hidden space and processed jointly by the decoder
Additional model details:
➤ Size: 975B (41B active) parameters
➤ Input modalities: Text, image, and audio (text output)
➤ Context window: 256K tokens on Tinker, open weights model supports 1M
➤ Pricing per 1M tokens (64K context window): $1.87 input / $0.374 cached / $4.68 output
➤ Pricing per 1M tokens (256K context window): $3.74 input / $0.748 cached / $9.36 output
pls stop making models and agents anthropomorphic
when it triggers the "this is a person" part of my brain I get unreasonably angry when it does something stupid that no person would do
e.g. claude
i rarely get angry at codex because it's much more mechanical and doesn't mess with the emotive part of your brain
First time Fable launched, I got a fever that morning and only got to use it for like an hour. It was gone by the time I was better. After using it for a full day of some complex tasks I’d been avoiding: all I can say is “WOW”.
BREAKING: Gemini Omni Flash by @GoogleDeepMind is 1st overall on Video Arena with an Elo of 1404.
Gemini Omni Flash establishes a 101 point Elo gap over Seedance 2.0 Mini by @BytePlusGlobal in 2nd place, one of the largest leaps we’ve ever seen on Video Arena.
This establishes Google as the world’s leading video generation lab, with a leap of 7 positions from their Veo series.
Congratulations to the @GoogleDeepMind team on this accomplishment!
@steipete Would you be willing to add this pipeline as an example to the documentation for https://t.co/7UDLZhOIS1? Would love a great example of how to setup a big workflow using summarize for extracting high-value knowledge from long-form videos.
We’ve received notice that the Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5.
We'll begin restoring access tomorrow, and will share an update soon.
We’re grateful to our users for their patience, and to everyone who worked with us on redeploying the models.