Today we release DeepSeek-TNG R1T2 Chimera.
This new Chimera is a Tri-Mind Assembly-of-Experts model with three parents, namely R1-0528, R1 and V3-0324.
R1T2 operates at a sweet spot in intelligence vs. output token length. It appears to be...
* about 20% faster than R1, and more than twice as fast as R1-0528
* significantly more intelligent than R1 in benchmarks such as GPQA Diamond and AIME-24/25, albeit not quite on R1-0528 level
* much more intelligent than our first R1T Chimera, and also think-token consistent, which is a major improvement
We perceive it as generally well-behaved and a nice persona to talk to. The weights are on @huggingface under the MIT licence. We are looking forward to your experiments and feedback!
Thanks to @deepseek_ai for giving their models to the world, to @chutes_ai and @openrouter for hosting R1T, to @WolframRvnwlf for benchmarking it, to @xlr8harder for beta-testing the new Chimera, and to @natolambert for constructive discussions at @aiDotEngineer.
Interesting experiment from Germany @tngtech: a model merge of DeepSeek V3-0324 and R1. Not sure about the method of merging (expert mix? SLERP/DARE?) but if it achieves R1 results with fewer tokens, that's worth deploying. Anyone up to test this?
Today we release DeepSeek-R1T-Chimera, an open weights model adding R1 reasoning to @deepseek_ai V3-0324 with a novel construction method.
In benchmarks, it appears to be as smart as R1 but much faster, using 40% fewer output tokens.
The Chimera is a child LLM, using V3s shared experts augmented with a custom merge of R1s and V3s routed experts. It is not a finetune or distillation, but constructed from neural network parts of both parent MoE models.
A bit surprisingly, we did not detect defects of the hybrid child model. Instead, its reasoning and thinking processes appear to be more compact and orderly than the sometimes very long and wandering thoughts of the R1 parent model.
Model weights are on @huggingface, just a little late for #ICLR2025. Kudos to @deepseek_ai for V3 and R1!
Presenting Mixture of Tunable Experts (MoTE): Behavior Modification of DeepSeek-R1 at Inference Time - a method extending the MoE architecture of LLMs. By tuning 10 key experts, it enables meaningful and focused behavior changes on-the-fly. Mon 18h CET: https://t.co/5cFW6dkZ9X
Ich breche nun zwei meiner Twitter regeln. Ich erzähle keine aktuellen Ereignisse aus der Klinik und werde nicht konkret, damit ich meine Anonymität wahren kann.
Es gibt aber Gründe, die am Ende klar werden.
Morgens habe ich erfahren, dass die Kollegen des letzten Tages die ganze
The hardest thing about #tngbtd is deciding which of the missed talks I want to catch up with over lunch... I heard a lot of good things about the talk about the LEGO Glomar Explorer: https://t.co/iNcnulDF7i
Sind Sie auch schon in der "Event-Location" angekommen? Dann steht dem https://t.co/DwfKHlRU0f nichts mehr im Wege. Wie man sieht, warten wir auch schon ganz gespannt auf die abwechslungsreichen Vorträge aus IT und Technik. #tngbtd
Slack-Registrierung:
https://t.co/EmfWoiR43z
@christophstock Klar, viele Vorträge von vergangenen Bigtechdays findet man auch auf https://t.co/2y1accYLPb, aber live dabei sein ist doch nochmal was anderes.
Check out #tngbtd at https://t.co/nbVDxDRYv4, where fear of missing out is a thing of the past because you can simply listen to multiple tracks at the same time 🤯
TNG is pleased to announce our annual Big Techday conference. It will take place in virtual form on May 8th, 2020. Check our website for more information on the event, the speakers and the programme. We will keep you informed under #tngbtd as well.
https://t.co/UgbRqdFLON
Tomorrow, @tngtech will attempt to set a new record in speed building this Lego Star Destroyer with more than 20 of our employees. We will live-stream our Legothon via Youtube (I will post a link tomorrow) starting around ~6PM CEST.
I learned a lot today about why fonts behave the way they do: https://t.co/EzyeDEeB12
Shoutout to @WestonThayer5 for this very well-written article and also for being super helpful with some accessibility questions I had earlier today!
Laurent Jaffart @ #tngbtd: OneWeb is essentialy taking satellite manufacturing from "Haute couture" to "Prete-a-porter" by increasing the production scale to 900 units.