Heidi Reichinnek bricht Interview ab, weil sie keine Antworten auf die Frage der Speicherung von erneuerbaren Energien hat.
Danach fordert sie den Journalisten auf, den Teil rauszuschneiden und der Pressesprecher greift ein.
Das ist die Fraktionsvorsitzende der Linken. 😂
@kimmonismus What if those very responsible labs internally dont slow “for research purposes” and employ ASI to push themselves to the top of everything? What if they already have models that are better in economic decisions than most CEOs?
Anyone else using DeepSeek V4.1 Flash having repeated "Let me go." self-reassurances in the reasoning stream?
Is this what model start thinking when you take "Wait, " from them? Eventually it goes though 😄
using @pidotdev
@ivanfioravanti Normally a closing </think> at end of your prompt, but its more of a dirty workaround. Depending on the chat template, you can also send a no think param. If you want less thinking, but thinking enabled, it depends if your server supports thinking budget.
@superalesha@NVIDIAAI More like “why compute attention on less important token sequences like ‘and the’ where a less capable attention mechanism could augment similar or same kv” Not sure if that makes sense 😅
At this point I am pretty convinced, that DeepSeek V4 Pro will push the frontier, if not on raw capability, then on parameter count to capability ratio. mHC might be a bigger win at scale than many thought... If the same architecture and 1,6T parameters is able to deliver, what others only can do with 6-10T - then wait for the 10T run of DeepSeek 🤠
Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows.
Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category, and is designed to run entirely on consumer hardware like a Mac or PCs with performant GPUs.
In keeping with our long tradition of sharing fundamental AI research, we’re releasing model weights under a permissive Apache 2.0 license.
🧵👇