Welcome to the CFIC, Lujiazui, Shanghai.
There is a 4K UHD Studio on the top of the building.
We open to all the Broadcasters.
#broadcast#broadcaster#television#business
When SceniX joined World Labs, we said spatial intelligence was never only about perceiving and generating virtual and physical worlds, but also interacting with them. Today, we’re sharing early results from that vision: building worlds that train robots. 🌎🤖↓
Across eight case studies spanning industry and academia, we explore what this shift means for scientific computing—and why human verification, stewardship, and long-term maintenance matter. https://t.co/LmZurP9NME
The Kimi K3 architecture figure for yesterday's big open-weight model release, along with some observations and thoughts.
1. Yes, it looks relatively complicated, but it's essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now)
2. The one new component compared to Kimi Linear is the LatentMoE. I omitted it in the figure below since it's already very crowded, but that's essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention.
3. Kimi K3's overall trend (similar to Nemotron 3, DeepSeek V4, and others) is also towards better inference efficiency. That is, there are many components that replace existing components with efficiency-tweaked versions. I.e., MoE -> LatentMoE, regular attention -> multi-head latent attention and Kimi Delta Attention. (I also have short tutorials and write-ups in my gallery if you are curious about additional details).
4. The one component change that is not an efficiency tweak is attention residuals. Like DeepSeek V4 improved the residual path with mHC (manifold-constrained Hyper-Connections), attention residuals are a way to improve the residual path, but it works a bit differently. I.e., mHC made the residual path wider. Attention residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an important/contribution weight. According to the report, it improves the validation loss and downstream performance (a bit) consistently and adds about 4% in training cost and 2% in inference cost.
5. Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead. (Again, this is inherited from Kimi Linear). In other architectures, the recent trend was towards RoPE in local attention layers (like sliding window attention) and NoPE in the global layers. There were a few architectures that only used NoPE everywhere, but this is the first frontier-level one as far as I know.
6. Kimi K3 now also has native multimodal support, which is great!
There are several other interesting training tidbits in the technical report, but that's it from the architecture front so far. A really great release overall.
MLB at Field of Dreams is one of those once-in-a-lifetime events to work. That fact is not lost of @MLBONFOX game director Matt Gangl (@a1director).
What are his favorite angles to cut to in this unique venue and why does he feel its been such a success with viewers?
"This takes us back to our roots, back to our Little League days."
Hall of Fame @Reds catcher @johnnybench_5 joins our broadcast booth to talk about his experience at the Field of Dreams.
📺: FOX and the FOX Sports App ➡️ https://t.co/5cfRya3sXd
As many as 828 million people worldwide are hungry.
@FAO helps put this staggering number into perspective & explains what can be done to help fill empty plates. https://t.co/B7EMnar7o2
How can businesses make sense of the new global reality and chart ideal paths toward a better future? Follow these six Signals. (Paid post for @Accenture) #ad https://t.co/drhLV1ze8x
Tesla is an AI company.
“This is what @Tesla Autopilot sees using neural networks that take 70,000 GPU hours to train and output 1,000 tensors (predictions) at each timestep.”
Dogecoin — it all started as a joke. So why would someone buy this? @Kr00ney breaks down everything you need to know about the meme-inspired digital currency and its meteoric rise.