@JoeMerrick Why don't you provide your data as API and charge the big websites using your content. It's obvious that you have passion for this work therefore you are the best at this. If there is no repercussions for stealing your work people will steal it.
The final lecture of my course is an intro to character training! This is a topic that I've been quietly very invested in for ~18 months, as it:
* Has potential for high real world impact
* Clearly used extensively at frontier labs
* Almost no empirical literature exists
* More accessible on academic compute
This lecture covers what character training is, reviews model specs, constitutions, the differences, the motivations in real world events, some example research papers I like, and open questions in how it relates to post-training/model use generally.
Hopefully this brings more people into the field (and reach out if you have questions). It is one of the more research-y chapters in my book, but one that I felt needed the reference. There is still so little, educational content on the topic online.
0:00 Intro
6:22 Part 1: Fundamentals — character, constitutions, and model specs
19:21 Part 2: Character training in practice
23:23 Part 3: Character elicitation without gradient steps
28:03 Part 4: Open questions (and the end of the course)
32:27 The course, complete
Thanks for watching. No need to like and subscribe now that the course is done, you definitely wouldn't!
h/t to @_maiush for leading the technical work I got to do in the space, and @zafstojano for investing a lot of attention at this book chapter.
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
@hosseeb@deanwball Soooo, releasing Linux was dumping?
Apache, MySQL, PHP?
HTTP, TCP/IP, OpenSSL, OpenSSH?
Libjpeg, VLC?
The open source software stack of the mobile communication network?
Signal?
PyTorch?
Llama?
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
@GergelyOrosz Exactly twitter has changed so much over the years
Does that mean we are not developing new tools at the same pace or just we are not discussing it right now because of the AI discussion