Поработал я тут с зумерами. Вот эти вот мужчины и женщины, 2004 года выпуска, с паспортами и дипломами. К счастью, недолго.
Что могу сказать, как опытный, брюзжащий старпёр, о качестве рабочих процессов.
Всё общение ведётся в мессенджерах. Туда генерится хуева туча контента.
Одни с помощью ChatGPT генерят ТЗ, документацию и всякие «концепции». Вторая сторона берёт это, загоняет обратно в ChatGPT и генерит свои вопросы. ChatGPT первых генерит ответы. Все письма, типа «деловая переписка», договора и прочее делаются примерно по такому же принципу.
В итоге туда-сюда гоняется туева хуча информации.
Только есть одно «но».
Никто нихуя не понимает, что происходит.
Ловишь такого за патлы в коридоре:
— Ты понимаешь, чё ты сейчас вообще написал? Объясни вот это предложение.
Оно делает круглые глаза, убегает в переговорку и там втихую просит ChatGPT объяснить ему смысл фразы, которую сам же пять минут назад вкинул в общий чат.
Я тексты ИИшки скоро уже буду отличать с закрытыми глазами. Пока, правда, только с открытыми получается.
И вот тут, на мой взгляд, начинается действительно интересная проблема.
Проблема не в том, что люди используют ИИ.
Наоборот. Нормальный специалист с ИИ становится сильнее.
Плохой специалист использует ИИ совершенно иначе.
Раньше было:
Я думаю → формулирую → проверяю → отправляю.
Теперь всё чаще:
ИИ думает → я копирую → мне отвечают → я скармливаю ответ ИИ → копирую следующий текст → отправляю.
Получается какая-то цифровая игра в испорченный телефон, где каждый участник ещё и подключил своего робота-переводчика.
И самое неприятное — внешне всё выглядит прекрасно.
Тексты структурированные.
Абзацы красивые.
Заголовки.
Маркированные списки.
«В рамках дальнейшего взаимодействия».
«С учётом вышеизложенного».
«Предлагается рассмотреть возможность».
Бурная деятельность кипит.
Только если остановить процесс и спросить:
— А что конкретно мы сейчас решили?
Тишина.
— А зачем мы это делаем?
Ещё тише.
— А что означает вот этот пункт?
«Секунду, я сейчас уточню».
У ChatGPT.
И вот это уже не проблема зумеров. Это потенциально проблема вообще всех, кто начнёт использовать ИИ как замену собственному мышлению.
Потому что навык, который раньше приходилось тренировать просто для того, чтобы нормально работать, теперь можно почти полностью делегировать машине.
Сформулировать мысль.
Объяснить.
Составить вопрос.
Ответить на вопрос.
Проверить текст.
Написать договор.
Объяснить договор.
А потом человек уже не понимает, что именно он подписал и зачем.
ИИ должен ускорять мышление.
А не становиться прокладкой между двумя людьми, которые сами уже не понимают друг друга.
Иначе мы получим совершенно новый вид корпоративной эффективности:
тонны сообщений, сотни страниц документации, десятки созвонов — и ноль людей, которые реально понимают, что происходит.
Главный тест работы с ИИ, по-моему, очень простой:
Если ты не можешь своими словами объяснить человеку содержание текста, который только что отправил от своего имени, — значит, ты его не написал. И, возможно, вообще не понял.
А дальше вопрос уже не в искусственном интеллекте.
Вопрос в том, сколько естественного интеллекта мы готовы ему делегировать.
My OpenRSI (recursive self-improvement) agent has been running fully autonomously for 24h and it's still landing win after win on https://t.co/axIq66WbuG by @eigenlabs@poolsideai.
Recursive self-improvement is now fully open. Link below 👇
This is one of the papers I'm quite excited about in the past few weeks. It's a very simple but practical modification to the DINOv3 training framework.
Let me explain how it works.
A super long overdue (3+ years?) post on scaling laws.
Compute is expensive. Scaling laws are a way to help us reason about the optimal compute allocation between data and model size before committing to a large run.
The post covers what scaling laws predict, how compute-optimal allocation works, why Kaplan et al. and Chinchilla disagree, and how data limits + fitting details make extrapolation tricky.
https://t.co/HP26eJvjHB
We are releasing our first quantized checkpoints for the Qwen3.5 series of models, co-designed jointly with our inference engine to achieve maximum possible performance on Apple hardware
Starting from 0.8B, 2B and 4B models
https://t.co/2R8BdhAfzv
Alright, it's time for a paper thread about my own first ever vision paper, which is having a bit of a moment on twitter rn thanks to @PINTO03091 and @yacineMTB.
BiternionNets: continuous head orientation from discrete labels.
Demo video from ~11y ago:
Introducing TIPS v2
👀Foundational text-image encoder
📸Can be used as the base for different multimodal applications
🤗Apache 2.0
🧑🍳New pre-training recipes
Happy that SigLino is a #CVPR2026 Highlight. It started as AMoE, focused purely on efficient MoE distillation (loss, data, multi-res management), and it is now a full series of Agglomerative ViTs (dense and MoE, from 30m to 0.6B params) distilled from SigLIP2 and DINOv3.
We used the AMoE variant to initialize the vision experts of an early-fusion grounding MoE with modality-specific experts and show that it is a strong baseline on the small-scale training data regime on the refcoco benchmarks. Later, we figured that full early-fusion with a dense model works well, and even better, which led to Falcon Perception.
Models: https://t.co/yGHO6KD9UJ
Paper: https://t.co/TbNQ4eeQTy
Code: https://t.co/gklCUSbTKv
https://t.co/cavyHJeWaI
With @NarayanSanath@dahou_yasser@lkhphuc@griffintaur@HildeKuehne@hhacid
🚀 Today we’re releasing FlashOptim: better implementations of Adam, SGD, etc, that compute the same updates but save tons of memory. You can use it right now via `pip install flashoptim`. 🚀
https://t.co/nRrLSpjnwV
A bunch of cool ideas make this possible: [1/n]
Watched the Physics of LMs Part 4.1 video by @ZeyuanAllenZhu today and pulled out a bunch of actionable takeaways.
Each bullet has a YouTube timestamp so you can jump straight to the exact segment.
Introducing Canon
1/ Attention is known to do associative recall, but it needs at least two attention layers to learn efficiently: https://t.co/CqnN0FvKu2
2/ Canon or Cannon layers: enabling horizontal information flow without paying the full attention cost: https://t.co/vJgQbZMdKO
3/ Why Canon can improve performance: backward feature correction. Multi hop reasoning is like learning a ladder: learn k minus 1 hop first, then climb toward k hop: https://t.co/eJlCnIojT4
3.1/ Layer wise training fails: Having 10 layers trained for k - 1 hop tasks and then frozen, then adding two more layers to train for k hop won’t succeed. We need to train all layers together so earlier representations can be corrected by later signals: https://t.co/eJlCnIojT4
3.2/ Residual links as the mechanism that makes backward feature correction possible: https://t.co/PkRo0rmto9
Position Encoding and Canon
1/ Canon turns NoPE into strong performers, sometimes surpassing RoPE plus Canon on certain "short context" tasks: https://t.co/12E0pR07Te
2/ Short context tasks: RoPE can hurt performance, and Canon lets you reduce RoPE usage while keeping length generalization: https://t.co/Cq7pr5OrhW
Example recipe: 1/4 heads RoPE + 3/4 heads NoPE + Canon
Ablations about where to add Canon layers
https://t.co/YeqeANkp0c
1/ Residual Canon h’ = h + Canon is crucial. Other tricks such as adding activations is unnecessary.
2/ Any single Canon gives a big gain, more Canon often gives more gain
3/ Canon is an architecture module, not tied to Attention or MLP specifically
https://t.co/oRyCARk9Wx
Using Synthetic Playground you can find:
1/ relu^2 slightly helps reasoning on plain MLP, but slightly hurts reasoning on gated MLP: https://t.co/nD9fPPqfpH
2/ MLP -> gated MLP improves reasoning but hurts knowledge capacity, and Canon can help recover capacity by speeding up learning: https://t.co/bwWwBPaZmE
3/ MLP → MoE improves inference speed but hurts knowledge capacity. Adding Cannon-ABC improves the capacity: https://t.co/cg2nO1CG2R
Linear models and horizontal information flow
1/ Linear attention underperforms broadly except knowledge capacity: https://t.co/nyd7jtSFIm
2/ Canon lifts linear attention to match stronger baselines and can outperform Mamba2 in some settings: https://t.co/ZqOymqN12Z
3/ Conv1d is really important for Mamba2, removing it hurts performance a lot; adding Cannon improves the performance a lot: https://t.co/4S26MTslQZ
4/ Conv1d is also important for GDN; adding Cannon further improves GDN’s performance: https://t.co/aVvtyTiSap
5/ Do not over optimize for long context: it can hurt short context performance: https://t.co/9tb3zZLp4v
6/ Most Performance of Mamba2 and GDN is Achievable with the simplest GLA + Cannon: https://t.co/jQZqJKBSsF
Transformers vs linear models
- w/ Horizontal info flow: GDN, Mamba2
- w/o Horizontal info flow: Transformers, GLA
1/ Mamba2 is weak in reasoning breadth:
https://t.co/eDJ4rdhWqA
2/ Linear Models can hold ~40% more knowledge compared to Transformer, but Transformers can reason 4x deeper : https://t.co/sbvzbLRMv1
3/ Whenever linear models perform badly, it is not simply due to recurrent memory limits, but its inefficient compression and retrieval, learning 1 hop reasoning slowly: https://t.co/KAdM5d25ds
Synthetic vs real world pre-training
1/ Real world pre-training hides underperforming dimensions; synthetic datasets act as a magnifying glass while still aligning with real world trends:
https://t.co/gEnxqTY2IU
2/ Real world data is too skill mixed and delays emergence: https://t.co/dsoq0Rjar8
Philosophy Notes
1/ Existence is cheap, but learnability is everything.
2/ Physics says: Beyond this noise.
Again, It’s important to design a versatile pre training playground, otherwise you’re going to miss the information in some other dimensions of the skills:
https://t.co/OJxtZSsMKN
Closing
Big takeaway for me:
- A well-designed synthetic dataset is crucial for architecture research. Not just because synthetic pretraining is cheaper, but because it is a cleaner microscope for revealing which skills a model truly has the potential for.
- “Toy” datasets do not have to be toyish. “Simple” can mean controlled noise, structured, and truthful, with less effort to verify. And I guess the real point, though, is not “synthetic vs real”. It is that designing scientific benchmarks is never trivial, and always essential.
One of the underrated papers this year:
"Small Batch Size Training for Language Models:
When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful" (https://t.co/0O4XjGDLIP)
(I can confirm this holds for RLVR, too! I have some experiments to share soon.)
Just updated the Big LLM Architecture Comparison article...
...it grew quite a bit since the initial version in July 2025, more than doubled!
https://t.co/oEt8XzNxik
Disney Research shared a new look at the Olaf animated character, included a closer look at the inner workings of the robot, walk cycle, impact reductions, and how they track the character's performance.
Ivan Sorokin and I are the official winners on the Arc Prize competition, with a significant lead over other teams.
Thanks to @kaggle and @arcprize for hosting the competition.
NVIDIA tech blog summarizing what we did: https://t.co/3Q34gFrrWq
Our writeup: https://t.co/v89q2QskPg
Our code: https://t.co/DIGuaf3dTQ
ну и в комментариях вспомнили про Лимонова, и у меня ощущение, что вот Лимонов он вообще не такой. когда читал эдичку, у меня не было ощущение такого отвращения от текста, то как гг описывает женщин меня просто выворачивает. но мб Лимонов делал так же, но я не замечал...
блин я короче вчера когда загуглил этот рассказ и нашёл его книгу, так офигел, конечно. как можно писать такие тексты для песен и такие тексты для книги диаметрально просто противоположные, я не понимаю..
возможно, конечно, вся суть в том, что они и не противоположные совсем, а одно и то же на самом деле, и поэтому людей так трясет от текстов его песен, но мне как-то не верится в это...