The 2nd bucket of highly-paid AI talent is emerging: the ones who are deep (enough) into not only pre-training and post-training of LLMs – but with the complementary skills to make LLMs actually work in messy real-world use cases (I don't mean SWE skills here, that's table-stakes).
Are they special? Not sure – just very rare right now – need ~10k hours of focused IC practice in real-world scenarios, when “GenAI” itself is barely 3 years old.
(Related to the recent “95%” Fortune article, the underlying MIT report, and all the chatter about the GenAI bubble)
My top 5 most memorable “LLMs” launches:
1. text-davinci-002 (first one that really "got it"/worked)
2. GPT4 (biggest step function jump seen till now)
3. Clause 3.5 Sonnet (first true dethroning)
4. o1-pro (clear glimpse of robust human-like reasoning)
5. DeepSeek-R1 (proof open can beat closed)
I am at #NeurIPS2024 this week! Key ML areas our group under @timshi_ai at @cresta is working on:
- AI Agents than can reason and troubleshoot effectively in complex enterprise domains
- Multimodal Knowledge Grounding
- LLM-as-a-judge framework that actually works
We are hiring!
This, w.r.t. emergence of consciousness. And I think the “deliberate” aspect is key here. GEB talked about it also in “Jumping out of the System” – the need to jump out of the task being performed, survey whats been done, and ensure key requirements (consistency AND efficiency).
In the brain, some neurons adapt easily, while others remain resistant to change. This can be likened to our beliefs, where certain deeply ingrained convictions, like the historical belief that the sun revolves around the earth, required substantial evidence and effort to alter.
Similarly, in machine learning, we could enhance efficiency by incorporating a "resistance to change" factor for each neuron during training. This factor would determine how readily a neuron or set of neurons can adapt or how firmly they maintain their learned patterns (or weights).
By drawing on the concept that some brain regions or neuron patterns encode hard-set beliefs through repeated reinforcement, we can apply this principle to improve the adaptability and stability of artificial neural networks. This could be a pathway for the Alignment Problem.
PS: Dropout is a somewhat overlapping, but a completely different concept, introduced for a different need.
I am at #NeurIPS2023 this week! Some ML areas our group under @timshi_ai at @cresta is working on:
- Domain specific instruction finetuning
- Retrieval Augmentation and Knowledge Grounding
- Reward modeling and conversation-level outcomes
Hit me up for a chat. We are hiring!