Are we in a new, 3rd paradigm for AI language models?
First, models predicted the most likely next word. Think 2018-2021, for transformer-based language models.
Second, they were rewarded for words that were helpful, harmless, and honest. Think RLHF or RLAIF, 2022-2023.
Now, with the o1 family, they are being rewarded for being objectively correct. Think 2024-???
Breakdown in video below:
5/5 In AlphaCodium @talrid23 introduced a well-engineered flow,
using simple tools with hard-coded algorithms,
where the AI was focused on analyzing & generating specs, pseudo-code, and code
Would an abstraction layer allow for easier further research?
https://t.co/t5FEkBb384
Based on StableNormal, @ychngji6 developed StableDelight for real-time reflection removal from textured surfaces.
Please check out our online demo!
https://t.co/XHNAXidKyJ
We should be awestruck by the fact that when asked to make something “beautiful” an LLM can produce it—in design, prose, or other forms.
Beauty may be an emergent property of these systems.
And if that's true, perhaps it's also a fundamental property of the universe itself.
🤯So real agents came faster than I thought
Without any Python ability. I signed up for Replit and used their new AI coder. I was able to build a working app that identifies sentiment at the paragraph level in 23 minutes. I only interacted 8 times. It did the design & debugging
Andrej Karpathy says as we expand our brains into an exocortex on a computing substrate, we will be renting our brains and open source will become more important because if it's "not your weights, not your brain"
My new mad theory that asking claude for the correct_answer ( or some equivalent ) instead of an answer has some new proof
1. Chain of thought is more consistent with final answer
2. Chain of thought is better quality
3. It repairs/corrects itself much better
Small thread of experiments today
wait, what? @cohere just dropped updated Command R plus & Command R - built for RAG and tool use, multilingual (23 languages), grounded generation, 128K content and much more! 🔥
> Command R Plus - 104B param, Command R - 35B param
> up-to 2x higher throughput & 2x lower latency
> Uses Grouped Query Attention (GQA)
> SFT + preference tuned model
> Massive 128K context window
> Trained in 23 languages, evaluated on 10
> Capable of code rewrites, explanations and snippets.
> Supports citation, tool execution, and structured outputs for grounded agentic use
> Model checkpoint available on the hub
> Works out of the box with transformers 🤗
This is a massive improvement from the last iteration of Command R+ (which is still one of my favourites on Hugging Chat).
Kudos, @cohere. I'm looking forward to giving it a spin.
Just tried Bland AI's new real-time conversational agent. Mind blowing🤯
This could disrupt call centers and outreach roles.
Talk to the AI agent on their website, it's wild.
https://t.co/FhFkSYUllo
High-quality image generation without the need for prompt-crafting is now ubiquitous (but ChatGPT's DALL-E 3 is lagging).
Here is "a high-fashion photoshoot of a knight wearing Monet-inspired armor" in Grok/Flux, Google's new system, Midjourney & inexplicably DALL-E. (best of 4)
Our CEO @DemisHassabis imagines a "international CERN for AI" – a collaborative hub where the world's top researchers work together to build AGI safely and responsibly. 🌐
Watch more on the latest episode of Google DeepMind: The Podcast ↓
https://t.co/LhF1ygWxzs
With prompt caching, you can reuse a book's worth of context across multiple API requests. This can also reduce latency by up to 85% on long prompts.
Use cases include coding assistants, large document processing, and agentic tool use. Get started: https://t.co/kC1ep1va6C
Holy shit. Without a doubt the most realistic AI images I've ever seen.
We are 99.7% of the way to completely indistinguishable-from-reality AI imagery.
(You can still see a few flaws when zooming in)
This is made with FLUX. Uncanny Valley.
Companies are just releasing next level models without telling anyone about it, just quiet, somewhat baffling roll-outs.
Gemini 1.5 Pro Experimental 0801 (really rolls of the tongue, there) is at the top of the major AI leaderboard. What's different about it? Who knows!