Congrats to @DeepgramAI on the GA launch of Flux TTS, now live in Pipecat 🎉
@JonPTaylor looks at how Flux TTS delivers a more consistent voice experience. Flux reads the whole conversation, not just the next line — adaptive tone, consistent pronunciation, clean interruption handling natively (no SSML markup, no style tags!) all at sub-200ms latency.
➡️ https://t.co/H8IsYicQ0z
Two months ago on No Priors he said he manifests his will to agents 16 hours a day and that coding isn’t even the right verb anymore.
Today he joins Anthropic.
When Karpathy picks his next move, you pay attention.
Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
We’re adding new ways for people to identify AI-generated images and understand where they came from.
In addition to C2PA Content Credentials, images now also contain a SynthID watermark, and can be identified using a public verification tool to check whether an image was made by OpenAI products.
https://t.co/qo0l4vyWli
sell me this devtool
@minitap_ai has launched minitest. Fully managed AI QA engineer for mobile, w/ zero authoring, zero maintenance. It catches the UI/UX, performance, payment issues your users would've hit, before they hit them.
comment for free 2 months unlimited access and a personal onboarding from the team
in the meantime enjoy my Oscar-worthy acting
This is what I was looking for in AI-native email, as we're still spending way too much time in our inbox... Not really the age of intelligence just yet.
Been an early tester, as I'm a strong believer that email should work through voice now.
We've raised $8.4m to make software that sharpens your attention.
We're starting with an email app that lets you handle your Gmail inbox in seconds.
We call it Avec. It's available now!
We’re saying goodbye to the Sora app. To everyone who created with Sora, shared it, and built community around it: thank you. What you made with Sora mattered, and we know this news is disappointing.
We’ll share more soon, including timelines for the app and API and details on preserving your work. – The Sora Team
Today we're releasing cara-3, our latest face-generation model. In an independent blind study, participants preferred Anam's interactive avatars over other providers across every metric measured. But why do we care about avatars to begin with?
https://t.co/e4RhlhIAN9
The work Alain is doing today is the third generation of a family tradition of embracing new technologies.
Alain's grandfather was a medical doctor, but became known for his award-winning photography. His father carried the craft forward with Super 8 films. Alain spent days in the studio as a kid, cutting and gluing film by hand.
That “make something from nothing” instinct showed up in other ways, too.
At 17, he convinced 10 people to trust him with their money for day trading. 10x’d it on the way up, then watched it fall and learned what risk really means on the way down.
By 24, he'd spotted an underdeveloped business unit at Merck & Co. and beat the forecast by 3x. €20M more than expected.
In 2011, most people saw the iPad as "an iPhone with a bigger screen." Alain saw an enterprise interface that could change how business gets done face-to-face.
That became Pitcher, his startup which he bootstrapped to an 8-figure exit in 2022, right after the Nasdaq dropped 50%.
Then he went into AI. Maybe too early.
In 2023, he was building AI agents before the market even had language for them. One insight kept surfacing: brands wanted known faces in their AI campaigns. But there was no easy way to license them.
Today, that's Twinity. Alain’s tech makes talent licensing scalable like software, so brands can move at the speed of AI.
Third generation of making something from nothing. Just a different medium.
Watch Alain tell his story below.
Y Combinator JUST announced what startups they want to fund next in 2025. And it's mostly AI that replaces $100k/year job functions.
My notes below in case it's helpful to you:
@giffmana believe o1 wouldn‘t have worked if it were released before the others. kinda like the moment i switched to mac and had to entrust photo mgmt to apple. bit of a black box that requires trust.
now that i know how great GPT4 is, i trust o1 to reason stuff for me.
@giffmana@hallerite Lucas, i just noticed that o1 started doing just that, i.e. asking if it understands correctly.
until today my record was 23s rhinking, today ot was thinking for 2min2s and then came back to ask me if it understood the goal correctly, then thought for another 1min38s… 🤯🤩
If it’s not clear yet, this is what I think will happen soon.
Every major tech company except Apple has announced their own LLM.
Apple have spent years perfecting their on-device neural engine. Capable of some absolute insane operations. Loads of compute in a small and energy efficient form factor.
With M1, M2 & soon M3 the neural engine is even more powerful than their A series mobile chipsets.
While we currently need the cloud to run ChatGPT and it’s clunky, I think Apple is going to blow everyone out of the water here. Both on desktop class hardware and mobile.
I think Apple will be launching their own secure and private LLM that runs on device (edge compute). And when necessary it offloads more heavy workloads to a cloud based LLM that’s optimized for heavier tasks. So we will initially have some hybrid.
Personal, with tight hardware and software integration this AI will be omnipresent. Apple will probably use this to sell a lot of new hardware that they claim is needed to run this. They will make a lot of moneys.
For me the LLM’s will form the new protocol level technology upon which most new software will be built. We will have to re-wire our core understanding about what an application is.
Single-use apps will be a huge thing. If you need to solve a unique problem, and nobody has ever done software for that because not enough market. With an LLM even a problem with only one user, will be doable, enter your ask, and code gets written, problem gets solved. Runtime ends, app dies. Done. Single use apps are born.
It’s hard to predict or try to understand how the world will look just 10 years from today. It will be very different, we have passed the inflection point, the rocket engines have been lit. We’ve taken off.
Add to all the above that every single field, category and market will be disrupted at the same time. And not only with text/coding but with any multi-media we have. Images, video & audio. Anything we can come up with can and will be enhanced or disrupted by AI.
Once we got more people that will have their AI A-ha moment the rate of change and adoption will continue to increase. This will continue until we have global access and coverage.
People will get left behind, and this will be one of the most important things to try to combat. Having a 0% left behind policy. We need to make sure AI benefits all.
We’re living through a paradigm shift, and we’re witness a new protocol level technology. We’re seeing it arrive in real-time and most people have no clue about what’s about to happen.
I’m not an AI alarmist, I’m an AI gardener, and optimist.
We will have time to adapt. Not as long as we had during the Industrial Revolution, but enough time to make sure we have a chance at a positive outcome.
We’re moving away from the Information Age into the Age of Intelligence. With unlimited access to intelligence anywhere, anytime.
18th Mars 2023 - Linus Ekenstam
In my decade spent on AI, I've never seen an algorithm that so many people fantasize about. Just from a name, no paper, no stats, no product. So let's reverse engineer the Q* fantasy. VERY LONG READ:
To understand the powerful marriage between Search and Learning, we need to go back to 2016 and revisit AlphaGo, a glorious moment in the AI history.
It's got 4 key ingredients:
1. Policy NN (Learning): responsible for selecting good moves. It estimates the probability of each move leading to a win.
2. Value NN (Learning): evaluates the board and predicts the winner from any given legal position in Go.
3. MCTS (Search): stands for "Monte Carlo Tree Search". It simulates many possible sequences of moves from the current position using the policy NN, and then aggregates the results of these simulations to decide on the most promising move. This is the "slow thinking" component that contrasts with the fast token sampling of LLMs.
4. A groundtruth signal to drive the whole system. In Go, it's as simple as the binary label "who wins", which is decided by an established set of game rules. You can think of it as a source of energy that *sustains* the learning progress.
How do the components above work together?
AlphaGo does self-play, i.e. playing against its own older checkpoints. As self-play continues, both Policy NN and Value NN are improved iteratively: as the policy gets better at selecting moves, the value NN obtains better data to learn from, and in turn it provides better feedback to the policy. A stronger policy also helps MCTS explore better strategies.
That completes an ingenious "perpetual motion machine". In this way, AlphaGo was able to bootstrap its own capabilities and beat the human world champion, Lee Sedol, 4-1 in 2016. An AI can never become super-human just by imitating human data alone.
-----
Now let's talk about Q*. What are the corresponding 4 components?
1. Policy NN: this will be OAI's most powerful internal GPT, responsible for actually implementing the thought traces that solve a math problem.
2. Value NN: another GPT that scores how likely each intermediate reasoning step is correct.
OAI published a paper in May 2023 called "Let's Verify Step by Step", coauthored by big names like @ilyasut@johnschulman2@janleike: https://t.co/iAvXNjjhcK
It's much lesser known than DALL-E or Whipser, but gives us quite a lot of hints.
This paper proposes "Process-supervised Reward Models", or PRMs, that gives feedback for each step in the chain-of-thought. In contrast, "Outcome-supervised reward models", or ORMs, only judge the entire output at the end.
ORMs are the original reward model formulation for RLHF, but it's too coarse-grained to properly judge the sub-parts of a long response. In other words, ORMs are not great for credit assignment. In RL literature, we call ORMs "sparse reward" (only given once at the end), and PRMs "dense reward" that smoothly shapes the LLM to our desired behavior.
3. Search: unlike AlphaGo's discrete states and actions, LLMs operate on a much more sophisticated space of "all reasonable strings". So we need new search procedures.
Expanding on Chain of Thought (CoT), the research community has developed a few nonlinear CoTs:
- Tree of Thought: literally combining CoT and tree search: https://t.co/KM1P2ZJrjG @ShunyuYao12
- Graph of Thought: yeah you guessed it already. Turn the tree into a graph and Voilà! You get an even more sophisticated search operator: https://t.co/5ncT5tuTOY
4. Groundtruth signal: a few possibilities:
(a) Each math problem comes with a known answer. OAI may have collected a huge corpus from existing math exams or competitions.
(b) The ORM itself can be used as a groundtruth signal, but then it could be exploited and "loses energy" to sustain learning.
(c) A formal verification system, such as Lean Theorem Prover, can turn math into a coding problem and provide compiler feedbacks: https://t.co/vpOBOI2FR5
And just like AlphaGo, the Policy LLM and Value LLM can improve each other iteratively, as well as learn from human expert annotations whenever available. A better Policy LLM will help the Tree of Thought Search explore better strategies, which in turn collect better data for the next round.
@demishassabis said a while back that DeepMind Gemini will use "AlphaGo-style algorithms" to boost reasoning. Even if Q* is not what we think, Google will certainly catch up with their own. If I can think of the above, they surely can.
Note that what I described is just about reasoning. Nothing says Q* will be more creative in writing poetry, telling jokes @grok, or role playing. Improving creativity is a fundamentally human thing, so I believe natural data will still outperform synthetic ones.
I welcome any thoughts or feedback!!
Announcing Grok!
Grok is an AI modeled after the Hitchhiker’s Guide to the Galaxy, so intended to answer almost anything and, far harder, even suggest what questions to ask!
Grok is designed to answer questions with a bit of wit and has a rebellious streak, so please don’t use it if you hate humor!
A unique and fundamental advantage of Grok is that it has real-time knowledge of the world via the 𝕏 platform. It will also answer spicy questions that are rejected by most other AI systems.
Grok is still a very early beta product – the best we could do with 2 months of training – so expect it to improve rapidly with each passing week with your help.
Thank you,
the xAI Team
https://t.co/iPqreWxmQh