Exciting news: DeepSeek-V4-Flash-High by @deepseek_ai has reshaped the Pareto Frontier in the Frontend Code Arena, with a score of 1586!
Priced at $0.14/$0.28 per MToken, it’s the best performance-per-dollar of any model in its class.
Congrats to the @deepseek_ai team!
Today is very exciting day for us.
I'm so happy and proud for what this means for every single one of our team who worked so hard, and I'm so incredibly thankful for everyone who supported us in any and every way along this journey, and also thankful to all our customers who always believed in us, both big & small.
This is a huge day, but of course our story continues, our regular product launch cadence will shortly :)
We’re excited to share that we just signed an agreement for @salesforce to acquire @fin_ai for ~$3.6B. The transaction is expected to close in the fourth quarter of Salesforce’s fiscal year 2027.
Fin started as Intercom 15 years ago. We changed our name to cap our transformation just weeks ago. We were a darling of the SaaS era and invented so many of the patterns you see in software today. Nearly four years ago, in need of a reboot, we jumped on weeks-old modern LLMs to create and define the category we know as Customer Agents today.
Salesforce invented modern software and SaaS. And @benioff is like the final boss of tech founder CEOs. In seat for 27 years, he’s one of the last of his era. Still pushing, pivoting, placing big bets. It’s a privilege for @destraynor and I to get to partner with him and join forces with Salesforce upon close at this most fascinating time. And will be very fun to get their help bringing Fin to magnitudes more consumers.
To our customers: Over the past few years we’ve been shipping intensely. Including recently our groundbreaking model, Apex, and our paradigm-defining internal agent, Operator. With the resources of Salesforce this will only accelerate. And yet little will practically change. I’ll still be CEO, Des will still be running R&D, we’ll both still be committed to continuing to lead this category. Thank you very sincerely and deeply for your belief in us.
To all of our friends, our families, and our employees, past and present: While this is not the end, it is a major, pivotal, special, and emotional moment for us. From the bottom of our hearts, thank you. For everything.
To my cofounders, my exec team: Look what we built. Four young lads with a dream and nothing to lose. And a home grown exec team who pulled off the greatest and arguably only late stage software company pivot to AI, and invented one of the most important categories in AI. Thank you for sticking through all of this with me.
And now, time to get back to work. See you at our next product launch in a couple weeks. (:
NEW YORK WENT ON A 44-11 RUN TO COMPLETE A 22-POINT COMEBACK WIN IN GAME 1 💨
DOWN 22 WITH UNDER 8 TO PLAY IN Q4.
30-8 RUN TO FORCE OT.
WON BY 11.
1-0 SERIES LEAD IN THE EAST FINALS 🍿
There is, of course, one AI Agent for Customer Support with public docs, self serve sign-up, public pricing, a CLI, API + more.
It's the highest performing one, and we share all our ideas+research too.
Product → fin .ai
Research → fin .ai/research
Ideas → ideas.fin .ai
"Universal One-third Time Scaling" - This paper attempts to explain why training LLMs is so slow, given that the loss converges according to a power-law whose origin is still debated. The authors show, through toy models and empirical evaluation on real LLMs, that this behavior arises intrinsically from the combined use of softmax and cross-entropy.
When learning peaked distributions like next-token distributions, these components produce losses and gradients that vanish as power-laws, yielding a universal scaling exponent of 1/3. A mechanistic explanation of neural scaling, with concrete directions for improving training efficiency.
https://t.co/xTOOaSk5Qb
Closed labs hide model sizes. They can't hide what their models know, and what a model knows is an indicator on how big it is.
Reasoning compresses. Factual knowledge doesn't. So you can size a frontier model from black-box API calls alone, and across releases you can literally watch a single fact arrive in the parameters over time.
For three years, my friends Jiyan He and Zihan Zheng have been asking frontier LLMs the same question: "what do you know about USTC Hackergame?", a CTF contest. May 2024: GPT-4o invented fake titles. Feb 2025: Claude 3.7 Sonnet listed 19 verified 2023 challenges. By April 2026, frontier models recall specific challenges across consecutive years.
After DeepSeek-V4 dropped, I instructed my agent to spend four days autonomously turning that habit into Incompressible Knowledge Probes (IKP) — 1,400 questions, 7 tiers of obscurity, 188 models, 27 vendors. Three findings:
1/ You can approximately size any black-box LLM from factual accuracy alone. Penalized accuracy is log-linear in log(params), R² = 0.917 on 89 open-weight models from 135M to 1.6T params. Project closed APIs onto the curve → GPT-5.5 ~9T, Claude Opus 4.7 ~4T, GPT-5.4 ~2.2T, Claude Sonnet 4.6 ~1.7T, Gemini 2.5 Pro ~1.2T (90% CI: 0.3-3x size).
2/ Citation count and h-index don't predict whether a frontier model recognizes a researcher. Two researchers with similar citation profiles get very different responses. Models memorize impact — work that shaped a field, not many incremental papers.
3/ Factual capacity doesn't compress over time. Across 96 open-weight models across 3 years, the IKP time coefficient is statistically zero, rejecting the Densing-Law prediction of +0.0117/month at p<10⁻¹⁵. Reasoning benchmarks saturate; factual capacity keeps scaling with parameters.
Website: https://t.co/CkwJsXqnsX
Paper: https://t.co/eNUdC9ye7w
We're announcing the most significant new strategic development since the start of the Customer Agent category: Fin now has specialized roles. And starting today anyone can sign up for and deploy the Fin Sales Role in minutes.
Fin is now by far the very best sales agent on the market, and it's been live with some of the most innovative digital brands for months, conversing with thousands of prospects, aiding discovery, building pipeline, booking meetings and starting trials.
We started the Customer Agent category with a focus on service by launching Fin three months after the launch of ChatGPT. Since then Fin has continued to lead the space, today delivering over 2 million resolutions a week for over 8k customers, including Anthropic, DoorDash, Snowflake, Asana, Mercury, Polymarket and many more exceptional brands.
But our vision for Fin has always stretched far beyond service. And for the past 6 months we’ve been building Fin as a single Customer Agent, delivering a seamless experience across all stages of the customer lifecycle.
We do not believe customers or businesses will want multiple agents for different parts of the lifecycle—because the agents won’t have shared memory and goals, and will have to compete with each other, delivering a poor customer experience and sub-optimal business results.
We’ve more roles to come, including another dropping in two weeks. For now please watch our launch video or visit our launch page to learn more: fin DOT ai SLASH sales
The AI Group building Fin introduced a new form of attention: Low-Rank Key Value attention.
It reduces KV cache memory by ~45–53% while improving model performance and training efficiency, by sharing a common KV base across heads with low-rank per-head specializations.
Read all the details in this article from @JamesONeil21 :
https://t.co/zdj0s6kGjy
Our AI group published a novel finding from pre-training research
In short: you can cut KV-cache memory in half by sharing most of the attention structure across heads + keeping small per-head differences—without hurting model quality or speed
I'll explain this as best I can
This is one of the most rewarding projects I've worked on. I am very grateful to the Intercom AI group for supporting this and for Fergal's leadership in enabling us to explore the frontier of AI research with practical benefits. See details of LRKV, our new attention mechanism
We’re excited to share our new form of attention, Low Rank Key Value attention.
This is a drop-in replacement to standard MHA that in our tests, reduces KV-cache by ~50%, with even lower test loss, across many scales of experiment.
Introducing M²RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
We bring back non-linear recurrence to language modeling and show it's been held back by small state sizes, not by non-linearity itself.
📄 Paper: https://t.co/AS8e2tNrRa
💻 Code: https://t.co/LMvBcI22Du
🤗 Models: https://t.co/NCmjrpNriq