"We just want what's best for you"
The purpose and objective of a system is determined by its actions on its environment, not by its internal communication. This is obvious. Thus, language cannot be the main source of intelligence.
Don't ask a system its purpose, observe its actions. Are you not tired of all the organizations with their mission statements when their impact on our world is completely different from it? Intuitively, I believe you already know this is true.
New paper: Scaling Laws for Behavioral Foundation Models over User Event Sequences
Behavioral FMs are models trained on sequences of user actions in domains like payments, transactions, recommendations, commerce, and fraud.
We run ~600 experiments on real interaction data, spanning 10^15 to 10^19 FLOPs, to derive compute-optimal training recipes for behavioral FMs.
We jointly study parameter split, batch size, model/data allocation, and sampled negatives.
Key findings:
• optimal event embedder share is small and stable, ~2%
• low-budget training is data-heavy, then moves toward Chinchilla-like D/N
• critical batch size scales with compute and depends on the deployed metric
• larger compute budgets increasingly prefer more sampled negatives
• by 10^19 FLOPs, negative sampling becomes memory-bound rather than FLOP-bound
• loss and deployed ranking metrics disagree in ways that scale
Paper: https://t.co/YBgsZP4u09
A lot of fun being interviewed by @Forbes about building AI ecosystems. As CEO of @UnboxAI_, I spoke about how behavioral data is reshaping the AI landscape through Large Behavioral Models and redefining what's possible. I’ll be speaking more about this at @HarvardHBS on Thursday at 4:00 pm. Sign up here: https://t.co/ZnWdxEEIQu
Exciting to see @Visa citing @UnboxAI_ multiple times in their new paper on transaction foundation models. 🚀
Love seeing Visa enter this space with "TransactionGPT." It’s a huge validation of the direction we’ve been pushing: treating behavioral and transactional data as first-class citizens in foundation models, not an afterthought.
The fact that our BehaviorGPT whitepapers are referenced throughout their work is a strong signal that this is now a real, recognized category.
Visa’s research confirms what we see daily: behavioral foundation models drastically increase performance, significantly outperforming finetuned LLMs in: (1) Accuracy, (2) Latency, and (3) Cost
That is a real MOAT.
We’re still just scratching the surface of what behavior-native models can unlock across risk, personalization, and product, but it’s clear the rest of the ecosystem is realizing this is the future.
Congratulations to @dozee_sim and the team on the paper!
Visa’s preprint: https://t.co/4pmJVKQ0K3
Our cited BehaviorGPT research: https://t.co/IWkSNnaCDX
"How come it’s so easy to lie to each other?" It’s even easier to lie to AI.
In this short excerpt from my @TEDxMIT talk, I explain why talk is cheap — for humans and for the large language models (LLMs) we build.
If our current AI systems are trained on cheap internet text, is the resulting intelligence also cheap? And what if AI could learn from something richer: our actions?
📺 Full TEDxMIT talk here: https://t.co/jhZ9Gq5Y2u
𝗘𝘅𝗰𝗶𝘁𝗲𝗱 𝘁𝗼 𝗶𝗻𝘁𝗿𝗼𝗱𝘂𝗰𝗲 𝘁𝗵𝗲 𝗳𝗶𝗿𝘀𝘁 𝗯𝗲𝗵𝗮𝘃𝗶𝗼𝗿𝗮𝗹 𝗙𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻 𝗠𝗼𝗱𝗲𝗹 𝗳𝗼𝗿 𝗩𝗶𝘀𝘂𝗮𝗹 𝗔𝗿𝘁 𝗮𝗻𝗱 𝗔𝗲𝘀𝘁𝗵𝗲𝘁𝗶𝗰𝘀 🚀
We trained on 𝟮𝟭𝟱 𝗯𝗶𝗹𝗹𝗶𝗼𝗻 𝗵𝘂𝗺𝗮𝗻 𝗶𝗻𝘁𝗲𝗿𝗮𝗰𝘁𝗶𝗼𝗻𝘀 (𝟰.𝟳 𝘁𝗿𝗶𝗹𝗹𝗶𝗼𝗻 𝘁𝗼𝗸𝗲𝗻𝘀) across art platforms. Unlike pixel- or object recognition based vision foundation models (CLIP, Dall-E, Stable Diffusion), BehaviorGPT learns from sequences of user actions—capturing intent, preferences, and nuances that drive real engagement.
Key highlights:
• +16% search conversion, +24% in recommendations, +11% dynamic categorization, +14% SEO/assortment optimization.
• 12x fewer human resources needed.
• Qualitative wins: Personalized motif generation, natural language navigation, and ethical attribution in AI art.
• Annual retraining yields 2.5% compounding gains. A first number for the "time-value of AI" in behavioral data loops.
This positions behavioral semantics as a game-changer over traditional VFMs, emphasizing actions over abstractions. Perfect for e-commerce, creative tools, and beyond!
Read the full paper here: https://t.co/voHZVAQbRx 📄
What do you think?
𝗪𝗼𝗿𝗹𝗱’𝘀 𝗳𝗶𝗿𝘀𝘁 𝗯𝗲𝗵𝗮𝘃𝗶𝗼𝗿𝗮𝗹 𝗺𝗼𝗱𝗲𝗹 𝗳𝗼𝗿 𝘁𝗵𝗲 𝘄𝗼𝗿𝗸𝗳𝗼𝗿𝗰𝗲
TL;DR: We treated four years of workforce behavior, 𝟰𝟯 𝗺𝗶𝗹𝗹𝗶𝗼𝗻 𝗲𝘃𝗲𝗻𝘁𝘀 from 𝟴𝟬 𝟬𝟬𝟬 𝗲𝗺𝗽𝗹𝗼𝘆𝗲𝗲𝘀, as a “language.” A Transformer (just like a large language model) learned to predict the next action and reached 𝟵𝟭 % 𝗮𝗰𝗰𝘂𝗿𝗮𝗰𝘆 (𝗙𝟭 = 𝟬.𝟵𝟯) in forecasting whether an employee will quit in the coming month and addressing $𝟲𝟯𝗠 𝗶𝗻 𝘀𝗮𝘃𝗶𝗻𝗴𝘀.
Check out the full white paper here: https://t.co/TVBXRXb8V5
𝗛𝗶𝗴𝗵𝗹𝗶𝗴𝗵𝘁𝘀:
1) Foundation over point solution: We didn’t train just for attrition. We first taught the model to predict any next action, then fine-tuned it for attrition, just like how language models learn general patterns before specializing.
2) Pre-training paid off: Starting from a general behavioral model gave us a 7% accuracy boost compared to training just for attrition from scratch.
3) Patterns in the data: When we visualized employee embeddings, we saw clear clusters by country, showing the model had learned deep behavioral differences without being told.
4) Real business impact: For a company with 30,000 employees and 70% turnover, reducing quits by just 15% would save $63 million per year, even with conservative assumptions.
5) A new kind of insight: Surveys and dashboards missed the signal. Modeling behavior directly worked. This is the future of understanding people at scale.
𝗙𝗿𝗼𝗺 𝗰𝗼𝗻𝘀𝘂𝗺𝗽𝘁𝗶𝗼𝗻 𝘁𝗼 𝗴𝗹𝗼𝗯𝗮𝗹 𝘄𝗼𝗿𝗸𝗳𝗼𝗿𝗰𝗲𝘀 𝗮𝗻𝗱 𝗯𝗲𝘆𝗼𝗻𝗱: Large‑Behavior Models are doing for real‑world actions what LLMs did for words. The frontier is no longer what people say but what people actually do, and the business upside is enormous.
𝗦𝗶𝗹𝗶𝗰𝗼𝗻 𝗩𝗮𝗹𝗵𝗮𝗹𝗹𝗮? 𝗜𝘀 𝗦𝘁𝗼𝗰𝗸𝗵𝗼𝗹𝗺 𝘁𝗵𝗲 𝗻𝗲𝘄 𝗦𝗮𝗻 𝗙𝗿𝗮𝗻𝗰𝗶𝘀𝗰𝗼?
I left Sweden for Silicon Valley when I was 19, almost because I had to. But every time I come back, I’m reminded why I’m so proud of this place. Maybe 19-year-olds today don’t need to leave to realize their startup dreams.
Today, there’s something special about the startup and AI community here. People genuinely care. About each other. About building great products. About solving real problems.
A recent dinner with @antonosika (@lovable), @joelhellermark (@sanalabs), and Andrey (@MiroHQ) was a great example of that spirit. Thoughtful, ambitious, and grounded.
Unbox AI has deep Swedish roots and a growing team in Stockholm. We’re proud to be part of this ecosystem and hope we’ve contributed to it in a meaningful way.
While Sweden has fewer startups than Silicon Valley, it makes up for it in quality. In the Valley, a higher fraction of entrepreneurs seem more focused on projecting the image of the next Steve Jobs than on solving real problems. In Sweden, I see more founders genuinely driven by the problems they’re trying to solve. Of course, no place is perfect, and each has its own unique strengths.
Stockholm continues to inspire me.
If you want to chat about startups, AI, or foundation models while I’m in town, just reach out!
The bet @UnboxAI_ did 6 years ago was exactly the same bet @OpenAI did 7 years ago. That with enough data and scale, transformers with autoregressive causal modeling will lead to superhuman intelligence. OpenAI did this for human language, we did this for behavior, actions, and transactions (human implicit language). Unbox AI is the OpenAI of behavior, and if it turns out that language is not all you need, we will at last get the credit we deserve.
Our vote: Large Behavioral Models (LBMs) will work in synergy with Large Language Models (LLMs)! 🚀 Just like humans have multiple sub-brains interacting, not just one type intelligence that does everything
Curious about people's take on this. Transformers + self-supervised learning works for any domain where there is lots of data (genomics, behavior, language, etc). Does those foundation models matter or will LLMs be the winner-takes-all?
If you believe that other foundation models will be critical, what will those models be and why are so few people working on them?
I think running a company is more of a sport than a science. It's about building muscle memory through repeated exposure. In this MIT entrepreneurship lecture, we use practical role-playing and case studies to exercise those muscles.
https://t.co/nfDavwObIs
To master a sport, the focus can't be theory. It has to be practice.
Heard about @stripe’s foundation model for transactions? Wanna know how we built one bigger, on more data, long before stripe did--with all the details of how we did it. Checkout our technical blog post: https://t.co/Ms9G7sawHl
TL;DR: We trained a transformer (just like a large language model) but instead of learning natural language, it learned the "language" of transactions. Specifically for purchasing behavior around groceries. It increased conversion by 1050%.
We took discrete payment data, searches, and clicks (massive amounts of transaction info) and trained the model to detect subtle patterns across billions of transactions. And to predict for every user, the next transaction based on previous transactions.
This gave them a foundation model for transaction and user actions: a general-purpose model that encodes rich transaction info and can be reused across many tasks.
One standout result: It increased conversion on recommendations by 1050%. Below you can see the resulting dense semantic vectors (embeddings) for products. Turns out that they contain a lot of insights. Products were split into two halves, based on something akin to “affordable” vs “premium” products, and then into subclusters from there.
Another standout result: The dense vectors (embeddings) of physical stores can also be seen below, where the same colors are similar in behaviour. Turns out the regional proximity isn’t sufficient to determine the store cluster. This allowed us to build new assortments based on customer behavior and transactional patterns in each region across Sweden, which increased sales by 2.2%.
Another standout fact: This was in 2020 and you can read more about it in the blog post. This was what started our journey on building BehaviorGPT for payments and transactions. We're now training a model over 6 trillion behavioral tokens and transactions, and we cannot believe the intelligence we're seeing. Stay tuned!
@TechCrunch Great to see this space heating up! @unboxai_ has been advocating & building payment foundation models for 5 years—now expanding to all purchase-related actions. Trained on trillions of data points 🚀
Video: https://t.co/2qamCBoe7v
White paper: https://t.co/cHSppuJqEc
AI thrived when it abandoned physics’ dream of one equation to explain everything. But now we chase ONE model to rule all knowledge. Reality is messy—we need an ecosystem of minds, not a monarch model. Agree? Disagree? Debate me! MIT lecture on this topic: https://t.co/WWRSU0Gwqi
Great to see @stripe accelerating, but we’re already a lap ahead. Our foundation model covers every payment touchpoint—fraud, clicks, searches, returns & purchases—trained on trillions (yes, trillions) of datapoints. Learn more at https://t.co/xEnhBrImBX
TL;DR: We built a transformer-based payments foundation model. It works.
For years, Stripe has been using machine learning models trained on discrete features (BIN, zip, payment method, etc.) to improve our products for users. And these feature-by-feature efforts have worked well: +15% conversion, -30% fraud.
But these models have limitations. We have to select (and therefore constrain) the features considered by the model. And each model requires task-specific training: for authorization, for fraud, for disputes, and so on.
Given the learning power of generalized transformer architectures, we wondered whether an LLM-style approach could work here. It wasn’t obvious that it would—payments is like language in some ways (structural patterns similar to syntax and semantics, temporally sequential) and extremely unlike language in others (fewer distinct ‘tokens’, contextual sparsity, fewer organizing principles akin to grammatical rules).
So we built a payments foundation model—a self-supervised network that learns dense, general-purpose vectors for every transaction, much like a language model embeds words. Trained on tens of billions of transactions, it distills each charge’s key signals into a single, versatile embedding.
You can think of the result as a vast distribution of payments in a high-dimensional vector space. The location of each embedding captures rich data, including how different elements relate to each other. Payments that share similarities naturally cluster together: transactions from the same card issuer are positioned closer together, those from the same bank even closer, and those sharing the same email address are nearly identical.
These rich embeddings make it significantly easier to spot nuanced, adversarial patterns of transactions; and to build more accurate classifiers based on both the features of an individual payment and its relationship to other payments in the sequence.
Take card-testing. Over the past couple of years traditional ML approaches (engineering new features, labeling emerging attack patterns, rapidly retraining our models) have reduced card testing for users on Stripe by 80%. But the most sophisticated card testers hide novel attack patterns in the volumes of the largest companies, so they’re hard to spot with these methods.
We built a classifier that ingests sequences of embeddings from the foundation model, and predicts if the traffic slice is under an attack. It leverages transformer architecture to detect subtle patterns across transaction sequences. And it does this all in real time so we can block attacks before they hit businesses.
This approach improved our detection rate for card-testing attacks on large users from 59% to 97% overnight.
This has an instant impact for our large users. But the real power of the foundation model is that these same embeddings can be applied across other tasks, like disputes or authorizations.
Perhaps even more fundamentally, it suggests that payments have semantic meaning. Just like words in a sentence, transactions possess complex sequential dependencies and latent feature interactions that simply can’t be captured by manual feature engineering.
Turns out attention was all payments needed!
My holiday gift to you: What is the AI revolution, and how can you come out winning? 🎁 2025 edition!
Thank you to @nttdatalatam for hosting me
https://t.co/e5LsW73JGm
🚀 Excited to share my fascinating panel on the Future of AI with Professor @manoliskellis. Will the new AI bring prosperity or doom? It was an absolute pleasure to delve into these topics with such distinguished experts: https://t.co/pbqLIbLj4N