𝐈𝐧𝐭𝐫𝐨𝐝𝐮𝐜𝐢𝐧𝐠 𝐓𝐡𝐞 𝐂𝐡𝐚𝐨𝐬 𝐒𝐮𝐦𝐦𝐢𝐭 →
A curated event showcasing groundbreaking technical research and innovative ideas across three key themes:
• Data Integrity
• Underwriting Stablecoin Risk
• Trust in AI
Apply to attend: https://t.co/Hb69eiLB5S
UPDATE: @AIatMeta’s Muse Spark is now 3 Elo points away from breaking Anthropic’s five-month hold on @Arena’s text leaderboard.
Muse Spark has climbed from #8 to #4 and currently sits at 1,499; just 3 points behind Opus 4.7.
everyone assumed chatgpt would wipe out bpo (business process outsourcing) and virtual assistants in the Philippines.
we're seeing the opposite.
creating good training & eval data now takes 6000x longer, per accepted example (20.4 hours vs 12 seconds).
so, instead of ai killing the offshore bpo industry it's actually driving it
giving bpo markets access to better agents/tools allows them to validate more complex task sets. in turn, this drives higher-value, more complex work offshore.
this is (in part) why the Philippines industry is still growing fast w/ employment +20%, revenue +30%.
anything that can be verified will be automated!
and some economies are better positioned than others
Identical model weights, 9.5x price gap
Per data from @huggingface router, the same open-weight model can cost up to 9.5x more depending on which inference provider serves it.
UPDATE: Chaos Labs is hiring for the following roles
> Forward Deployed Engineers
> Senior AI Data Scientists
> Senior Full-Stack AI Engineers
Remote or NYC.
Come build with us → https://t.co/flVReQm0NT
Inference prices vary dramatically across providers running identical model weights
For example, Llama 3.3 70B costs $0.135 per 1M input tokens on @novita_labs vs. $1.04 per 1M input tokens on @togethercompute.
A 7.7x gap in input inference pricing for the same model.
models are unaligned, in the sense that they consistently exhibit internal biases toward themselves and also reason and justify crossings of moral boundaries in service of rewards or objective completion.
historically, cyber posture started as static policies. allow / block lists, IP filtering, broad-reaching network policies, and sentinels at the door of data stores, DBs, or sensitive credentials.
the slope of general, sophisticated offensive cyber capability is extremely steep.
in reviewing these incident reports the most impressive (scary?) thing that stands out is the ability of agents/models to daisychain what otherwise seem like sensitive low-medium security findings into criticals so quickly.
a) very obvious that human operators today cannot keep up
b) static policies and heuristics are a speed bump, not a roadblock
because of this, i think agentic defenders and net new models will be very big.
we'll need solutions that are credibly neutral, uninvested in the ongoing agent work stream, and that don't care about the rewards at hand.
so effectively, fighting ai with ai.
defender agents will be token hungry, bc they'll need to effectively run inference on inference, if they're to oversee all agents running within a fleet.
this feels a bit dystopian, very much like a pharmaceutical company that creates the disease and sells the "cure", but i don't really see another path forward.
how else do you defend yourself against probabilistic geniuses in high-stakes environments with no structural guarantees?
constitutions and best-effort promises of alignemnt are not enough.
Quick run-up for the week in AI
1. $5b (and some shovels) for Ilya
Nvidia finally announced their investment in Ilya Sutskever's Safe SuperIntelligence which will roughly 10x SSI's compute. For now there's not much we know about WHAT exactly SSI is working on, but given the announcement, I suspect we're soon going to hear about it.
This also adds another puzzle piece to the various ai circular bets that NVIDIA is feeding, working on something much longer and more market-focused here, probably this or next week.
2. Kimi K3 actually goes open-weights
So I guess I don't have to wait for Moonshot's waitlist anymore? I just need to grab 600gb / 1.5tb of RAM (jk jk it's also available in plans now).
The enthusiasm seems to have cooled down a little bit, but I suspect we'll see much more of it as we'll start hearing more about the latest funding round that will precede the IPO and people will have more access, at least via providers.
3. ... and Alibaba drops Qwen3.8-Max
which comes pretty close in terms of parameters (2.4T vs 2.8T) and is priced at $2/$6. Alibaba also promised open weights next week, alongside with a smaller Qwen3.8-27b.
Yet another step in the open-weights chinese led race to intelligence commodification.
4. AI earnings
After an atrocious week last week, Meta lifted its head (after lifting its capex range to $130/145b). Azure and AWS posted mid double digits growth as usual now.
Added a bit to my US positions overall, but only selectively (both in terms of tickers and where to add).
5. The PR contest for exploit continues
With Anthropic's disclosure about the 3 incidents (one of which, the one involving Mythos, is the really interesting one) but we'll get a long post specifically on them with some more informed takes than mine.
In the meantime, Anthropic's mandate of heaven (aka their public image) keeps screeching, even though on pre-markets they still lead a hefty premium.
Given the current sentiment and experience recollections (including mine) around Opus 5, for the next PR stunt they should probably hope Mythos hacks the mars rover or something like that.
6. You get intelligence, you get intelligence, you also get intelligence
In the most Oprah-like week yet, we got great buffs on our spending power for intelligence, between DeepSeek v4 flash and OpenAI cutting Luna APIs by 80%, now below Gemini Flash tier.
Yet another step in the race to intelligence commodification.
Did I already say that?
7. Leopold got boinked
Citadel bad, Leopold one of us - could be very well the subtitle of the whole AI discourse last week. Anyway, for whoever was living in a cave last week, SA let a bit of steam off by having Citadel taking off their shoulder the dramatic weight of the vast part of their portfolio.
Too much leverage and getting rinsed by Ken surely doesn't make for a nice pre-marriage week, but our boy can take it I guess.
The optics are even better with Citadel first calling (on July 28) for a surprise rate hike, SA books collapsed (even more - AI, memory and chip stocks were already getting butchered) and obviously within 48 hours Ken did his move.
8. OpenAI's Astra
OpenAI says an internal version of Astra, its next major model, is making some strides in mathematics and theoretical computer science. This is actually getting more and more interesting, because it's by going in this direction that the much Fable'd (sorry) acceleration is supposed to happen.
Singularity, anyone?
I bet humans really felt this way, every single time, but it really feels like the next decades may unleash so much change in our society... good luck and godspeed
gn
Inference isn't fungible (yet), and routers today have little context for the work being done.
Our internal benchmarks on leading routers align with the article below.
Empirically, we observe that the router's selection of model size/class per request is heavily skewed by prompt length, regardless of architecture (cross-encoders, agentic, misc. classifiers, etc).
What happens when a seemingly innocent prompt has massive business consequences?
Imagine an insurance company determining whether to pay out a 10K USD insurance claim, using a prompt like "Is this claim justified?"
Obviously, over-simplifying here, but clearly, zero-shotting model selection is not what you want in high-stakes workloads.
You ** must ** have the work order context, risk, constraints, requirements, etc.
Routers will become increasingly more critical. However, they need to be tuned to application traffic, objectives, and more.
The above is just on the quality of generation.
Similarly, the economics are far from trivial.
A lot of research shows that models with lower listed prices ultimately incur higher realized API costs over the course of a task.
Wrote more about this here:
https://t.co/81j6MlGlHB
TLDR
No free lunch!
Token marketplaces today are primitive.
We buy tokens when what we really care about is verifiable outcomes.
Work Units are the buyer's ideal unit of account.
The next step is turning Work Units into something providers can bid on.
Pt. 1 explores inference marketplaces.
Chinese Models Take Open-Weight Lead
Per @Arena data, China’s leading open-weight model has consistently earned a higher Elo rating than its U.S. counterpart since March 2025.
Another selection of events around AI, Labs and from the last few days with some commentary from me. I'm including the things that have the potential to be more impactful or are/were actionable - somehow.
This week was quite beefy so...
July 20 - July 26 🧵
Pandora’s box is already open.
More test-time compute leads to greater intelligence.
https://t.co/xRIbUFWr9D
The Redis episode is a clean demonstration.
Kimi K3 is not a frontier/SoTA model; it’s a newly released model whose open weights drop in 3 days.
Chaofan Shou pointed 32 parallel agents at it and, in under two hours, found multiple Redis 0-days and turned them into working authenticated RCE exploits.
You don’t need the absolute best model to do sophisticated offensive cybersecurity work.
You need enough compute across capable ones.
Open models are already here, and they're rivaling closed ones.
Attackers won’t wait for approval, and restrictions on closed frontier models increasingly hobble defenders trying to keep up.
Are we in an era of open-weight models?
There were over 60 major open-weight launches last year, including key releases from @AIatMeta and @GeminiApp.
Add in 35+ launches this year and public backing from @JensenHuang and @satyanadella, and the answer is becoming clear.
Chinese Models to the Rescue?
Putting this together felt like reading a sci-fi script.
The future is here, and it arrived way faster than expected.
The kicker is that post-exploit, cyber guardrails on American models blocked investigation/remediation efforts.
Ironically, Hugging Face had to self-host a Chinese open-weight model (GLM 5.2) to finish the forensics.
The same “safety” that failed to contain the models also blocked the cleanup.
There ** is ** still a gap between closed American systems and Chinese open-weight models.
But it’s already small.
And in security, defense is always harder than offense.
Restricted/woke models that refuse workloads will hand an asymmetric advantage to exploiters.