NEW: malware developers added nuclear & biological weapons text to to their spyware.
Goal? To trigger LLM safety refusals... so that their spyware wouldn't be analyzed by an AI security scanner.
Cleanest practical example I can think of for why over-indexing on first order safety alignment is risky.
When closed (and open) models ship with aggressive refusals, they will be sprinkled with second-order blindspots that attackers will discover...and exploit.
We are only in the earliest days of attackers leveraging these features, and it wouldn't surprise me if users systems that need to handle complex cybersecurity issues demand that models be less safety-blunted.
In the weeds: @SocketSecurity's post also shows why intention matters in how you design a malware analysis pipeline to avoid prompt manipulation.
H/T to colleagues that shared this with me https://t.co/f3Aj9TYxU4
The real power of forward deployed engineering has always been putting strong technical people directly alongside the operators who own the outcome. That proximity forces the work to solve the actual problem instead of some sanitized version of it.
In the AI era this principle has become even more valuable. Agents can now sit inside real workflows and improve from actual decisions, which means the highest-leverage work is extracting the tacit knowledge that lives with subject matter experts, building evaluations that reflect how things actually break, and closing the production feedback loop so agents get better from real outcomes.
The part that makes this so lucrative is most CEOs know AI matters but don't trust their own team to tell them the truth about where they actually are.
The fractional CAIO isn't selling expertise.
They're selling an honest outside perspective that nobody internally has the incentive to give.
The person who can walk into a company, skip the politics, and say 'here's what's real and here's what's theater'…that's worth $4M.
HEADLESS
The reason more “dashboard” software companies will push to become headless stores (that is, become “pipes” companies) is for two reasons:
(A) one of the most durable opportunities in software is to become the singular data store for all structured or unstructured context for a certain domain or function (or heck, for the entire organization).
(B) once a function or org embraces you as the data store, you price based on compute and storage, and beautifully compound and grow as the org context inevitably explodes over time, the way other compute/storage businesses like Databricks and Datadog have durably and steadily compounded over time.
Now, of course, the challenge is that as soon as you open up to 3rd party agents, some of these agents, being venture funded companies themselves and run by ambitious entrepreneurs, will try to suck all the context out of you, reduce you to a CRUD database, and ultimately offer a version of you to their customer for free, bundled with their agents. Meanwhile, you will likely offer your own bundled free or low priced agents as alternatives to the most commonly used third party agents.
This war between headless context stores and agentic workflow startups is just starting. Will be interesting to see who can commoditize their complement first.
The analogy to chess openings is interesting. Seems (and let me know if you agree / disagree) like a lot of math proofs fall into this category of 'exhaustive search + being able to think X steps ahead + some heuristics to discriminate whether a branch is worth exploring' type of strategy, which is the same for chess engines. If you can evaluate two million positions per second and apply a filter, you can evaluate a lot of positions that human brains simply wouldn't have capacity to do in that amount of time cause they're not calculators.
It's an important distinction, because human brains evolved to do certain types of calculations extremely efficiently (pattern recognition, kinetics, modeling) while being pretty poor at others (e.g. multiplying two large numbers together).
But in some cases, it's the pattern recognition that actually holds us back, in a way. It works extremely well up to a point, and then blinds us to some other solution that doesn't fit the pattern.
And that's where the 'wow' moments have happened with chess engines. It's not the times they are able to calculate 12 moves deep while humans can only do 10 moves. Like okay, fine, we get it, and it's expected. It's the times that all our intuition says that a certain move makes no sense, because it breaks from the pattern, while the chess engine wasn't so reliant on pattern recognition and therefore didn't write off that branch / line right away before exploring it deeper.
And if this maps to math proofs, then we would expect the entire population of math proofs which only require breaking away from the pattern to be solved by AI in the near future. At the limit, we might say 'If it's provable using known methods, AI will find the solution given enough time / computational effort.'
And then, eventually, we would have a bunch of conjectures that can't be proven using known methods. And we would run out of methods that everyone agrees on.
And that's where it could get super interesting. Because then we could say 'These additional conjectures can be proven if one allows these additional methods'. So in a way, those methods become like new axioms. We'd be discovering new potential axioms through the proof process rather than thinking them up.
What could this lead to? Who knows. But which systems / models of axioms we find interesting / meaningful is generally informed by our experiences. So we (royal we, I mean you, the mathematicians, I do not consider myself one) bring the meaning and can pluck any additional axioms right from the tree. The axioms which most pertain to our reality.
A system (in this case, our universe) can never prove its own axioms. But it can make some educated guesses.
If I had to bet, I think we ultimately keep the axiom of choice but we do away with the axiom of infinity, but we do add some other axiom(s) to replace infinity, and I have no idea what those are.
But to wrap it up, I don't think proofs are where the art of math lies. They're a necessary step to get at the foundations and truly understand what isn't reachable to us otherwise - what's below the floorboards.
This is an amazing blogpost but the conclusions that fall out of it are quite grim.
1. If pre-training from scratch is required, then robotic capabilities will scale as a function of operations expense - linear with respect to number of people employed, number of collect hardware - etc. This leaves most American companies in a hard place as China is better positioned to scale ops due to large scale manufacturing and cheap labor.
2. That we cannot bootstrap off of priors from AI models like VLMs and World Models foreshadows a grim foretelling for the whole field of robotics. If we cannot re-use the internet, leverage the collective human experience already accumulated, we are cooked. Collecting the entirety of human experience on hardware again via tele-operators will take many years if not decades and this would mean physical AI will scale far too slowly. Breakthroughs that allow us to use human data / YouTube / RL/ models like Gemini and Veo - that scales capabilities as a function of FLOPs, data tokens(even if not from robot) - will then become very valuable to break us out of a slow scaling paradigm.
Very cool! 1. Regarding accelerating clinical trials please see the ideas of @RuxandraTeslo for example https://t.co/VBQzu8CBg0 and the clinical trial abundance effort. 2. There are right to try laws in Montana and New Hampshire to let people try drugs after they are proven safe (phase I). 3. For sure! I'm taking a personalized mRNA and we're making custom ADCs, see https://t.co/NbzApDUqop 4. ctDNA can be a faster way to measure disease, see https://t.co/ep3jYxfvie > MRD for my data
I believe we've found the best AI-native coding interview
We call it the “Composer 1 interview”
Candidates get 1 hour to build a real, medium-sized project live
The only constraint: they have to use Cursor’s Composer 1 model
8 articles I read last week that changed how I think about what it means to build a product in the AI age:
1. Building for trillions of agents by @levie https://t.co/YiC3l6cYWO
2. How Coding Agents Are Reshaping Engineering, Product and Design by @hwchase17
https://t.co/eY0MYATuY0
3. Services: The New Software by @JulienBek https://t.co/ZnQtR5nxXI
4. Lessons from Building Claude Code: Seeing like an Agent
By @trq212 https://t.co/dlgzLNYrtC
5. How apps don’t get killed by Claude By @michlimlim https://t.co/xzyzTEN2K5
6. The 10x Lawyer by @zackbshapiro https://t.co/oqdgRKK2gr
7. The Claude-Native Law Firm by @zackbshapiro https://t.co/ECQd2C1wp6
8. When Your Life’s Work Becomes Free and Abundant by @adityaag https://t.co/Ad51Pm7apH
Got tired of managing all client context in my head, now an agent read the transcript after the meeting, creates the linear issues and spin up the cursor to fix it.
Will let you know how it plays out