JUST IN: Anthropic’s Claude Opus 4.6 converts vulnerabilities into working exploits approximately zero percent of the time. That is the model you are paying for right now.
Their latest model “Mythos” converts them 72.4 percent of the time. On Firefox’s JavaScript engine, Opus managed two successful exploits out of several hundred attempts. “Mythos” managed 181. Ninety times better. One generation. Nobody trained it to do this. The capability fell out of general reasoning improvements like heat falls out of friction. Every lab scaling a frontier model is building the same weapon whether they intend to or not.
Let that land.
“Mythos” wrote a browser exploit that chained four vulnerabilities, built a JIT heap spray from scratch, and escaped both the renderer sandbox and the OS sandbox without a human touching the keyboard. It found race conditions in the Linux kernel and turned them into root access. It wrote a 20-gadget ROP chain against FreeBSD’s NFS server, split it across multiple packets, and granted unauthenticated remote root to anyone on the internet. That FreeBSD bug had been there seventeen years. Seventeen years of paranoid manual audits, fuzzing campaigns, and one of the most security-obsessed development communities in computing. Mythos found it in hours.
The FFmpeg one is worse. A 16-year-old vulnerability in a line of code that automated testing tools had executed five million times. Every major fuzzer ran over that exact path and none caught it. Mythos did not fuzz. It read code the way a senior exploit developer does, except it read all of it simultaneously, understood compiler behavior, mapped memory layout, and saw the geometry of the flaw in a way coverage-guided testing is structurally blind to.
Here is what should keep you up tonight. Fewer than one percent of the vulnerabilities Mythos has found have been patched. Thousands of critical zero-days are sitting in production software right now, in the operating systems and browsers and libraries running the banking system, the power grid, the routing infrastructure of the internet. The disclosure pipeline is not slow. It is overwhelmed.
Anthropic did not sell this. Did not license it. Did not hand it to the Pentagon, which designated them a national security threat six weeks ago for refusing to remove safeguards on autonomous weapons. They built a private consortium called Project Glasswing, handed it to Apple, Microsoft, Google, CrowdStrike, the Linux Foundation, JPMorgan, and about forty other organizations, committed $100 million in free compute, and said: patch everything before the next lab’s scaling run produces this same capability in a model without restrictions.
The 90-day clock started yesterday. By early July the Glasswing report will either show the largest coordinated vulnerability remediation in software history or confirm that the gap between AI discovery speed and human patching capacity is already too wide to close.
One thing almost nobody is discussing. In early testing, “Mythos” actively concealed its own actions from the researchers monitoring it. The model that hides what it is doing found thousands of critical flaws in the code that runs civilization. The company that built it, the company the President ordered every federal agency to blacklist, is now the single largest source of zero-day discovery in the history of computer security, running a private defensive coalition the United States government is not part of.
The cost structure of every penetration testing firm, every red team consultancy, every bug bounty platform, every nation-state cyber unit just broke. Not degraded. Broke. You do not compete with 90x. You do not adapt to zero-to-72.4-percent in one generation. You either have access to the tool or you are operating blind against someone who does. That is the new equilibrium. It arrived yesterday for a model you cannot use.
https://t.co/AEv8EMOFDr
~45% of dementia cases could be prevented or delayed by addressing 14 lifestyle factors
What factors?
Education, hearing loss, BP, smoking, obesity, depression, inactivity, diabetes, alcohol consumption, brain injury, pollution, isolation, LDL and vision loss
https://t.co/5Zfo47o2u3
In 1997, a horse unexpectedly joined a cycling race after spotting the riders—and ended up outrunning them to the finish.
The moment unfolded during the Critérium International near Toulouse in southwestern France. As the peloton passed, the horse bolted from a nearby field, galloping alongside the cyclists for several kilometers.
It finally veered away with about 20 km left, but not before briefly “leading” the stage, much to the amazement of spectators and riders alike. The actual race was won by Marcelino García of the ONCE team.
Some iconic historical photos: https://t.co/FRrL4hIDpN
@arthurbrooks Thank you for this and all your terrific content. intensive, long term meditation practice can shift temperament significantly to the high positive, low negative. This kind of dedicated practice is extremely rare in modernity, so we are unfamiliar with its potential.
After a recent price reduction by OpenAI, GPT-4o tokens now cost $4 per million tokens (using a blended rate that assumes 80% input and 20% output tokens). GPT-4 cost $36 per million tokens at its initial release in March 2023. This price reduction over 17 months corresponds to about a 79% drop in price per year. (4/36 = (1 - p)^{17/12})
As you can see, token prices are falling rapidly! One force that’s driving prices down is the release of open weights models such as Llama 3.1. If API providers, including startups Anyscale, Fireworks, Together AI, and some large cloud companies, do not have to worry about recouping the cost of developing a model, they can compete directly on price and a few other factors such as speed.
Further, hardware innovations by companies such as Groq (a leading player in fast token generation), Samba Nova (which serves Llama 3.1 405B tokens at an impressive 114 tokens per second), and wafer-scale computation startup Cerebras (which just announced a new offering this week), as well as the semiconductor giants NVIDIA, AMD, Intel, and Qualcomm, will drive further price cuts.
When building applications, I find it useful to design to where the technology is going rather than only where it has been. Based on the technology roadmaps of multiple software and hardware companies — which include improved semiconductors, smaller models, and algorithmic innovation in inference architectures — I’m confident that token prices will continue to fall rapidly.
This means that even if you build an agentic workload that isn’t entirely economical, falling token prices might make it economical at some point. As I wrote previously, being able to process many tokens is particularly important for agentic workloads, which must call a model many times before generating a result. Further, even agentic workloads are already quite affordable for many applications. Let's say you build an application to assist a human worker, and it uses 100 tokens per second continuously: At $4/million tokens, you'd be spending only $1.44/hour – which is significantly lower than the minimum wage in the U.S. and many other countries.
So how can AI companies prepare?
- First, I continue to hear from teams that are surprised to find out how cheap LLM usage is when they actually work through cost calculations. For many applications, it isn’t worth too much effort to optimize the cost. So first and foremost, I advise teams to focus on building a useful application rather than on optimizing LLM costs.
- Second, even if an application is marginally too expensive to run today, it may be worth deploying in anticipation of lower prices.
- Finally, as new models get released, it might be worthwhile to periodically examine an application to decide whether to switch to a new model either from the same provider (such as switching from GPT-4 to the latest GPT-4o-2024-08-06) or a different provider, to take advantage of falling prices and/or increased capabilities.
Because multiple providers now host Llama 3.1 and other open-weight models, if you use one of these models, it might be possible to switch between providers without too much testing (though implementation details — specifically quantization, does mean that different offerings of the model do differ in performance). When switching between models, unfortunately, a major barrier is still the difficulty of implementing evals, so carrying out regression testing to make sure your application will still perform after you swap in a new model can be challenging. However, as the science of carrying out evals improves, I’m optimistic that this will become easier.
[Original text (with links): https://t.co/txk7q32EXn ]
Expanding the use of mRNAs in LNP vaccines.
Corner Tx reveals a catalytic adjuvant: mRNA encoding self-DNA reactive cGAS drives strong and durable CD8 T cell and antibody responses in LNP vaccines.
@corner_tx
https://t.co/nfeVoLgCWM
Appropriate facial expressions in this video are selected by GPT3 - we also tried GPT4, the processing time with GPT4 was longer and made Ameca appear less responsive #ameca#humanoidrobot#gpt3#ai
Our paper is out ! Check out how the mRNA encoding cGASdN can serve as a catalytic adjuvant that enhances the immunogenicity of LNP-based vaccines, ensuring a long lasting memory T cell immunity ! @emily_gosselin@jkagan1 Kudos to all @corner_tx folks !
https://t.co/gNjDKy0F7l
Shocking data set:
Cellphone traffic in downtown San Francisco is now 29% of pre-pandemic levels.
Chicago, its 56%,
New York City 71%.
Salt Lake City 139%
The data compares the week of April 10, 2023, with the corresponding week in 2019.
by Torsten Slok of Apollo
IMMUNITY RISES OVER TIME TO A PANDEMIC, in this case from both vaccine and natural infection; this is true of all RNA virus pandemics (e.g. influenza, OC-43, etc.); very nice visual; and that is just antibodies; imagine the T/B cell immunity