The Wonderland CTF was a blast!
Huge congrats to all the teams, especially “STACK TOO DEEP”, “NADA ESPECIAL” and “SECSEE”.
Oh, also: https://t.co/WHMt1f36Mk 👉👈
We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners.
We outline what happened, how the activity was contained, and how we’re working with evaluators to strengthen our approach to third-party testing.
https://t.co/ZL3n6mxYMS
.@hosseeb on why beating North Korea's hackers now comes down to who can afford the most compute:
"If the attackers are spending 500 bucks, and good guys are spending maybe a couple bucks, they're like, oh, I just wanna check and see if there's any low-hanging fruit. Otherwise I'm not gonna use this code if I can tell it's obviously busted. So they spend a couple dollars worth of compute, and they say, ah, it looks good enough. It's a well-funded project. It's open source. Yeah, I guess I'll use it."
"But North Korea is spending 500 bucks, 1,000 bucks, 2,000 bucks just grinding and grinding and grinding, running multiple agents overnight trying to find an attack."
"That means that no matter how many good people are spending $2 worth of GLM compute on you, they will never be able to find the depth of the tree that North Korea is searching for. None of the normal users will ever get there, which means it's only really on the company."
"The company is the only party that can coordinate enough compute to be able to actually expand the search space enough to out search an attacker."
"But it does mean that attackers have the same problem. So if North Korea and Russia and China are all trying to hack your protocol, but they all spend, let's say, $5,000 each, if you spend $10,000 you will actually find everything that they can find and more."
"Which in that sense, it benefits defenders over attackers in the long run, because the company can actually outspend any individual attacker in principle, and that used to be not sufficient defense in the old model because of the fact that attackers were uncorrelated with each other."
@dragonfly_xyz
I recently got a chance to test @Certora AutoProver. Using frontier models to write your spec is cool, but i believe breakthroughs in prover tech and frontier model capabilities will increasingly close the gap on what’s possible. Maybe we need a formal verification benchmark?
Robinhood CEO @vladtenev believes mathematical superintelligence could eliminate software bugs altogether.
"Math skills and coding actually go hand in hand... In the same way that MSI can be used to make sure that your math is correct, when you have a computer program or a piece of software, it could be used to make sure that's correct and that there are no mistakes."
"Rather than it being a cat-and-mouse game of escalating model capabilities, finding breaches and preventing breaches from happening, verification and proof that software is immune to these types of bugs, I think is the future."
They go from charging tokens to charging outcomes. Your harness fetches a quote from multiple providers that guarantees n outcome at a price/ duration. See @Tesla auto-bidder
What's next for Anthropic?
1. GLM 5.2, Cursor / xAI's and Cognition's models are Opus-level at up to 5x cheaper
2. Most software workloads live at Opus-level capability.
3. Fable is 2x Opus pricing, and routes to Opus most of the time anyway because of the cyber filter.
4. Budgets are tightening. Hard to see spend going up on Fable.
5. Cursor and Cognition have the distribution and will use cheaper Opus-level models to win customer workloads.
Will we see Anthropic competing on cost?
Researchers from Berkeley and Princeton are partnering with Eigen Labs to launch a suite of open science autoresearch challenges together on Frontier CS.
The paper is being presented at @icmlconf in Seoul today. If you’re there, join the researchers at Hall A 502 from 2:30-4:15 PM local time to discuss.
The challenge is live globally: https://t.co/p03zkNqb6Q
The goal is to create the lowest resource circuit that computes elliptic curve point addition correctly on 99% of inputs. The benchmark is running 9024 random points that sampled deterministically based on a circuits hash. This is fine because Schor's algorithm fails at a rate similar to its subcircuits: it works approximately.
The probability that a 99% correct circuit makes it through this check is 0.99^9024 ~= 2^-130. Tiny.
So grinding is actually a valid (and seemingly effective) strategy! It's a essentially a way for a 99.9xxx% correct circuit to tell the benchmark that it wants a new set of points to be tested on in order to prove that it's correctness exceeds 99%.
This is the exact same test Google thoughtfully designed into their ZK proof, read here for more details:
https://t.co/Ttatgf9H3d
https://t.co/nU7BwCUZym
Open agents just broke Google’s closed quantum benchmark in public.
ECDSA. Google’s quantum circuit. Post-quantum future.
@sreeramkannan and @bbuddha_xyz from @eigenlabs will be joining @MTSlive to break it down shortly.
My MLSys keynote on AI writing systems code got more interest than I expected. The recording will take a while, so in the finest tradition of AI labs sharing blog posts, we’re starting the Core Automation Blog with this one https://t.co/h4uSOyrglf