@_Suresh2 I agree, also no production level schemes were broken, but i guess long horizon (16 hours) in tier 2 at 20% means we’ll probably see models at that level by the end of this year
This work is not quite getting any attention. They evaluate Opus 4.8, Sonnet 5, Mythos 5, GPT-5.5, & GLM-5.2 on different primitive families (hash functions, block ciphers, AEADs, KEMs, PKEs, and digital signatures). Models achieve 86% on tier 1 (schemes with known attacks)
AI is advancing fast in math and cyber. Cryptanalysis sits at their intersection and underpins our digital security.
So can LLMs do cryptanalysis?
Increasingly, yes!
Introducing CryptanalysisBench where we test if AI can break 191 real schemes from past standardization efforts
On Tier 2 ( contains algorithms for which no attacks are known or attacks are too slow to run in practice), models achieve up to 10%, and models keep getting better at longer horizons
Dear BitMEX Users,
Today, we share with a very heavy heart that BitMEX exchange will shut down its operations, effective 23 September 2026 at 04:00:00 UTC.
The owner and operator of BitMEX, HDR Global Trading Limited, has made the difficult decision to close operations following a strategic review of the business.
It may not look the same today, but we are proud of our 11+ year legacy and the role we played in shaping the crypto industry. We invented the 100x leverage perpetual swap, which for most of you, was the first step to your crypto trading journey. It is now the most traded financial product in the crypto industry, adopted by thousands of users and exchanges. And we remain proud of our robust security infrastructure, which has allowed us to maintain a flawless track record of 0 customer funds lost to hacks in our entire operating history.
We want to reassure you that your assets remain fully safe and under your control during this transition period. This announcement is just to give enough time to ensure a smooth withdrawal process for everyone.
From today we strongly encourage all users to close their positions and withdraw their funds as soon as convenient. For more details on the full process, please read our blog: https://t.co/OOHeh6xHm8
BitMEX was once home to some of the greatest traders today. Our team has dedicated tremendous effort and passion into building the platform into what it is, and we are glad to have reached some of you during your time with us. To everyone who has traded, supported, and grown alongside us - thank you for your trust over the last 11 years.
The BitMEX Team
Lot of applications in TCS though, and as far as applied math goes you see it pop up everywhere. Graph theory as a pure math subject maybe is silly, but without appel & haken maybe the computer aided proof discussions in math would’ve been a decade late?
Tom leighton in his maths for cs course has very cool lectures https://t.co/XoptRYJehj
i dont get the doomers of knowledge work generally, as if having an oracle at a certain level doesnt mean you can ask more ambitious questions, maybe they are revealing their own inability
@stoicsavage 1) cant make good enough models that their own cluster gets saturated
2) there seems to be enough demand (we are currently on our way to consume a quadrillion tokens a month in a few years)
3) at their scale they get multiple year deals so good cash flow projection for business
The GenAI economy has generated $110 billion in sales over the past 12 months. It is growing fast. On an annualized basis, the revenue run rate exceeds $175 billion.
These numbers took us several months to construct, and as far as we know, it’s the first bottom-up, deduplicated measure of consumer and enterprise AI spending across the full stack.
We are releasing this research today in our first The State of the AI Economy report.
https://t.co/cJwZb0T99C
@saliencexbt I’ve heard some people have a large suite of tests and make sure the tests directory doesn’t have modifications/deletions, and additional behaviour/ux changes can be smoke tested. But i still don’t trust a loop with this
Im not entirely sure about this framing, feels like hindsight bias to some extent.
There was some exploration in 2003 on this front https://t.co/ydKIRvxGii which is the first llm in some sense but it was geared towards machine translation. But this was mostly trained with mapreduce, then you have alexnet in 2011 which sort of showed the feasibility of gpus, in 2014 you have more exploration of scaling working https://t.co/1w81sF7BzI and this was scaled further in 2016 iirc (dario was a part of baidu then). Whats missing here is two things, the absence of hardware and the general concept of scaling text. By mid 2017 you had v100, the transformer was released but it was originally released for machine translation not text modeling. You could say the elmo paper and then the gpt paper which sort of gave you the architecture, the downstream effects of text training and the hardware which leads you to gpt 2 and 3. I think its convenient to frame it as something missing, but its more like all of it came together, and you has the right team at openai which took the bet to scale it to 175b params. In 2016, what would they have scaled for? Machine translation, summarization and question answerinh had specific architectures, theres also the evaluation angle here where the need for a different architecture/idea came into place because of 2 benchmarks glue and decaNLP. And gpt/elmo/bert showed improvement on these benchmarks just by scaling text.
We evaluated recent open models on KellyBench.
Here is what we found:
🏆 GLM 5.2 is new open source SoTA, but still loses -30% on average over 5 runs.
📈 We estimate GLM 5.2 is 6+ months behind the frontier based on KellyBench and internal quant evaluations. (Note: we have not evaluated Fable)
🌗 Kimi K2.6 slightly improves on Kimi K2.5 but still struggles at -60% average RoI.
🐈 Recent Mistral models struggle, obtaining mean RoIs of -78% and -99% respectively.
Leaderboard link and more graphs below.
None, depends on how much you were counting on gemini as a part of future revenue. People in ai land always like to flip flop with leaderboards every few weeks, while consumers have shown no real stickiness and are happy to switch at the drop of a ball. Their cloud rev and compute keeps them in the game i think, and they always have the anthropic stake as a small hedge
@AcerFur if you look at lean for software verification, could be thought of as cyber adjacent, so they just block all of it, this is probably lesser than 0.001% of their usage