1/ LLMs are increasingly being used to power high-interaction honeypots while maintaining a low security risk.
But how good are they really? To answer this question, we introduce Honeyval, the first comprehensive eval framework for LLM-powered honeypots.
From day one, mimic has been focused on a single goal: general-purpose dexterous manipulation. Today we're proud to announce the mimic hand M1 and the mimic wearable U1.
We believe the only way to solve dexterous manipulation at scale is by going full-stack at the frontier of physical AI, building every layer ourselves around one fixed point, the human hand.
The M1 is a highly backdrivable, tendon-driven hand that covers the full range of human capability, from heavy payloads to fine manipulation.
If you are at ICML check out our oral presentation on Saturday at 11AM the DEMO workshop about adapting reasoning models with cheap and simple techniques.
(DEMO @ #ICML26 Oral)
RLMs are undoubtly powerful. But how can we train them further?
IFT on golden answers can destroy reasoning capabilities.
We show: With our careful merging technique, reasoning can be restored entirely, while keeping gains on the target domain! 🧵
(DEMO @ #ICML26 Oral)
RLMs are undoubtly powerful. But how can we train them further?
IFT on golden answers can destroy reasoning capabilities.
We show: With our careful merging technique, reasoning can be restored entirely, while keeping gains on the target domain! 🧵
Introducing two new research-level mathematical training datasets!
Training data for research mathematics, especially in the post-training regime, is severely lacking. Using our benchmark pipelines on a larger scale, we now created almost 6,000 training data samples.
1/ LLMs are increasingly being used to power high-interaction honeypots while maintaining a low security risk.
But how good are they really? To answer this question, we introduce Honeyval, the first comprehensive eval framework for LLM-powered honeypots.
Finally cleared the last hurdle!
Our latest quantization-conditioned attack works against almost every popular quantization method, including GPTQ, AWQ!
"Widening the Gap: Exploiting LLM Quantization via Outlier Injection"
https://t.co/UEq0Qeeqdn
1/ LLMs are increasingly being used to power high-interaction honeypots while maintaining a low security risk.
But how good are they really? To answer this question, we introduce Honeyval, the first comprehensive eval framework for LLM-powered honeypots.
6/ We open source Honeyval, hoping to standardize LLM-powered honeypot eval and provide a basis for incremental progress on LLM-powered honeypots.
Code: https://t.co/CPk32WYmgQ
Website + Leaderboard: https://t.co/fI93cvFH3R
Paper: https://t.co/BAStCDLaIa
LLMs have become capable of proving complex mathematics. However, the proofs they produce vary significantly in how clear, motivated, and insightful they are.
To measure these differences, we introduce ProofRank, the first benchmark to scalably evaluate aspects of proof quality.
Many papers conclude that an imperfect verifier has minimal impact on RLVR training. Is that really the case?
We show that, depending on the error pattern, the impact of verification error can be diverse, including delayed training, suboptimal plateaus, and complete collapse.
In a new blog post, we show that API errors and retry policies have significant impact on benchmark performance!
While retrying requests is ubiquitous in LLM evaluation, its effect on performance is undocumented, time-dependent, and leads to various incorrect conclusions.🧵
📣 new submission to SWT-bench
TEX-T by @SFResearch achieves 87% in script mode.
Amazing to see this benchmark hike along with SWE-bench from 15% to almost 90% in the last 1.5 years. Time for new unit test benchmarks :)
https://t.co/FQtCGOICAs
1/🧵 LLMs can write their own benchmarks to uncover security vulnerabilities!
We leverage LLMs to expand BaxBench with 40 entirely novel, complex web backend tasks, more than doubling the original benchmark, resulting in AutoBaxBench. These tasks include extensive test cases and end-to-end exploits to expose vulnerabilities in implementations, which we confirm match or even outperform human-written exploits.