Excited to release Hyra! It’s been less than 3 months since I joined Hunyuan. We turned Hyra into reality in a short time.
kudos to great colead @chengjiale4 and great boss @ShunyuYao12 and great Hyra team.
Looking forward to seeing Hyra optimize Hy models and Tencent products!
Introducing Hyra-1.0, the first version of Hunyuan Research Agent. 💡💡💡
Built to recursively improve solutions for performance-driven research and engineering tasks.
Explore our demos in AI4AI, AI4Science, and AI4Fun:
https://t.co/GClJMuT8XN
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work.
Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort.
Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
Introducing Hyra-1.0, the first version of Hunyuan Research Agent. 💡💡💡
Built to recursively improve solutions for performance-driven research and engineering tasks.
Explore our demos in AI4AI, AI4Science, and AI4Fun:
https://t.co/GClJMuT8XN
Introducing Harbor-Index, a compact, diverse, and high-quality benchmark built to challenge frontier agents.
We carefully select, audit and fix 82 high-signal tasks out of 6,627 candidates spanning 54 benchmarks.
No agent gets above 30%. (1/5)
HPC-AI is heading to #ICML2026! 🚀
Looking for top-tier compute efficiency or building the next generation of AI Agents? We’ve got you covered! Stop by Booth B301 to talk cutting-edge technology, supercharge your projects, and grab awesome perks!
🎁 Exclusive ICML Swag & Perks:
📱 Interact & Win: Engage with this post to grab our custom caps & eco-friendly canvas bags for FREE! 🧢💼
📝 1-Min Survey: Complete a quick questionnaire and claim your exclusive ICML badges & tech stickers! 🧲✨
🔋 Top-up Bonus: Top up your account during ICML and take home a premium Bluetooth speaker! 🔊
📝 Accepted Author Special: Got a paper at ICML? Unlock our Author Incentive Program to get exclusive vouchers for GPU compute and Model APIs! 🎓💡
⚡ Why Choose HPC-AI?
1️⃣ GPU Cloud Compute | High Performance & Elasticity
- Full-Spectrum GPU Coverage: From H200 to B200, we provide top-tier GPUs for training, fine-tuning, and inference.
- Flexible & Cost-Effective: Pay-as-you-go billing with ultimate price-to-performance ratio.
- Stable & Ready-to-Use: Instant deployment, elastic scaling, and enterprise-grade reliability for both corporate R&D and academic research.
2️⃣ Model APIs | One-Click Access to Frontier LLMs
- Mainstream Models Covered: Seamlessly access DeepSeek, GLM, Kimi, MiMo, and more, with continuous updates.
- High-Performance Inference: Built on enterprise GPU infrastructure, ensuring accelerated inference and high availability for massive workloads.
- Seamless Ecosystem Integration: Compatible with 40+ mainstream AI frameworks and tools (Dify, Cursor, Open Claw, Open Code, etc.). Connect in 3 easy steps!
📍 Booth B301 📅 July 7-10 🏢 ICML Conference (Korea Seoul)
#ICML2026 #HPCAI #MachineLearning #DeepLearning #GPU #ModelAPIs #LLM #AICompute
📣 Announcing Terminal-Bench Science: benchmarking AI agents on real scientific workflows – now open for task contributions👇
https://t.co/MSPMwnbhVt
@AnthropicAI, @OpenAI, and @GoogleDeepMind use Terminal-Bench to evaluate AI on coding tasks. We're now extending it to scientific workflows.
1/6🧵
Today, we’re releasing Continual Learning Bench 1.0: the first, realistic benchmark for measuring how AI systems can improve in online settings.
Benchmarks today assume models are stateless. Each example is independent, and once a system finishes a task, it moves on as if nothing happened.
But deployed AI systems should learn from experience. We tested 10+ frontier systems against novel, expert-validated tasks and find there’s still plenty of headroom for learning. (1/n)
Excited to release SimpleTES: a better open-sourced AlphaEvolve!
With gpt-oss, SimpleTES discovers SOTA solutions across 21 tasks: quantum compilation, LASSO speedup, scaling law discovery, kenel optimization...
Project: https://t.co/izrhyCSNsc
Code: https://t.co/K8cnKEfGlX
New paper: Spend Less, Fit Better
Fitting scaling laws for LLMs can cost millions💰-but what if you can get the same insights with just ~10% of the budget?
We frame scaling-law fitting as budget-aware experimental design and propose a method to pick the most valuable runs.#LLM
Winning the award for Highest Scored Rejected Paper at @iclr_conf . The AC is a harsh reviewer from NeurIPS and overrides all positive reviews and ignores every improvement made since NeurIPS.
Devastated because this is genuinely my best work yet.
https://t.co/LakaVdjJep
PS. The only negative scored review is the only "Fully AI Generated" review flagged by Pangram (https://t.co/X0PXhccI3l) . Other positive reviews are "Fully Human written". What an amazing rejection.
Excited to share that SLD has been accepted to ICLR 2026! 🎉
Also: SLDBench is now integrated into Harbor (https://t.co/H627UOpPm7), so you can quickly run any agents × LLMs with minimal setup. It’s much lighter-weight than SWE-Bench / PaperBench. Give it a try!
Sharing our interesting study on AI discovering its own physics! 🧪 Can AI scientists automatically discover scaling laws better than human experts?
We found that AI-discovered laws are not just more accurate: they are surprisingly interpretable and reveal patterns humans missed. 🤯
Check our blog for details: https://t.co/CZTER32jxO
Brilliant paper from Stanford + Tsinghua + Peking University + Wizard Quant
Shows an evolution style LLM agent can discover scaling laws that predict performance better than humans.
The big deal is that it turns scaling law writing from slow expert guesswork into an automated search that can guide expensive training and fine tuning decisions.
Scaling laws are simple formulas that guess how an LLM will do as it gets bigger, but experts still craft them by hand and they can fail in new settings.
The authors build SLDBench from 5,000 or more past training runs, and each task asks for 1 formula that predicts well on larger, unseen runs.
They propose SLDAgent, which keeps rewriting both the formula code and the parameter fitting code, testing each new version and keeping the best like an evolution loop.
This helps because the formula and the fitting method depend on each other, so improving only 1 often gives shaky predictions.
Across 8 tasks it beats human formulas on extrapolation, meaning prediction beyond the seen scale, and with GPT-5 its average R2 rises from 0.517 to 0.748.
The payoff is practical because it helps pick learning rate (step size) and batch size (examples per update) with fewer sweeps, and it helps choose which pretrained model to fine tune from small trial runs.
AI discovers its own scaling laws!📈
SLDAgent uses symbolic reasoning + evolution to find new laws that are more predictive/intuitive than widely-used expert-fit formulas. And they improve pretraining and fine-tuning!
Great work led by @AndyLin2001@haotian_yeee!
The Terminal-Bench paper is here! Read it to learn where frontier models still fail and the secrets of how we sourced hundreds of high quality environments from our open source community. 🧵
🤔Want a principled way to RL your diffusion model?
Check Data-regularized Reinforcement Learning (DDRL)! Post-train @nvidia#Cosmos World Foundation models with a million GPU hours! 🤯
Novel formulation ➡️ Theoretically integrates SFT into RL ➡️ Robust to Reward Hacking 🛑
Details: https://t.co/1A9q8ho2xb
#DDRL #Diffusion #RL #NVIDIA #Cosmos
Check our new blog post in collaboration with Algorithmic SuperIntelligence Labs about using OpenEvolve to discover scaling laws: https://t.co/zE5pO9N5pG