Excited to share what we’ve been building — Hyra-1.0, our first step toward AI systems that can improve research and engineering solutions through iteration. 🚀
Introducing Hyra-1.0, the first version of Hunyuan Research Agent. 💡💡💡
Built to recursively improve solutions for performance-driven research and engineering tasks.
Explore our demos in AI4AI, AI4Science, and AI4Fun:
https://t.co/GClJMuT8XN
Love Claude Design, hate daily limits? 🛑
Meet Open CoDesign — the open-source, desktop-native AI design and web dev workspace.
🚀 Use the world’s most advanced models: GPT-5.4, Claude Opus 4.7, GLM-5.1, Gemini-3.1, or your own local models.
BYOK. No vendor lock-in. No usage-limit popups.
It’s Claude Design, but fully open. 🔓
🔗 Star on GitHub: https://t.co/KUxBuszXts
#OpenSource #AI #WebDev #BuildInPublic
Introducing GLM-5V-Turbo: Vision Coding Model
- Native Multimodal Coding: Natively understands multimodal inputs including images, videos, design drafts, and document layouts.
- Balanced Visual and Programming Capabilities: Achieves leading performance across core benchmarks for multimodal coding, tool use, and GUI Agents.
- Deep Adaptation for Claude Code and Claw Scenarios: Works in deep synergy with Agents like Claude Code and OpenClaw.
Try it now: https://t.co/WCqWT0qCQb
API: https://t.co/xDy1O6ZPcz
Coding Plan trial applications: https://t.co/qCM6cri0KK
What you see in the video:
1️⃣ Autonomous Deep Research: Using Tavily to scan 2024-2025 papers.
2️⃣ Code Execution: Writing and running its own evaluation harness.
3️⃣ Analysis: Generating plots and Markdown reports.
This is why we call it the AI R&D operating layer.
🔗 Project: https://t.co/Wa5NxZVEQi
Can an AI discover better prompts than a human? 🤖👨🔬
I asked PhD-Zero to investigate the optimal prompting strategies for Qwen-1.7B.
4 hours of autonomous research later, it didn't just find tricks—it produced a 10-page report with counter-intuitive insights that I hadn't even considered.
🎥 Watch the full end-to-end R&D loop:
#LLMs #autoresearch
The most fascinating finding? "Thinking Compression." 🧠⚡
While we often default to long Chain-of-Thought (CoT), PhD-Zero proved that for 1.7B models, concise reasoning actually outperforms verbose paths.
In small models, excessive "thinking" introduces logical drift. It literally out-researched me on the "physics" of its own kind. 🤯
The goal of PhD-Zero is to build the AI operating layer for R&D—equipping agents with specialized research skills instead of just writing code.
90% autonomous workload, but 100% human-directed. Check out the project and our vision here: 🔗 https://t.co/Wa5NxZV70K
The era of Autonomous AI R&D is here. 🚀
What happens when an AI Agent leads the R&D?
0% → 20.0% on AIME25 in just 48 hours. 🧠🚀
I deployed PhD-Zero (powered by Codex) as a 24/7 AI research intern to optimize Qwen-1.7B-Base.
11 iterations. 90% hands-off R&D, requiring only human confirmation at key milestones. Zero sleep, just high-speed execution. 🧵 on how this intern "thought" its way through:
cc @karpathy@bindureddy@Alibaba_Qwen
2/ The "Aha!" moment: Step 5. 🛠️
The agent autonomously identified a loss_mask mismatch that I would overlooked.
After the fix, supervised length jumped from 688 to 6552 tokens. The model finally "clicked" and started learning the actual reasoning. This self-healing is what separates an Agent from a script.
@karpathy Love the loop! I’m building PhD-Zero to give Claude Code/Codex a "PhD Skill-set."
It turns them into 24/7 autonomous researchers by standardizing the R&D workflow.
Let the agents handle the 2 a.m. experiments. 🧠⚡
🔗 https://t.co/KEQBLTgisM
What if Claude Cowork became truly open — and worked on both Win/macOS? 🧰🚀
Meet Open Cowork: an open-source Agent SDK desktop app with sandbox permissions, skills for PPTX/DOCX/Excel, and traceable execution.
Repo: https://t.co/pv6EL0udGD