Today we're releasing IQuest-Q1 and opening the model weights. 320B total. 15B active. Built for code, software engineering, and complex agentic tasks. Technical report/HuggingFace/GitHub — available now.
📦 Deliveries still stalled: robots that had finished unloading were blocking the packing station. Its first fix moved the jam to the entrance—so Q1 went back and revised it again.
On unseen orders, the regular-day and blocked-aisle runs delivered every order with zero collisions. On the 420-order peak run, Q1 delivered 87% within the time limit.
Hugging Face:https://t.co/YxjyR7Bm3Q
GitHub:https://t.co/W3xkaabQMf
🤖 Next up from the community: 8 warehouse robots, hundreds of orders, one packing station, and an aisle that closes mid-run. Could IQuest-Q1 keep deliveries moving?
Q1 built a dispatcher and ran it. The logs showed robots colliding head-on, so it fixed the right-of-way rule.
🚑 Q1 measured ~18s between lights, used the offset to build a green wave, staggered the second row N–S, and inserted ambulance priority on top.
🚦 After 54 minutes of testing and revision, the reported average delay fell from 6 mins to 22 secs per vehicle. The ambulance crossed town in just over a minute, down from nearly four.
Hugging Face:https://t.co/YxjyR7Bm3Q
GitHub:https://t.co/W3xkaabQMf
🥰 Since releasing IQuest-Q1, we’ve loved seeing developers put it to work and share their experiments with us.
One community developer challenged Q1 to improve a traffic sim: 6 intersections flipping in lockstep, 30s green each way. Asked it to do better.
Reward curve drops after an env update. Ambiguous — model regression, or broken infra?
Model clusters failures. Traces them to pre-patch env faults being scored as task failures — invalid signal entering policy updates.
Filters bad evals out. Signal recovers. Training resumes.
RL training run misbehaves. The reward curve looks wrong.
The model reads the curves, pulls the logs, forms a hypothesis, finds the bug — a stray whitespace breaking the generation chain — and patches it.
Mean reward: 0.704 → 0.769.
IQuest-Q1 is trained in three stages:
Pretraining → mid-training → post-training.
Post-training focuses on software engineering, long-horizon agentic tasks and general reasoning, combining SFT + reinforcement learning.
In post-training, IQuest-Q1 uses MOPD(Multi-Teacher On-Policy Distillation) to consolidate strengths from multiple specialist teachers through the student model’s own on-policy rollouts.
This brings specialized capabilities into one model without simply inheriting any single teacher’s bias profile.
Numerically solved the equations of motion (100k steps). Rendered the three orbits as glowing silk threads in Three.js.
Klimt-style. Gold, ochre, dark brown. Click anybody to inspect mass and velocity.
Cold, sharp, sci-fi. Deep black space, glowing neon track, anti-gravity craft banking hard through corners.
Full game loop. Inertia camera. Thruster FX. Pure front-end. WebGL + Three.js.
One prompt.
The challenge isn't just building a racing game. The world has to extend forward coherently as the vehicle moves — foreground and background elements swapping in and out.
IQuest-Q1 generates that kind of spatial continuity end-to-end.
Image + natural-language description → full 3D scene.
Explicit spatial structure, scene hierarchy, environmental relationships. Paths, landscape features, guided routes, scene interaction.
One prompt.
Prompt: a multiplayer 3D FPS set in a children's bedroom.
One generation: Full game. Boss raid, in-game economy, HUD, lobby, building-block waiting room, paper-airplane spectator view after death.
No scaffolding. No iteration.😀
Today we're releasing IQuest-Q1 and opening the model weights. 320B total. 15B active. Built for code, software engineering, and complex agentic tasks. Technical report/HuggingFace/GitHub — available now.
Can medical AI ground its reasoning in evidence?
UniReason-Med, accepted to the Main Conference of #EMNLP 2026, brings Grounded Chain-of-Thought to 2D medical images and 3D CT volumes.
UniReason-Med interleaves textual reasoning with spatial coordinates and region-level visual evidence-no external detectors or segmenters at inference.
We also release UniMed-CoT, a 220K-sample grounded reasoning dataset for 2D and 3D medical imaging.