Can video generation models do for vision what LLMs did for language?
Introducing GenCeption from @GoogleDeepMind: one feed-forward video model for various vision tasks — SOTA, data-efficient, and emerging behaviors (ECCV 2026)
🌐 https://t.co/3rwVgTolZz
(1/8)
🎊Excited to share that our research on validating AI-generated social science data is now online at @PNASNews !
Check it out: https://t.co/ljqoTBVBxB
Great thanks to my collaborators!
💡In this work, we highlight population-level statistical realism as a core criterion.
Our research has been well-represented at ICLR'26 🇧🇷, with the following papers on reasoning and agentic paradigm improvements. 🎉
1️⃣Dancing in Chains: Strategic Persuasion in Academic Rebuttal via Theory of Mind
2️⃣Diversity-enhanced reasoning for subjective questions
3️⃣Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4
4️⃣Webwatcher: Breaking new frontier of vision-language deep research agent
5️⃣Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
Check it out further offline in case you missed the poster session earlier~ 🍻
Introducing Nemotron-Terminal: a systematic data engineering pipeline for scaling LLM Terminal Agents.
We bridge the gap between open models and proprietary models with a fully open synthetic-to-real trajectory pipeline.
🤯The payoff: SFT on our Nemotron-Terminal-Corpus boosts Qwen3-32B from 3.4% → 27.4% on Terminal-Bench 2.0 (+24.0), rivaling models multiple its size.
What makes it work?
🌟Terminal-Task-Gen: A lightweight data curation pipeline that seamlessly combines the adaptation of existing datasets with robust synthetic task construction.
🌟Nemotron-Terminal-Corpus: A massive, open-source dataset covering diverse terminal interactions, which contains explicit planning and execution traces for complex long-horizon tasks.
And we’re releasing everything:
📦 Nemotron-Terminal-Corpus (Large-scale dataset)
🤖 Nemotron-Terminal models (8B, 14B, 32B)
Paper: https://t.co/ORIZ01sav1
HF Daily: https://t.co/nSH4hu7I5D
Models & Data: https://t.co/J1Zc22M95r
Our tech report just hit the #1 spot on Hugging Face Daily Papers!
We're also incredibly excited to see the open-source community putting our work to the test, with the Nemotron-Terminal-Corpus dataset currently trending at over 1,800 downloads and counting.
We can't wait to see what the community build with it!
Excited to share our new paper "One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control" on arXiv. With our novel designs of Unified Masked Conditioning (UMC) and Decoupled LoRA Control (DLC), One4D can seamlessly handle single-image-to-4D, sparse-frame-to-4D, and full-video-to-4D tasks in a unified model by generating high-quality RGB frames and accurate pointmaps.
Project page: https://t.co/QMGkHd7xWK
arXiv: https://t.co/v1urLAMkPw
Code to be released at: https://t.co/th4P7gdmdM
Huggingface Paper: https://t.co/bZYtt83OGY
Many thanks to our collaborators Prof. Dan Xu @danxuhk and Yuxin Wang.
The source code and checkpoints of ThinkDiff are finally released! We also released a modified vLLM for embedding. Feel free to use them and raise an issue if you have any problem!
Github: https://t.co/PgvHdytanj
Checkpoints: https://t.co/NtPBk9Pkpb
vLLM: https://t.co/jbyiIfVhWH
This new system called AlignGuard helps make sure AI image generators create pictures that are safe and appropriate. It uses a team of smaller models working together and suggests better training methods to keep things respectful. Tests show it works well, making images safer without sacrificing quality. The creators recommend using it after training an AI model as a final safety check before releasing it to the public.
#AI #ImageGeneration #Ethics
https://t.co/TOOTrzZ2SJ
#ArtificialIntelligence
🔥🚀 Native image generation of Gemini is just on fire!
Our #ThinkDiff paper bridges vision-language models with diffusion models to unlock native image generation from VLMs. While not yet as perfect as #Gemini, it offers a glimpse of the groundbreaking potential!
🔥 Check out:
🔗 Project: https://t.co/QKjEHeiIBy
📄 Paper: https://t.co/M6JxdgcS3H | https://t.co/u5VvAP1k44
💻 Code to be released at: https://t.co/NVMi3x02gV
Check out our awesome test cases:
1. Generate a style consistent and logic-correct image by in-context reasoning:
🌟 Amazing work! SuperGPQA spans 285 graduate-level disciplines with 26,529 multiple-choice questions, each averaging 9 options per question. 📚💡 Congrats to @GeZhang86038849
🎉 Huge congrats to UTMath for being selected as one of the 15 reliable resources by SuperGPQA experts! From our UTMath test set, 401 questions were included, with 290 high-difficulty questions—that's nearly 1/5 of the toughest math problems in the dataset. 🎓✨ An evidence to the quality of UTMath! 🔢👏
Is verifying a correct answer through multiple-choice options truly sufficient for measuring graduate-level knowledge? 🤔 What about the risk of oversimplifying complexities, or the challenges of designing high-quality distractors? These are questions worth exploring. 📚🌟
#AI #Education #SuperGPQA #Mathematics #UTMath
Great work! 🎉 Our UTMath also identified this but it skips manual transformation. Each problem has input-output pairs forming unit tests ✅ for correctness.
Tested on UTMath GPT4o scores 26% 📉, o1-mini 29% 📊, proving the test's difficulty. 🔥Pls check: https://t.co/EeFW9PThvl
o1-preview shows a whopping 30% reduction in accuracy when Putnam math problems are slightly variated.
Not sure about the reliability of the results but I think it's a good paper worth checking out to understand LLM robustness on complex math problems better.
A hot research topic right now is to understand better whether these models can "reason" robustly or if they are just relying on memorization.
Robustness is important as it's key to model reliability.
@omarsar0 Great work! 🎉 Our UTMath also identified this but it skips manual transformation. Each problem has input-output pairs forming unit tests ✅ for correctness.
Tested on UTMath GPT4o scores 26% 📉, o1-mini 29% 📊, proving the test's difficulty. 🔥 Check: https://t.co/EeFW9PThvl
🚀Excited to introduce our work UTMath: Math Evaluation with UNIT TEST via RCoT🧠💡
🧠With an average of 68 test cases per problem, UTMath ensures that the model truly solves the problem.
💡We release our benchmark and training dataset (~70k).
Explore: https://t.co/Ga3QyzI487
🔥Introducing Latent Guard from Oxford University and HKUST, a customizable T2I safety framework at test time. We release the largest T2I safety dataset with 50k~170k prompts of 723 unsafe concepts in 7 categories.
🌟Code&Data:https://t.co/4VTfui7Yu2
🌟Project: https://t.co/SOxw4rbQ51
🌟Paper:https://t.co/p6mtq49316