I founded Apex Intelligence today! We are targeting building self-evolving models for unlocking undiscovered discovery by AI autonomously. We believe the future research will be done mostly by AI Scientists!
MemoryArena is live at ICML! Our collaborators @YzhuML and @rlyin0171 are presenting our work — go say hi and grab them for a coffee chat if you’re around 😉
🤟MemoryArena got 16K downloads on Hugging Face last month, so I guess the community has spoken:
memory eval should be more than static recall -- agentic memory needs to be test in temporally/casually dependent, multi-session tasks as how they are actually used in the real world.
Data: https://t.co/zCSPb2qQQV
Paper: https://t.co/slZyxgmMAO
Project page: https://t.co/2OvkxXhrH2
#ICML2026
🤔Does ~99% accuracy on OmniDocBench mean VLMs have solved document parsing and understanding?
❗️Or is OmniDocBench just too easy?
Check out our Dr.DocBench — a “pro version” testbed designed to diagnose VLMs brittleness on more challenging document parsing and understanding tasks. Full evaluation pictures along the axises of performance and token efficiency.
Dr. DocBench's evaluation harness is now live on GitHub.
It extends OmniDocBench with a multipage sliding-window pipeline and subject-level granularity, scoring text, tables, formulas, and reading order across a configurable page window. Ready-to-run inference for 9 models, plus three block-matching strategies for prediction-to-GT alignment.
💻 https://t.co/Oj18q1nWv6
📄 https://t.co/4BPc1SKRs1
#ACL2026 #ICML2026
How should we train LLMs for user simulation? A good user simulator shouldn't merely imitate what someone said. It should be indistinguishable from what users could have said.
Introducing 🎭Turing-RL: training user simulators with a Turing-Test reward. The model gets rewarded when an LLM judge can't tell its response from the real users’.
Thanks for sharing our 𝙈𝙚𝙢𝙤𝙧𝙮𝘼𝙧𝙚𝙣𝙖!
If you’re feeling that recall-style benchmarks (e.g., locating facts in a static long chat history) are insufficient for evaluating agent memory … that’s exactly the gap we're targeting!
𝙈𝙚𝙢𝙤𝙧𝙮𝘼𝙧𝙚𝙣𝙖 evaluates memory inside a real memory–agent–environment loop in:
🛒 Bundled web shopping,
🔎 Progressive web search,
🧳 Preference-constrained planning
🧠 Formal Math/Phys reasoning
Not just about retrieving the right sentence — it’s about using past experience to guide future actions across interdependent subtasks.
Paper: https://t.co/slZyxgmMAO
HF Dataset: https://t.co/fMkBCCcSbE
Agent memory benchmarks are misleading.
Scoring well on memory recall doesn't mean an agent can actually use that memory to take correct actions across sessions.
Models that achieve near-saturated performance on existing long-context memory benchmarks like LoCoMo perform poorly when tested in real agentic scenarios.
This new research introduces MemoryArena, a benchmark designed to evaluate agent memory across interdependent multi-session tasks.
Unlike existing benchmarks that test memorization separately from action or focus on single sessions, MemoryArena uses human-crafted agentic tasks where agents must learn from prior interactions and apply that knowledge to solve subsequent challenges.
Why it matters: as agents handle longer, multi-session workflows, memory isn't just about retrieval. It's about applying the right context at the right time to make good decisions.
Paper: https://t.co/PQpmsZVCvr
Learn to build effective AI agents in our academy: https://t.co/LRnpZN7L4c
Thrilled to share that three of our papers were accepted at #ICLR2026 🎉 (#RL)
They’re connected by one theme: learning systems shouldn’t need perfect supervision or hand-designed knobs—we can make them learn more efficiently and more automatically.
1) A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning
https://t.co/cLuOgIzovz
Multi-objective RL is often treated as “pick a scalarization and optimize it.” We propose a different lens: bring reward-free RL insights—especially forward–backward representations—to make MORL more efficient and reduce the brittleness of committing to one scalarized objective.
2) Composition-Grounded Instruction Synthesis for Visual Reasoning
https://t.co/jTVAJRfoA9
Process rewards are incredibly useful in RL, but VLM visual reasoning usually lacks them. We introduce a synthetic data recipe that generates compositional subproblems and associated process rewards, turning hard visual reasoning into learnable step-by-step supervision at scale.
3) BOAD: Discovering Hierarchical Software Engineering Agents via Bandit Optimization
https://t.co/q3P3ippj3q
Agent design is painful because agents are black-box and debugging becomes prompt trial-and-error. BOAD uses bandit optimization to automatically discover hierarchical agent structures and coordination, enabling smaller models + finetuning to achieve strong performance—often rivaling frontier approaches.
Massive thanks to my collaborators: @iris_j_xu@GuangtaoZ@charlesjin_@gan_chuang@ZexueHe@PCHsiehtw 🙏
If any of these directions resonate, I’d love to connect—happy to discuss details, share notes, or brainstorm follow-ups.
🚀 Single-agent coding assistants hit a wall on long-horizon tasks. Even frontier AI IDEs like Claude Code rely on manual sub-agent design—which is tedious and suboptimal.
What if we could discover the optimal agent hierarchy automatically?
Introducing BOAD: Bandit Optimization for Agent Design.
🔥 Outperforms GPT-4 & Claude-3.7-Sonnet (36B model) 🤖 Uses Multi-Armed Bandits for principled agent discovery
Stop guessing prompts. Start optimizing. 🧵👇
This work is led by @iris_j_xu during her internship at MIT-IBM Watson AI Lab.
Paper: https://t.co/U576P0979m
Code: https://t.co/R2FfcBoSKu
#LLM #RL #AIagents #CodingAgents #MultiArmedBandit
🎉The MemVis @ICCVConference workshop was a wonderful success!
👏Huge thanks to all our speakers and panelists for the inspiring talks and discussions: Kristen Grauman @ManlingLi_@albertobietti@RanjayKrishna@Ben_Hoov;
💪And thanks to all participants for the engaging discussions that made the day so lively!
🌺Finally, deep gratitude to our organizing team @maojiayuan , @kondic_jovana , @RogerioFeris@DimaKrotov for making it happen and making everything run smoothly.
💡Bridging vision and memory drives long-horizon, human-aligned intelligence. We look forward to continuing the conversation at the next MemVis @MemVis_ICCV25!
🚀 Our Memory and Vision (MemVis) Workshop is happening this Sunday 8:30am-1pm, Oct 19 at #ICCV2025 in Honolulu, Hawaii!
📍 Room #304B
🕘 Full schedule is live
Join us and our amazing speakers to explore how memory connects with vision models through inspiring talks, panels, and posters!
💫 Finally, huge thanks to our generous sponsor @abaka_ai, and join us in Happy Hour here: https://t.co/YSgagDD85V
🚀 Our Memory and Vision (MemVis) Workshop is happening this Sunday 8:30am-1pm, Oct 19 at #ICCV2025 in Honolulu, Hawaii!
📍 Room #304B
🕘 Full schedule is live
Join us and our amazing speakers to explore how memory connects with vision models through inspiring talks, panels, and posters!
💫 Finally, huge thanks to our generous sponsor @abaka_ai, and join us in Happy Hour here: https://t.co/YSgagDD85V
⏰ Just 2 days left to submit to #MemVis @ #ICCV2025!
We welcome works on state-space models, diffusion, retrieval, lifelong learning, multimodal memory, and more.
📝 Formats: ≤4p abstract or full ICCV paper.
Don’t miss it ▶️ https://t.co/ttDaNdd7qA
How to build a factual but creative system? It is a question surrounding memory and creativity in modern ML systems. My colleagues from @IBMResearch and @MITIBMLab are hosting the @MemVis_ICCV25 workshop at #ICCV2025, which explores the intersection between memory and generative models.
Link: https://t.co/ibbeGhRz4L
2/3 Also, many thanks to all authors who submitted their great work, gave oral talks & poster presentations!
Grateful to our dedicated reviewers — this workshop wouldn’t have been possible without you! 🙏
🔥 M+ is at #ICML2025 now!
We combine long-term memory (on CPU) with short-term memory (on GPU) for LLMs, pushing efficient long-context modeling to 160k+ tokens.
📍 My co-author @YzhuML will present M+ in person tomorrow 4:30pm.
👋 Come chat about scaling memory for LLMs!
🎉 Our paper “M+: Extending MemoryLLM with Scalable Long-Term Memory” is accepted to ICML 2025!
🔹 Co-trained retriever + latent memory
🔹 Retains info across 160k+ tokens
🔹 Much Lower GPU cost compared to backbone LLM
https://t.co/UZTSBINSij