On RefCOCO, EGM-Qwen3-VL-8B reaches 91.4 IoU with 737ms latency, outperforming Qwen3-VL-235B at 90.5 IoU and 4,320ms, while running 5.9x faster.
EGM points toward deployable, efficient visual grounding models that reason before they localize.
Excited to share that our paper EGM: Efficient Visual Grounding Language Models has been accepted to ECCV 2026!
arXiv: https://t.co/AVO8vAs06J
GitHub: https://t.co/jLBlV4odZy
Project: https://t.co/q5f5vd2jQU
#ECCV2026#VLM#VisualGrounding#EfficientAI
EGM equips small VLMs with test-time reasoning. A small VLM generates richer intermediate reasoning before predicting coordinates, improving both standard and amodal grounding while keeping latency low.
What if visual agents could do more than look at an image?
PERIA introduces a tool-augmented VLM workflow for spatial reasoning: perceive task-relevant evidence, interact with the visual context, and reason over accumulated observations.
#VLM#AI#SpatialReasoning#Agent
The takeaway is simple: tool access is not enough. Trained tool use turns visual evidence into spatial intelligence.
From map reasoning to visual probing, PERIA shows how visual agents can actively gather evidence before they answer.
The result: stronger grounded spatial reasoning from a compact model. Across 13 benchmarks, PERIA-8B improves over its Qwen3-8B backbone by 10.0% in-distribution and 4.4% out-of-distribution, while reaching performance comparable to much larger models——Qwen3-235B and GPT-5.