Apparently spatial intelligence is worth ๐.๐๐ now. We settled for a ๐๐๐ฎ๐ซ๐๐๐ acceptance ๐
Excited that ๐๐๐๐๐๐๐๐๐ was accepted to the ๐๐๐ฎ๐ซ๐๐๐ ๐๐๐๐ E&D Track ๐
We study how VLMs perceive, reason, and act in interactive 3D worlds ๐. We find that they often reason well once spatial structure is given explicitly, but still struggle to extract that structure from visual input ๐. We also find that the interaction setup itself can significantly change the outcome โ๏ธ.
Big thanks to @YanZhengtexas & Atlas Wang ๐
More soon!
#NeurIPS2026 #SpatialIntelligence #VLMs
๐Our paper "Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety" has been accepted to EMNLP 2025 Main Track! @emnlpmeeting
๐First survey connecting LLM interpretation & safety
#IJCAI2025 Matthew Hull, Georgia Tech, delivering their talk #8557: RenderBender: A Survey on Adversarial Attacks Using Differentiable Rendering at the CV: attacks session.
๐จ New work: We rethink how we finetune safer LLMs โ not by filtering after the generation, but by tracking safety risk token by token during training.
We repurpose guardrail models like ๐ก๏ธ Llama Guard and Granite Guardian to score evolving risk across each response ๐ โ giving rise to the STAR โญ score, a fine-grained safety signal that enables more targeted safety supervision.
On top of this, we introduce โญDSS (STAR-Guided Dynamic Safety Shaping) โ a training method that ๐ซ suppresses unsafe patterns, ๐ช preserves capability, and generalizes across LLMs, guardrails, harm levels, and datasets.
Our method outperforms "Deep Token," the method from this yearโs #iclr2025 Best Paper ๐ โ remaining robust against key finetuning-as-a-service threats like ๐ response adaptation, ๐งช prompt poisoning, and ๐ harmful prefilling.
#MachineLearning #DeepLearning #LLM #AISafety #Alignment #Finetuning
@ChatGPTapp has a new sub-window that seems intended to display reasoning chain-of-thought but the display truncates and is blanked after a few output words on MacOS
Highly niche problems - @ChatGPTapp app always wants to emphasize parts of an LaTeX equation using \underbrace but it can't render with display math so it just displays as raw chars. When I ask it to re-generate w/out \underbrace, it works great!
Diffusion models leverage a variety of samplers.
Deterministic methods like DDIM produce orderly paths. In contrast, stochastic samplers like DDPM produce chaotic trajectories.
Despite their differences, both methods draw valid samples from the underlying distribution.
Create heatmaps that localize text concepts in generated videos.
We discovered that our approach, ConceptAttention, can be directly extended from image generation to video generation models!
It's amazing how simple techniques often generalize way better than more complex ones.
Diffusion Transformers aren't just generative models, but also powerful multi-modal encoders.
ConceptAttention creates rich heatmaps of text concepts in images from DiT representations.
This even works on real images, and can be applied to tasks like segmentation!
Demo ๐
Transformer Explainer
Interactive Learning of Text-Generative Models
discuss: https://t.co/p1GMsyhsEI
Transformers have revolutionized machine learning, yet their inner workings remain opaque to many. We present Transformer Explainer, an interactive visualization tool designed for non-experts to learn about Transformers through the GPT-2 model. Our tool helps users understand complex Transformer concepts by integrating a model overview and enabling smooth transitions across abstraction levels of mathematical operations and model structures. It runs a live GPT-2 instance locally in the user's browser, empowering users to experiment with their own input and observe in real-time how the internal components and parameters of the Transformer work together to predict the next tokens. Our tool requires no installation or special hardware, broadening the public's education access to modern generative AI techniques.
Introducing ClickDiffusion!
We developed a system for precise image manipulation and generation that combines natural language instructions with visual feedback provided by the user through a direct manipulation interface.
@DJSnM@neiltyson I can't believe how many people taught _only_ the equal transit theory (still widely accepted!). I was also taught this way when I started flying and have only recently realized that there is more to it.
Looking to convert your tables from images and PDFs to HTML? Check out our UniTable! It outperforms both GPT-4V and LLaVA-1.6 in table recognition. #TableRecognition
Paper and code in thread ๐งต