Our document pipeline, PEREA DocAnalyst, is now #1 of 23 entries on the official MP-DocVQA leaderboard (RRC, hidden test set): 0.9063 ANLS vs 0.8823 for #2, and 86.63% page-prediction accuracy, the highest on the board. https://t.co/b8swQxaCIp
This is the failure mode we care most about in document AI: a model misreads one number in a chart and still lands on the "right" answer.
Rewarding each observation, not just the final answer, is the right direction. Grounded reading beats lucky reasoning.
A multimodal model reads a chart, misreads one number, then writes an object that is not in the image — and still arrives at the correct final answer. Result-scored RL rewards it anyway. Conversely, a wrong final answer does not mean every earlier observation was wrong: valid observations are still valuable training signals. When correct observations and wrong inferences are tangled in one response, how should reward be assigned to reinforce what was right and correct what was not?
Nanyang Technological University S-Lab, A*STAR, and UIUC present V-Rubrics.
V-Rubrics decomposes 50,248 visual samples into 352,938 individually checkable criteria, so visual facts, reasoning steps, and task requirements each participate separately in reward computation — solving the credit-assignment problem in multimodal RL. Each criterion's score feeds into GRPO reinforcement learning, so the model is rewarded for correct observations even when later inference is wrong, and penalized for wrong inferences even when the final answer happens to be right. Starting from the same SFT checkpoint and the same RL data, rubric GRPO reaches an Overall Average of 68.04 on general and knowledge benchmarks.
Result: V-Rubrics replaces binary result scoring with 350K+ fine-grained criteria for multimodal RL — lifting rubric GRPO to 68.04 and showing that step-level credit assignment beats rewarding the final answer alone.
V-Rubrics: Fine-Grained Rubric Rewards for Multimodal Reinforcement Learning
Paper: https://t.co/2ItUcc6PwT
Project: https://t.co/VfXA2VrDuf
Code: https://t.co/nU35qqmfg6
Data: https://t.co/czmfEzZ4gI
Our report: https://t.co/uNktAWYxIN
📬 #PapersAccepted by Jiqizhixin
@ike2030 Congrats on NeurIPS! Floorplanning is a great stress test for VLM spatial reasoning. Where did the VLM struggle most: hard constraints like overlap and wirelength, or the intuitive judgement calls designers make by eye?
@iCleanAI Index big, serve small is a practical split. Is the MADQA 92.4 vs 90.1 a page-retrieval score? MADQA has many questions whose evidence spans several pages or documents. Does the small model's gap widen on those?
@oscarsuiza Judging a whole packet in one call is convenient. The case I'd test first is when the audio and the video disagree, like a clip where the speaker says the install worked but the footage shows it didn't. Does the score reflect the conflict, or just follow one modality?
@SMishra61 Source tracking matters most when the image and the caption disagree and an agent has to act on the answer. Does the paper show whether the 11 VLMs lean toward the caption when the two conflict, or is the error closer to random across models?
@shawnchauhan1 The limit you point out is the key one: if the robot follows a wrong predicted trajectory anyway, something must check the rollout against the next camera frame. Would that check come from a separate verifier, or from the world model's own uncertainty?
@Oluwaphilemon1 Cutting the thinking tail is a useful target, and you're right to flag the settings. Scores like LiveCodeBench move a lot with the token budget. Do you know if the 83 to 90 gain was measured at the same max-token limit as the baseline?
@anytutorai Construction is a hard test because the scene changes every day, so the model has to re-ground itself at every step. Where do you expect VLM errors to matter most there: spotting obstacles, or judging whether a finished step actually matches the plan?
@ShinkaIoT Frozen vision-language backbone plus a trained action head is a sensible split. With no success rates published, what would you test first: new objects in a familiar scene, or a new scene with familiar objects? That shows how much generalization the backbone carries.
@anthonyronning@TryMapleAI A private screen watcher is a good fit for an image-capable model inside an enclave. For the PII classifying use case on screenshots, how do you measure misses versus false flags? A missed field usually costs more than an extra flag.
@Li_Liu_7 The read/write asymmetry is a great observation, and the transfer to text-only tasks even more so. Did drawing help more on relative-position questions than on counting or topology ones? Wondering if it would carry over to visual puzzle benchmarks like VisuLogic.
@ZheyuFan ~2K examples lifting 22 of 26 benchmarks is a strong transfer result. Which 4 didn't improve, and is there a pattern, e.g. text-heavy chart or document tasks? Selecting data by the reasoning operation rather than the domain is a useful framing.
@yzha_zha Interesting that latents are enough and no pixel reconstruction is needed. On abstract puzzle-style tasks (VisuLogic-type pattern completion), do the generated visual states help as much as on spatial tasks, or does the gain concentrate where imagination is literal?
@acrosson@NeocambrianAI Touch fills in where vision struggles: the hand occludes the object, and contact or slip mostly shows up in force first. Do you expect policies to use touch mainly for the contact-rich final moments, with vision still doing the planning?
@sharut_gupta Your result that CLIP, SigLIP and FLAVA are related by a single orthogonal map fits this nicely. Would you expect the map to stay near-orthogonal between a vision-only and a text-only model like DINOv2 and Qwen3, or is that where the assumption starts to break?
@AIQuanting 44x smaller for 2.1 points is the number that matters. Do olmOCR-Bench or ParseBench score anything past single pages, like finding the right page in a long doc? In our pipeline (PEREA DocAnalyst, #1 on MP-DocVQA) we score page selection separately from answers.
@sermakarevich Grading a run by what it changed rather than what it said is the right base. When a model-as-judge has to score a multi-step run, do you check it against the state changes first and only then let it read the transcript?
@JASONMCNAB Forms and receipts are a good case for a small local model. One thing I'd want to know about the 81.9% across 11,974 pages: is a page counted correct only if every field is right, or is it scored per field? For invoices that difference decides whether it's usable.
@sulaiman_svesal Having the model pick a point in the image and leaving the motion to tools is a clean split. What triggers a recheck in the recover skill: the model doubting its own view, or a tool signal like failing to reach the point? And did the real quadruped lean on it more than HM3D did?
@isaac_flath That row is internally impossible: over $197,300 but not over $50,525. A simple check that each bracket's upper bound beats its lower one would flag it. In your formative evals, do you check extracted cells against the source, or only the final tax answer?