OpenCompass just released RISEBench, the first benchmark on Reasoning-Informed Visual Editing (RISE).
GPT-4o Image Generation only scores 36% on this challenging task!
Technical Report: https://t.co/wYvyZdhjKk
#GPT4o
Uni-1 is a decoder-only autoregressive transformer. Text and images are represented in a single interleaved sequence, acting both as input and as output. This enables Uni-1 to think and render in the same forward pass, achieving a new benchmark of intelligence and quality.
Visual recognition (especially regarding abstract contents) remains a significant challenge for current VLMs. Evaluations based on VisFactor will drive improvements in models' capabilities in this area.
Welcome to the exam for your VLM on VisFactor! https://t.co/1oj6lASTri
Just created a Gallery to display all generation results on RISEBench (by powerful models including GPT-4o Image, Gemini-2.0, Bagel, etc.). Please contact me if you want the results of your new model to be included!
Tech Report: https://t.co/UfxfXFFDhq
OpenCompass just released RISEBench, the first benchmark on Reasoning-Informed Visual Editing (RISE).
GPT-4o Image Generation only scores 36% on this challenging task!
Technical Report: https://t.co/wYvyZdhjKk
#GPT4o
- VisualPRM for Test Time Scaling of Visual Reasoning Problems: https://t.co/dvHgGoubOG
- 5%~10% Avg. Accuracy Improvement over 7 mainstream benchmarks.
- This work is released with 400K Tuning Data & 3K Benchmark Problems
we are analyzing the top papers on @huggingface (~4000 papers mostly related to LLMs) and here is a list of the top 20 authors with the most papers published in less than 2 years.
all of them Asian!
(not necessarily in Asia)
this is no competition, these alphas OWN the game.
DeepSeek is a wake up call for America, but it doesn’t change the strategy:
- USA must out-innovate &race faster, as we have done in the entire history of AI
- Tighten export controls on chips so that we can maintain future leads
Every major breakthrough in AI has been American
After 1yr of Building
VLMEvalKit now reaches 100+ Contributors
On the journey of exploring LMM capabilities, we will go further
https://t.co/1oj6lASTri
OpenCompass has established a leaderboard to evaluate complex reasoning capability of LMMs, consisting of four advanced multi-modal math reasoning benchmarks. Currently, Gemini-2.0-Flash took the 1st place. DM me to suggest more benchmarks and models to this LB.
As my kids are singing APT non-stop these days, I did a bit of reverse engineering of the APT music video and tried to understand why the MV is so addictive.
Here is what I learned.
Mitigating racial bias from LLMs is a lot easier than removing it from humans!
Can’t believe this happened at the best AI conference @NeurIPSConf
We have ethical reviews for authors, but missed it for invited speakers? 😡