๐ข๐ฝ๐ฒ๐ป-๐๐ผ๐๐ฟ๐ฐ๐ถ๐ป๐ด ๐ฆ๐ฒ๐ป๐๐ฒ๐ก๐ผ๐๐ฎ-๐ฉ๐ถ๐๐ถ๐ผ๐ป-7๐-๐ ๐ผ๐ง โ ๐ผ๐ป๐ฒ ๐บ๐ผ๐ฑ๐ฒ๐น, ๐ฎ๐น๐น ๐บ๐ฎ๐ท๐ผ๐ฟ ๐๐ถ๐๐ถ๐ผ๐ป ๐๐ฎ๐๐ธ๐.
Reimagines CV as multimodal generation. Steerable by language or visual prompts. ๐ก๐ผ ๐๐ฎ๐๐ธ-๐๐ฝ๐ฒ๐ฐ๐ถ๐ณ๐ถ๐ฐ ๐ต๐ฒ๐ฎ๐ฑ๐ โ ๐๐๐ถ๐น๐น ๐ฆ๐ข๐ง๐, ๐๐๐ฟ๐ฝ๐ฎ๐๐๐ถ๐ป๐ด ๐๐ผ๐ผ๐ด๐น๐ฒ ๐๐ฒ๐ฒ๐ฝ๐ ๐ถ๐ป๐ฑ'๐ ๐ฉ๐ถ๐๐ถ๐ผ๐ป ๐๐ฎ๐ป๐ฎ๐ป๐ฎ ๐ผ๐ป ๐๐ฒ๐ด๐บ๐ฒ๐ป๐๐ฎ๐๐ถ๐ผ๐ป ๐ฎ๐ป๐ฑ ๐ฑ๐ฒ๐ป๐๐ฒ ๐ด๐ฒ๐ผ๐บ๐ฒ๐๐ฟ๐.
๐ ๐๐. ๐ฉ๐ถ๐๐ถ๐ผ๐ป ๐๐ฎ๐ป๐ฎ๐ป๐ฎ (๐๐ผ๐ผ๐ด๐น๐ฒ ๐๐ฒ๐ฒ๐ฝ๐ ๐ถ๐ป๐ฑ):
๐นReferring segmentation (RefCOCOg cIoU): 80.3 vs 73.8
๐นSemantic segmentation (Cityscapes mIoU): 71.2 vs 69.9
๐นDepth estimation (NYUv2 ฮด1): 98.1 vs 94.8
๐นSurface normal (NYUv2 mean error): 14.4 vs 17.8
Built on a decade of SenseTime's CV leadership โ #1 in China's Vision AI market for 10 consecutive years, and named a Tech Innovator in global GenAI CV by Gartner. ๐ฉ๐ถ๐๐ถ๐ผ๐ป, ๐ป๐ผ๐ ๐ป๐ฎ๐๐ถ๐๐ฒ ๐๐ผ ๐ณ๐ผ๐๐ป๐ฑ๐ฎ๐๐ถ๐ผ๐ป ๐บ๐ผ๐ฑ๐ฒ๐น๐. New tasks, defined by instruction โ not by new model heads.
๐ฆ๐ฒ๐ป๐๐ฒ๐ก๐ผ๐๐ฎ-๐ฉ๐ถ๐๐ถ๐ผ๐ป-7๐-๐ ๐ผ๐ง, ๐ณ๐๐น๐น๐ ๐ผ๐ฝ๐ฒ๐ป-๐๐ผ๐๐ฟ๐ฐ๐ฒ๐ฑ: ๐ผ๐ป๐ฒ ๐บ๐ผ๐ฑ๐ฒ๐น, ๐ฒ๐๐ฒ๐ฟ๐ ๐บ๐ฎ๐ท๐ผ๐ฟ ๐๐ถ๐๐ถ๐ผ๐ป ๐๐ฎ๐๐ธ ๐ฏ๐ฒ๐น๐ผ๐:
๐น๐๐ฒ๐๐ฒ๐ฐ๐๐ถ๐ผ๐ป/๐ข๐๐ฅ/๐๐จ๐
๐น๐๐ฒ๐ฝ๐๐ต & ๐ป๐ผ๐ฟ๐บ๐ฎ๐น
๐น๐ฆ๐ฒ๐ด๐บ๐ฒ๐ป๐๐ฎ๐๐ถ๐ผ๐ป
๐น๐ ๐๐น๐๐ถ-๐๐ถ๐ฒ๐
It can also define new vision-task variants through natural language โ recombining visual capabilities across traditional task boundaries.
Open-sourced: model weights + the SenseNova-Vision Corpus (50M-example subset, plus full toolkit to reproduce the remaining public-source data) for research & development
๐ค HF: https://t.co/dJRnmJ9tPJ
๐งGitHub: https://t.co/CRTDeAigXn
๐ป Try the demo: https://t.co/KDSf7xtjnj
๐Technical Report: https://t.co/82wolPMLMR
๐พDiscord: https://t.co/NdqTQVFANp
๐ฅVision as Unified Multimodal Generation๐ฅ
๐ฏSenseNova-Vision๐ฏ unifies vision tasks (e.g., detection, keypoints, segmentation, depth, surface normals, point maps, and camera pose) as unified multimodal model (UMM) generation *with SOTA results*
- Code: https://t.co/a1WEmDLVe4
SenseNova-Vision-7B-MoT is live on ModelScope: a unified multimodal generation model for computer vision. ๐
๐ค https://t.co/25qQEJ2Mvn
๐ https://t.co/zH10Vcdvcf
๐ Strong against generalist vision models: compared with Youtu-VL, it reports higher COCO detection mAP, Cityscapes mIoU, RefCOCO / RefCOCOg grounding, and NYUv2 depth ฮด1.
๐งฉ One model, many output types: structured text for boxes, points, OCR, GUI grounding, keypoints, and camera records; image-like outputs for masks, depth, normals, and point maps.
๐ง No task-specific heads: visual tasks are expressed through text, image, or mixed text-image generation, with outputs decoded back into benchmark formats.
๐ Trained on SenseNova-Vision-Corpus-50M, covering structured visual understanding, dense geometry, segmentation, and multi-view visual geometry.
License: CC BY-NC 4.0.
Thanks for reposting, Songyou! Vision Banana is great and really inspiring! Happy to take this direction one step further by bringing vision closer to foundation models. Looking forward to more discussions!
SenseNova-Vision ๐ฅ SenseTime's new model treats all of computer vision as generation
- 7B
- CC BY-NC 4.0 ( non commercial )
- Model/ 50M instruction corpus/ benchmark/ paper/ demo, all open ๐ซถ