π Excited to unveil our latest work: WaLa! A simple yet powerful approach. We're also open-sourcing billion-parameter models trained on 10M 3D shapes across various conditioning modalities! Can't wait to see what the community creates! π
π https://t.co/BvwuzkEnL5
#AI #MachineLearning #OpenSource #3DGeneration
π Excited to unveil our latest work: WaLa! A simple yet powerful approach. We're also open-sourcing billion-parameter models trained on 10M 3D shapes across various conditioning modalities! Can't wait to see what the community creates! π
π https://t.co/BvwuzkEnL5
#AI #MachineLearning #OpenSource #3DGeneration
Heat map-based ML explainability can be complex. Check out our work explaining learned visual features with pre-trained language models, using a translator network to convert visuals to text and uncover insights from data/model. #ML#AI#Explainability
https://t.co/vPSdb25I80
SLiMe: Segment Like Me
paper page: https://t.co/SyR9JBKMZY
Significant strides have been made using large vision-language models, like Stable Diffusion (SD), for a variety of downstream tasks, including image editing, image correspondence, and 3D shape generation. Inspired by these advancements, we explore leveraging these extensive vision-language models for segmenting images at any desired granularity using as few as one annotated sample by proposing SLiMe. SLiMe frames this problem as an optimization task. Specifically, given a single training image and its segmentation mask, we first extract attention maps, including our novel "weighted accumulated self-attention map" from the SD prior. Then, using the extracted attention maps, the text embeddings of Stable Diffusion are optimized such that, each of them, learn about a single segmented region from the training image. These learned embeddings then highlight the segmented region in the attention maps, which in turn can then be used to derive the segmentation map. This enables SLiMe to segment any real-world image during inference with the granularity of the segmented region in the training image, using just one example. Moreover, leveraging additional training data when available, i.e. few-shot, improves the performance of SLiMe. We carried out a knowledge-rich set of experiments examining various design factors and showed that SLiMe outperforms other existing one-shot and few-shot segmentation methods.
SLiMe: Segment Like Me
paper page: https://t.co/SyR9JBKMZY
Significant strides have been made using large vision-language models, like Stable Diffusion (SD), for a variety of downstream tasks, including image editing, image correspondence, and 3D shape generation. Inspired by these advancements, we explore leveraging these extensive vision-language models for segmenting images at any desired granularity using as few as one annotated sample by proposing SLiMe. SLiMe frames this problem as an optimization task. Specifically, given a single training image and its segmentation mask, we first extract attention maps, including our novel "weighted accumulated self-attention map" from the SD prior. Then, using the extracted attention maps, the text embeddings of Stable Diffusion are optimized such that, each of them, learn about a single segmented region from the training image. These learned embeddings then highlight the segmented region in the attention maps, which in turn can then be used to derive the segmentation map. This enables SLiMe to segment any real-world image during inference with the granularity of the segmented region in the training image, using just one example. Moreover, leveraging additional training data when available, i.e. few-shot, improves the performance of SLiMe. We carried out a knowledge-rich set of experiments examining various design factors and showed that SLiMe outperforms other existing one-shot and few-shot segmentation methods.
Semantic segmentation with an arbitrary granularity is a challenging task. We introduce SLiMe: Segment Like Me, which can segment an image according to a given sample by leveraging diffusion models' cross/self-attention and prompt optimization. https://t.co/d5THFQHdLK
Diffusion for virtual worlds: we've trained a new model to create 3D objects from text. And it's 50x faster than any alternative.
We made each of these cute characters in just one minute with Genmo's new text-to-3D model.
Build your own world at https://t.co/O1AFhhf0WC