It’s an easter egg inside our latest Editions page, and was handcrafted by @letkma , @maca_graphics and myself. With support from the Editions team @shopify
You have to try it out:
https://t.co/jMJAQBUb4X
Goodbye VEO 2.
SkyworkAI just launched an OPEN SOURCE Human-Centric Video Foundation Model.
Why is no one talking about this?
Examples and link below 👇🏼
@virattt@AanandBajaj1@findatasets Love this mindset. Also, the tools are available for anyone RIGHT NOW and can at least do some of the work of the highly paid analyst. That's a win in my book
@ibamarief Pengalaman ngomong sama temen2 influencer, at least di Indonesia. Value followers itu x < tiktok < IG < Youtube. Jadi kalau punya large following di TikTok, ga mudah untuk monetize nya dibandingin IG or Youtube.
New short course Multimodal RAG: Chat with Videos, developed with @intel and taught by @vasudev_lal!
In this course, you’ll work with LLaVA (Large Language and Vision Assistant), a Large Vision Language Model (LVLM) that can process both images and text. For example, given an image of a person doing a handstand on a skateboard at the beach, LLaVA doesn't just caption the scene, it’s able to predict possible outcomes, like the person losing balance or falling off. By understanding not just what's in a video frame, but what might happen next, your application can provide more insightful answers to questions about video.
You'll build a full multimodal RAG pipeline that can chat about video content:
- Use the BridgeTower model to create joint text-image embeddings in a 512-dimensional multimodal semantic space.
- Learn video processing techniques to extract keyframes, generate transcripts using Whisper, and create captions.
- Use the LanceDB vector database to store and retrieve high-dimensional multimodal embeddings.
- Integrate the LLaVA model, combining CLIP's (Contrastive Language Image Pretraining) vision transformer with Llama, for advanced visual-textual reasoning.
Your final system will ingest video data, generate embeddings for frames and text, perform similarity searches for relevant content, and use the retrieved multimodal context to inform LVLM-based response generation. The result is a system capable of answering nuanced questions about video content, effectively chatting about the video it has processed.
Please sign up here! https://t.co/cjUHPK3rK2
@alexcarliera@reshotAI Is this a productized version of advanced liveportrait for comfyui? https://t.co/c4CrPvzRb8
Aside from better UI/UX, how is the tech used here better/different?
It's hilarious that Scarlett Johansson is pictured here, but there's no:
Sam Altman
Dario Amodei
Ilya Sutskever
Andrej Karpathy
or even...Mark Zuckerberg
@LeoSuppya @sgzsh269@venturetwins@Photoshop Even if IDM VTON is better, it specifically prohibits commercial use, while the one from Kolors can be used commercially (Apache 2 license).
@HalimAlrasihi@Kling_ai I think IDM VTON still better, especially when using Flux (though it’s heavier on the resources).
But good to have competition popping up like this.
@levelsio Not a coder but know some very basic python. Claude and ChatGPT is a lifesaver that enables me to create MVP real quick even as a non-tech founder. Obviously will need proper engineer to scale it up.
Meta presents Sapiens
Foundation for Human Vision Models
discuss: https://t.co/LH0tEgJvnX
We present Sapiens, a family of models for four fundamental human-centric vision tasks - 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction. Our models natively support 1K high-resolution inference and are extremely easy to adapt for individual tasks by simply fine-tuning models pretrained on over 300 million in-the-wild human images. We observe that, given the same computational budget, self-supervised pretraining on a curated dataset of human images significantly boosts the performance for a diverse set of human-centric tasks. The resulting models exhibit remarkable generalization to in-the-wild data, even when labeled data is scarce or entirely synthetic. Our simple model design also brings scalability - model performance across tasks improves as we scale the number of parameters from 0.3 to 2 billion. Sapiens consistently surpasses existing baselines across various human-centric benchmarks. We achieve significant improvements over the prior state-of-the-art on Humans-5K (pose) by 7.6 mAP, Humans-2K (part-seg) by 17.1 mIoU, Hi4D (depth) by 22.4% relative RMSE, and THuman2 (normal) by 53.5% relative angular error.
Actually I was reading the book "A Poison Like No Other: How Microplastics Corrupted Our Planet and Our Bodies" just last week.
I didn't realize the extent to which plastics have come to permeate and mess with our entire environment. It's not just about the polymer granules of the plastic, which is problematic by itself when during their breakdown they get small enough to make their way everywhere, including inside our organs, brains, etc.
It's about the ~thousands of exotic chemicals that get mixed into the plastics to tune them: plasticizers (to make them more flexible/durable), stabilizers (to help them resist heat, light), flame retardants, colorants, fillers, antioxidants, UV stabilizers, antistatic agents, lubricants, biocides, etc etc. These chemicals leach from the plastics over time (by default, but especially when you e.g. when you microwave your food). The vast majority of these chemicals have never been evaluated for safety.
There's many other fun facts in the book. We already knew "recycling" of plastic is basically fiction. It also turns out that e.g. when you see "biodegradable" on your plastic, that doesn't mean in normal natural conditions - they only degrade via specific processing plants that are equipped to degrade them.
Toxic, indestructible, synthetic molecules are mixing through the organic environments and the food chain and quite likely poisoning the environment and us.
It definitely feels like we've allowed the convenience of plastics to get way ahead of our understanding of their global effects and that there are some major unpriced externalities in the industry.
wow FLUX the image-generation model from @bfl_ml has taken the open-source ai world totally by storm
never seen so many derivative/spaces/demos of a model trending at the same time 🤯