FP8 HippoAttention is coming: Up to 3X faster than FlashAttentionV2 on H100. More than 700T achieved FLOPS on MI-300X. #Transformers#GPU https://t.co/ZOlweFIKxJ
Writing prompts is tricky and tedius. We at @LeptonAI recently worked with our friends at @hippoml_com to release api for PromptLLM - better prompts for AIGC🖼️
Remember the song "down by the bay"? Read on to guess which image is generated by AI prompt :)
Our latest blog dives into the future serving system beyond Docker, showcasing our unified solution for data centers & local AI. https://t.co/2lDAx2AUro
Just one day after its release, a user has achieved what I've been envisioning for a while: every family should have a 'Brain Computer' where all AI requests can be processed at home, ensuring maximum privacy with no data sent to third parties. This user runs PrivateCanvas backend on an RTX 3090 and the PrivateCanvas UI on an iPad, creating with an Apple Pencil. It's a small step towards decentralized AI inference. :)
Introducing PrivateCanvas. Harness the power of your local GPU for contiguous editing and generating with cutting-edge models like Large Language Model, SDXL, Segment Anything, and GANs. Experience top-tier performance with minimal hardware demands. https://t.co/oo62q2LNcL
Hello world! Enable LLM on NVIDIA Datacenter & Orin DriveKit GPU and Apple GPU with SOTA Performance. On Apple M2 Max our new engine encode/decode is 13.8X/2.4X faster than llama.cpp #LLM#GPU#NVIDIA#Apple https://t.co/lmlxFXQNEo