.@AIatMeta's open-weight Muse Glimmer model is optimized for @Snapdragon powered devices, bringing industry-leading on-device agentic AI experiences and expanding what's possible for developers in an open AI ecosystem.
AI has a memory problem; moving data over long distances burns more energy and time, running into what we call the "memory wall".
Qualcomm High Bandwidth Compute (HBC) takes a different route: deliver near-memory computing to break the memory bottleneckAI has a memory problem; moving data over long distances burns more energy and time, running into what we call the "memory wall".
Qualcomm High Bandwidth Compute (HBC) takes a different route: deliver near-memory computing to break the memory bottleneckAI has a memory problem; moving data over long distances burns more energy and time, running into what we call the "memory wall".
Qualcomm High Bandwidth Compute (HBC) takes a different route: deliver near-memory computing to break the memory bottleneckAI has a memory problem; moving data over long distances burns more energy and time, running into what we call the "memory wall".
Qualcomm High Bandwidth Compute (HBC) takes a different route: deliver near-memory computing to break the memory bottleneck.
This Week in AI:
🔵 Qualcomm High Bandwidth Compute (HBC) moves compute closer to data, improving AI inference performance, efficiency, and total cost of ownership: https://t.co/RgXMyCbN4c
🔵 We expanded our collaboration with @BMWGroup to help power the next generation of AI-driven mobility experiences through @Snapdragon Digital Chassis solutions: https://t.co/k6lJ1VYO1b
🔵 We completed our acquisition of @Modular, strengthening our AI software capabilities and accelerating generative and agentic AI across different environments: https://t.co/jBTU80RnTI
This Week in AI:
🔵 We launched GenieX, a unified runtime that simplifies running modern generative AI models on-device across Windows, Android, and Linux: https://t.co/um4EVeDVFO
🔵 @durga_malladi, EVP & GM of Technology Planning, Edge Solutions, and Data Center at Qualcomm, sits down with @Newsweek to discuss the rise of agentic AI and how hybrid AI will shape the future of computing: https://t.co/X0WyjGEoM9
🔵 Qualcomm High Bandwidth Compute advances faster, more efficient AI scaling: https://t.co/5eYcbX8p5s
🔵 Qualcomm AI Hub enables smarter model optimization with quantization: https://t.co/qRUkkZzWeG
Modular has joined @Qualcomm, and it's a big moment for MAX, Mojo, and our developer community.
Join us on Monday, 8/10 for a live Ask Us Anything, featuring:
@clattner_llvm, Modular cofounder & CEO, now EVP of Advanced AI Software & Platforms at Qualcomm
@iamtimdavis, Modular cofounder & president, now SVP & GM of Modular at Qualcomm
@durga_malladi, EVP and GM, Technology Planning, Edge Solutions & Data Center at Qualcomm
Submit questions in advance: https://t.co/y3vxoUlcZS
RSVP to the livestream: https://t.co/zT5XitimoE
🚀Summer Fest Day 3: Cost-Effective MoE Inference on CPU from Intel PyTorch team
Deploying 671B DeepSeek R1 with zero GPUs? SGLang now supports high-performance CPU-only inference on Intel Xeon 6—enabling billion-scale MoE models like DeepSeek to run on commodity CPU servers.
Key highlights:
1. Full CPU backend for SGLang with Intel AMX
2. Native BF16 / INT8 / FP8 support for both Dense and Sparse FFNs
3. 6–14× TTFT and 2–4× TPOT speedup vs. llama.cpp
4. 85%+ memory bandwidth efficiency with optimized MoE kernels
5. Flash Attention V2 + MLA + MoE all optimized for CPU
6. Multi-NUMA parallelism mapped from GPU-style Tensor Parallelism
This work is now fully upstreamed to SGLang main—read how we achieved it, and how far you can go without a GPU 👇
#LLMInfra #ModelServing #MoE #Xeon6 #SGLang #FP8 #INT8 #CPUInference
💡Intel Neural Compressor v3.4 is released, supporting more quantization recipes, e.g., W4A8 (FP8). In the past weeks, we've contributed the algorithm AutoRound to HF Transformers and vLLM, and now we are making the contribution to SGLang. Stay tuned.😀
🎯https://t.co/XklzQFSqo1
Ever tried searching for “more ladybugs than flowers”?
We did. And the AI nailed it.
Fine-tuned LLMs can really deliver when trained on the right datasets. This demo shows what happens when #Qwen3 models are optimized and deployed on Intel hardware.
Read the full article: https://t.co/PcAHJu7kwu
Automated prompt engineering on-device—no fine-tuning, no RAG.
This new guide shows how to use #DSPy with Intel #oneAPI and llama.cpp to boost task accuracy from 📉 35% → 📈 78%
Run LLMs locally, optimize efficiently.
Read the guide → https://t.co/NgkfQn9kMw
Learn how #PyTorch 2.7 and @Intel GPUs can accelerate your #AI workloads on Windows and Linux.💡 Read the latest blog from the Intel PyTorch team: https://t.co/JOiHw2Xa7M
Dive into Automated Prompt Engineering! Learn how to create an efficient pipeline tailored for specific tasks and discover techniques to optimize your prompts on Intel AI PC.
#AI#GenerativeAI#LLMs#DeepLearning#PromptEngineering#weareintel
https://t.co/8tRTVkJbve
Its little brother is already up and working. Will take work to use XMX and make really fast, but already everything is correct and usable.
PyTorch -> tinygrad -> OpenCL -> i915
Create job recommendations tailored to the user with Intel® + @TensorFlow ! Personalized, efficient, and faster results. Let’s build smarter systems today. Learn how: https://t.co/6C1iU0QEmR
#AI#TensorFlow#JobRecommendation#MachineLearning
Unlock faster ML model training with JAX! Get started with high-performance computing using a simple NumPy interface. Learn how: https://t.co/uQ3A2LIrIh
#MachineLearning#AI#DeepLearning#JAX
Curious about the latest optimizations in #PyTorch 2.6 for @intel Corporation platforms? The Intel PyTorch Team has detailed key improvements for Intel x86 CPUs and GPUs, enhancing performance and efficiency.
👉 Explore the latest advancements and how they can improve your PyTorch workloads: https://t.co/UkedEDaWl1
#pytorch #machinelearning #cpus #gpus
Dive into language identification with #PyTorch and Intel. Follow along and learn to develop a solution using the Hugging Face SpeechBrain toolkit optimized for Intel hardware to identify up to 133 languages. https://t.co/x745jL9vvg