In this paper is presented BaseRT, a LLM inference runtime built directly on Apple’s Metal GPU API, without any intermediate framework.
https://t.co/B6NpmeE2Ca
We’ve built the fastest inference runtime for LLMs on Apple Silicon, faster than MLX and llama.cpp.
BaseRT is up to 35% faster than Apple's MLX on token generation (decode) for small models, and 78% faster on prefill for bigger models.
Details below.
Super cool to be partnered with @unity to deliver AI gameplay solutions in Unity Sentis.
AI driven NPCs in video games will help Creators build much bigger and more immersive worlds filled with 100s of characters. This is just the beginning..!
The AI fun continues ✨ You can use Unity Sentis in open beta to build AI-powered experiences in the Unity runtime. Sentis reduces the complexity of importing and running AI models and enables creating lifelike NPCs and new interactions for richer gameplay.
See it in the demo below that uses four chained AI models to create an open-ended dialogue and quest experience.
Learn more about Sentis: https://t.co/ZvR5qmOtR2
#Unite2023 #UnitySentis
It wasn't until I'd *taught* algorithms a few times that I finally understood why sorting is in the CS curriculum. Unfortunately, most curricula don't explain this!
It is NOT because sorting is an important algorithm to learn to implement...
We’re releasing GPT-4 — a large multimodal model (image & text in, text out) which is a significant advance in both capability and alignment.
Still limited in many ways, but passes many qualification benchmarks like the bar exam & AP Calculus: https://t.co/L6VGJ0WfFv