🎉 mlx-omni-server v0.4.0 is out!
✨ New /v1/embeddings service (via mlx-embeddings, thx @0ssamaak0) - generate embeddings easily!
🔊 More TTS models (kokoro, bark & more!) with mlx-audio (thx @ zboyles) !
🧠 mlx-lm upgraded for models like qwen3!
https://t.co/iGK9vfbRuK
I created an open source mac app that mocks the usage of OpenAI API by routing the messages to chatgpt desktop app so it can be used without API key
You can simply change the api base (like if u are using ollama) and select any of the models that u can access from chatgpt app
I have been exploring @cohere Multimodal Embed 3 last week. It's like magic! 🪄
I was able to replace the whole pipeline of CLIPPyxX (CLIP + OCR + Text embedding model) with just Embed 3.
fp
I created the index once and now I'm can do image search very easily. Before I had to create 2 indices (1 for image embeddings and 1 for text embeddings) and search using 2 routes.
The results are amazing You can try it now using Cohere-embed branch https://t.co/hJPgxb6e5H
I integrated locally running mistral with Siri, far better than standard Siri. no more "Here's what I found"
Thanks to @ollama for making this possible!
GitHub: https://t.co/ILydRonUGT
I integrated locally running mistral with Siri, far better than standard Siri. no more "Here's what I found"
Thanks to @ollama for making this possible!
GitHub: https://t.co/ILydRonUGT
Once CLIPPyX server is running, you can access it from any UI. I made a simple html page and Flow launcher plugin. and @Microsoft Powertoys run plugin will be ready soon!
CLIPPyX: The new AI powered image search tool
🔵Search by Image Caption
🟡Search by Textual Content in Images
🔵Search by Image Similarity
Video at 1x speed on my laptop (1660 Ti / 16 GB RAM)
GitHub: https://t.co/hJPgxb6e5H
Tool Overview
Models can be used directly from🤗Transformers, or more efficient options like @Apple Mobile CLIP and GGUF text embedding models through llama.cpp
I integrated locally running mistral with Siri, far better than standard Siri. no more "Here's what I found"
Thanks to @ollama for making this possible!
GitHub: https://t.co/ILydRonUGT
@Mo7amed_Ma3rouf@MatthewBerman@GroqInc@pinecone I think it's not just about calling the LLM itself. there're other things like memory, input image, and adding features like RAG. requires a python app.
On-device model is a very good idea (I think apple is going to do it next June) however now, model options are limited.....
@MatthewBerman@GroqInc@pinecone I have already tried similar approach (no RAG yet) but it supports multimodality too
Check it out: https://t.co/ILydRonUGT
@rezk1222 في حاجات في أساسيات الـCS معمولة بشكل كويس جدا، المشكلة إنه بدأ يحاول يعمل دور documentation بديل لـlibraries أو frameworks، ودا شيء شديد الغباء بصراحة