this is real but pretty misleading. airllm has been around since late 2023, not breaking news. the tweet conveniently skips the biggest problem, speed. it loads every layer from disk per token, so a 70b model at fp16 means reading 140gb from disk for each token. even with a good nvme thats like 20 seconds per token. your old gaming laptop technically runs it but you wait 30 minutes for a short answer. if you have a good gpu you dont need this, llama.cpp with quantized models is way faster. if you dont have a good gpu youre stuck waiting forever. cool proof of concept but the bottleneck is disk io not gpu, and no hardware upgrade fixes that
openclaw icin 60 saniyede kurulum yapan bir deployment tool u olusturdum gectigimiz haftalarda, benim icin hafta sonu projesiydi ve ardindan cok fazla alternatifi ciktigi icin pazarlamaya pek ugrasmadim. Free Trial var istersen deneyebilirsin. @clawstackapp
Su siralar runpoddan kiraladigim GPU uzerinde local modeller deploylayarak denemeler yapiyorum fakat hala fiyat/performans bir yontem bulamadim. Kucuk modellerin tool calling becerileri kotu, guzel olanlarinda VRAM ihtiyaci oldukca yuksek. Ondan oturu pek memnun kalmadim.
Dun MiniMax M2,5 un bir paketini gordum, Aylik 50$ gibi ilginc bir fiyati var ve 1000 prompt/4 saat seklinde bir paketi mevcut. Redditte oldukca ovuldugu yazilari gordum ve yeterli gibi gozukuyor. Belki onuda inceleyebilirsin.
https://t.co/UnjzoNB9J7
@iamfra5er Maybe u should try other models which has higher tool usage talents. I am currently using is Claude Haiku 3.5 daily. Its doing most of tasks just fine.
https://t.co/IIW65LHNq6
Hackathons aren’t just about code, they’re about people who help you push further. ✨
At ETHIstanbul, 30+ mentors are joining to support builders, share insights, and make sure no idea gets stuck.
Meet the ETHIstanbul mentors 👇
we can build your next million dollar idea!
yeah, you heard it right! just schedule a meeting with us, share your idea and we will bring your vision to life together.