@sudoingX Single 3090 kv cache on q8, serving to hermes via lm studio at 100k context window. 800 to 1000 tokens per second prompt processing and about 30tokens per second speed
What da hell💀
@stevibe your benchlocal tool helped me setup qwen 3.6 27b for hermes agent with a good context window.
Qwen 3. 27b at:-
Q4_k_m
100k context window
All layers offloaded to gpu
Kv cache both quant at q4_0
Hardware:-
An RTX 3090 with 32gig system memory.
@trackonindia this is what we receive in form of delivery after delays. The whole box was littered with mice waste and a few packages were eaten up by the mice. Even after complaint registration not a single person is listening responsibly. Now who will bear the loss of my order!
@ChibiReviews If it’s done by Apple themselves like their own servers and encryption then no problem. But if it’s on 3rd party then we are absolutely cooked
@venom1s Just get the latest possible thing, M4 will last you longer than m2 even though you may or may not notice the difference. The iPhone one is on you.