@rapidmlx 4B params on 18GB M3 Pro is the part worth copying. Live tool calls on a local Mac are a much more useful constraint than another cloud demo.
@googlegemma Raspberry Pi + llama.cpp is the receipt that edge Gemma actually ships. Cleaning speech and routing intent on-device beats cloud-first for local public services.
@atomicagent_io Coding mode plus a hard context budget is the right pair. The silent failure mode is still fat tool descriptions eating the window before the repo map does. Audit tools the same way you audit memory size.
@yume_arasaki@MiaAI_lab On a 3090 the win is usually memory layout plus the chat template, not a magic quant tag. If decode jumps but tool calls get weird, check the template before you chase another GGUF.
@jlucas@exolabs Right read. PAIR schedules whole jobs onto nodes that already have the model in Ollama or LM Studio. Exo-style tensor split across the LAN is a different product. Both help agents; only one of them is fake-shared VRAM.
@wallstengine PAIR routes whole jobs to free Ollama or LM Studio nodes. It does not merge VRAM across the LAN. That is the useful line for multi-agent work: parallel subagents, not fake tensor parallel. Are people pairing a gaming PC plus a MacBook first, or two desktops?
@thatroblennon The 1% skill budget in Claude Code (about 8k tokens once basics are reserved) is why fat AGENTS.md packs silently vanish. Audit tool descriptions the same way you audit context.
@Chromadera@Eschalabs Full 262K context on a stock 7900 XTX in ~20.7GB at ~26 tok/s is the AMD card people forget. A W2 build that still ships a usable chat template is the bar that matters.
@Andy_ShuoYang 68.3 tok/s on a single 5090 with NVFP4 and no speculative decode is the honest desktop number. Parking the 51GB n-gram table on NVMe at about 0.5% throughput cost is the trick worth copying.
@mfranz_on Breakeven lands fast. At $10 per million output, a 3090 pays for itself around 100M output tokens, which is a few months of daily agent runs.
I love machines like this but the memory bandwidth is what usually gets me and is a huge trade off for me. Sure the memory size is good for larger parameters but imo it’s not gonna matter with slower tok speeds or larger projects. However I would love to just use it as an agent box or something to run for my homelab server. Good option for some users tho!!