@WescheNex1q That's also one of the reasons why models like Luna and Gemini Flash are so popular among enterprise customers, especially in domains other than just coding. They don't have to be "the best". They just need to be cost-efficient and "good enough".
@WescheNex1q No need to pretend everything requires SOTA intelligence models. When you do certain things on scale, like extraction/classification/search and process billions of tokens at a given time, small models do those things just fine, but way faster and cheaper than K3/Sol/Fable.
@MiaAI_lab@theo Yep, especially if cost is a part of the equation. Not every task needs a 2.8T model that costs $3/$15 (in/out), if a cheaper one at $0.2/$1.2 is enough. When you need to crunch a lot of data and don't necessarily need peak intelligence, Luna and DSV4F 0731 are hard to beat.
@elshayib_ Individual accounts churn like crazy on both sides, so it might not be the best metric to use for such predictions. Would be a totally different story if those were primarily team/enterprise accounts tbh, but companies are way slower with moving between vendors.
@rohanpaul_ai Doesn't matter BF16 weighs 55.6 GB, that person runs Q4_K_M that's like 17.8 GB instead. NVFP4 would be a bit faster than that, despite it being closer to 20 GB.
@TheAhmadOsman "That means Tech Debt is a thing of the past as well" - strongly disagree. The definition of tech debt is gonna evolve over time. Every time a noticeably more capable model drops, most (if not all) things made before that date will be considered tech debt.
@amkza007@TeksEdge (I agree with the general sentiment though - it'd be way more viable if the bandwidth was closer to the 70-class GPUs or above, or if it used HBM for at least a part of its memory)
@amkza007@TeksEdge Gorgon Halo is slightly faster than its predecessor - not by much though (~6.7%). It uses 8533 MT/s RAM instead of 8000 MT/s (Strix Halo) - so the bandwidth ends up being exactly the same as on DGX Spark, aka 273 GB/s.
@alexocheema all the other filters have an option to clear them, but not the slider. Also, the readability with many overlapping data points is kinda weak. Sure, tooltips help a bit, but maybe it can be changed to make the line containing focused data point stand out a bit more?
@alexocheema More hardware: Arc Pro B60/B70 and R9700 for sure. Maybe even RTX5060Ti 16 GB (or two) and RX9060XT (or two) to showcase more entry level options, as used hardware or workstation cards aren't always an option.
@alexocheema as for the hardware, maybe also Strix Halo with 64 GB and 128 GB RAM, since the former one is often less than half the price of GB10-powered devices