@CTOAdvisor@nvidia As a quick warning, keep an eye at longer context on if your engine tries to do speculative on the whole ctx - tends to be bad, especially for dflash (where with the wrong setup you can get huge gains at low ctx but negative results at long; speculative_window etc here)
@CTOAdvisor@Apple@nvidia It's VERY easy to just run mlx-convert and get a quant. quick check: size of checkpoint? eg, if you have 13GB of mlx weights and 16GB of NVFP4, odds are your nvidia quant is much stronger because that 3GB is a lot of critical layers kept at FP16/32
@CTOAdvisor@Apple@nvidia places to check
1. quantizations: different formats do not REQUIRE but authors will often avoid quantizing some things (eg, quantize sparse experts but not ffn layers or exeprt routers)
2. kv cache dtype - the loss from weight quants is constant; it snowballs in kv; use fp16/bf16
@jimcramer 1. trillions of dollars is not "a few bucks"
2. when has protectionism ever worked? imagine if China sold oil 87.5% cheaper and you wanted American companies to pay 8x fuel cost for everything. Does that make them more globally competitive?
Off my normal AI topics, but if you're into M:tG or gaming in general, my friend Brian (also a co-founder of GGG / @pathofexile ) interviews Richard Garfield on his burgeoning gaming channel. https://t.co/lOnsMAcFu8