@EthosVentures@MiTypeScript Image tokens as continuous embeddings can hold more information than quantized tokens. So the efficiency gain can be real and then it doesn’t mean they are mispriced wrt the amount of compute.
@NicholasBardy@VictorTaelin@zhijianliu_ Wouldn’t any change in the model require retrained dflash? What makes a lora fine tuned model different? Dflash wouldn’t be trained for the new models output distribution.