A real revolution in local models will be made by the one that, with quantization of 4, will occupy about 25 GB and have 1-2 billion active parameters. This will allow you to run it on almost any PC with 32 GB of RAM, regardless of the presence of a graphics card.
A real revolution in local models will be made by the one that, with quantization of 4, will occupy about 25 GB and have 1-2 billion active parameters. This will allow you to run it on almost any PC with 32 GB of RAM, regardless of the presence of a graphics card.
A real revolution in local models will be made by the one that, with quantization of 4, will occupy about 25 GB and have 1-2 billion active parameters. This will allow you to run it on almost any PC with 32 GB of RAM, regardless of the presence of a graphics card.
@thorstenball The difference, in fact, is only in either writing or pressing the plan button. And yet, in working mode, he can begin to change something at his discretion, which is not good