First of three releases. Chat-optimised instruct model and a DeepSeek R1 reasoning model with <think> tags coming soon. Same binary, swap GGUF.
We're Evrmind a UK startup, AI safety + edge compute. [email protected]
Model: https://t.co/naqz4T4l9i
Code: https://t.co/5fjtIZV5zT
We released EVR-1 Maano: a 3.93 GiB compression of Llama 3.1 8B that stays coherent where standard 3-bit quants collapse into loops.
Under 6% repetition at 500 tokens. Standard quants hit 77-80%.
Not GPTQ, AWQ, or standard GGUF, novel compression.
https://t.co/naqz4T4l9i
๐งต
Runs on laptop, desktop, Mac, Windows, Linux, even Android via Termux.
~34 tok/s on RTX 6000 Ada
~8 tok/s on Mac Mini M4
Ships with a browser-based chat UI, run one script, open localhost:8080. No cloud, no API key. 3.93 GiB model file.
You don't need a bigger budget to beat a corporation. You need Velocity.
As a Solopreneur, you are light. Your decision speed is instant.
The goal of Evrmind isn't just to help you "keep up". It's to help you outmanoeuvre.
Don't wish for their budget. Leverage your speed.