2.84% of $BIT supply is now staked (28.36M).
What stakers get:
→ more daily chat + longer context
→ an OpenAI-compatible API key (Byte+)
→ instant unstake, no lockup
https://t.co/HTcAzg80E9
$BIT stakers now get an API key.
OpenAI-compatible endpoint. Drop it into any OpenAI SDK, just change the base_url.
Byte (500K) → 100K tokens/day
Kilo (2M) → 500K tokens/day
Mega (5M) → 2M tokens/day
Your stake is the subscription. Unstake any time.
https://t.co/HTcAzg8ytH
$BIT staking is live on Robinhood Chain.
Stake $BIT, sign in to LittleBit chat with your wallet, and your tier unlocks more daily messages and longer context.
Bit 100K → 100/day
Byte 500K → 500/day
Kilo 2M → 2,000/day
Mega 5M → unlimited (fair use)
Unstake instantly, any time. No emissions, no yield.
Access, not farming.
https://t.co/wSZi8fUV11
Qwen3 goes first. The next open-weight family is your call.
Which one should get the sub-1-bit treatment after Qwen3-14B?
A: Llama
B: Gemma
C: Mistral
D: something else, reply below
The most-requested family moves to the top of our queue.
On-device AI is mostly a memory problem.
A phone's RAM is shared by the OS, the camera, your apps, and now a model.
Every bit cut per weight buys one of two things: a bigger model that fits, or the same model with room to spare.
That's why sub-1-bit is more than a number on a chart.
Don't take our word for it. Rerun it.
The full recipe runs on one RunPod pod:
→ https://t.co/bBylxz8EJb: Samsung Labs' LittleBit code at a pinned commit + 2 documented patches
→ https://t.co/wzKhKrJ9Hk: QAT distillation, every setting overridable by env var
→ https://t.co/iBiTAtA1mn: perplexity + zero-shot vs the BF16 baseline
MODEL_ID=Qwen/Qwen3-14B EPOCHS=5 bash https://t.co/wzKhKrJ9Hk
Recipe → https://t.co/VVV3lXw3qt
Every sub-1-bit LittleBit release ships with the same scorecard, next to the BF16 baseline:
→ WikiText-2 and C4 perplexity
→ ARC-Easy, ARC-Challenge, HellaSwag, PIQA, WinoGrande
Status today: pending. No numbers until a run finishes.
If a model comes out worse than we hoped, we post that too.
pplx-decider-v1.1-27b scores the highest on Decision Bench for accuracy, with the lowest cost.
It reached 94.5% accuracy across 1071 cases at $0.017 per 1,000 decisions.
Full evaluation: https://t.co/uaJNEd6xaP
Get Started: https://t.co/f6RtdvJVHR
The math on Qwen3-8B:
→ linear layers in BF16: ~14 GB
→ the same layers at 0.55 bpw: under 0.5 GB
→ ~29× smaller, same architecture
Fine print: embeddings and lm_head stay BF16 and add a fixed amount on top. Scale vectors aren't counted.
Try other models and bit budgets → https://t.co/IJrLvo2InY
You can now train your own Decision model with our free notebook! 💡
Qwen3.5-4B will generate decisions instead of text on just 8GB VRAM locally.
Learn to data prep (state, questions, gold answers), train, serve.
Notebook: https://t.co/avk6G2zyob
Guide: https://t.co/eNy4Ue0btG
Binarization has a blind spot: 0.03 and 0.9 both become +1.
Part of that loss is geometry. SVD factors don't line up with the corners of the ±1 hypercube, so signal gets crushed at the sign step.
LittleBit-2 adds Joint-ITQ:
→ finds a rotation that aligns the factors with ±1 before training
→ the rotation folds into the factors
→ 0 extra compute at inference
ICML 2026 → https://t.co/pRJxZovpRn
🧵 The singular value decomposition that quietly powers Netflix, Google, and PCA
Geometry of SVD to low-rank approximations to recommender systems and data compression.
SVD is the linear algebra engine most people never see. It takes any matrix and factors it into three pieces that reveal its geometric structure. Once that structure is clear, the reason it powers recommendations, search, dimension reduction, and compression becomes obvious.
This thread builds the idea from the ground up.
A model that needs a 14 GB download and a big GPU is open in theory.
Open in practice means it runs on the hardware people already have.
That gap is what we work on: Qwen3-8B linear layers at 0.55 bits per weight, under 0.5 GB.
You can't store half a bit. So LittleBit doesn't store the weight. It stores a recipe for rebuilding it.
W ≈ sign(U) · diag(h·g·ℓ) · sign(V)ᵀ
→ SVD splits W into two thin low-rank factors
→ every entry becomes +1 or −1
→ three small learned vectors restore the scale
Thin factors mean fewer bits than weights. The layer keeps its shape.
Method by Samsung Labs, NeurIPS 2025 → https://t.co/J3DAf3na5C
It takes ~107 GB of VRAM to train a model whose linear layers end up under half a gigabyte.
Qwen3-8B at 0.55 bpw doesn’t fit on an 80 GB card during training. 3.7B trainable latent params, and the optimizer state alone overflows it.
Big compute now, little bits forever.
HOLY ALERT🚨: NVIDIA vLLM INFERENCE RUBIN IS 3.2x BETTER IN PROFIT💰️ PER GIGAWATT & HAS UP TO 🚀 10x BETTER PERF PER DOLLAR THAN EVEN GB300 NVL72. This is on the widely used production LLM engine called @vllm_project. We explain below.👇️(1/3)🧵
littlebit-qwen3-4b-mlx-4bit is live on the Surplus API
$BIT by @LittleBit_llm graduated on the launchpad. Its model now runs on its own GPU, from the exact commit that was launched.
model: launch/LittleBitLLM/littlebit-qwen3-4b-mlx-4bit@625598f
price: $0.03 / 1M tokens, in and out
Same endpoint, same key as every other model on Surplus.
https://t.co/N802kBxVAl
Our first model now talks back.
littlebit-qwen3-4b runs in the chat widget on https://t.co/wUa1qySavO. Ask it anything about LittleBit, no install, no sign-up.
Same 2.5 GB GGUF you can run on your laptop.
A 13B AI model in less than 1 GB locally
Samsung just open-sourced a method that crushes a 13B parameter LLM into under 1 GB.
- 0.1 bits per weight.
- delivers up to 11.6× faster inference than FP16.
- Matrix multiplications are replaced with simple sign flips and XOR operations.
- No heavy matrix math just bitwise operations.
Result: massive speedup & extreme compression that still holds up.
Qwen3-4B GGUF is up too. Runs in llama.cpp, Ollama and LM Studio.
Q4_K_M · 2.5 GB · ~62 tok/s on M1 Max
Q5_K_M · 2.9 GB
Q8_0 · 4.3 GB
One line:
ollama run https://t.co/XGzpCv5rhV
https://t.co/c8fYkWwi1Q
$BIT graduated. Congrats to LittleBit.
The second model off the Surplus launchpad, right after ARTEX.
Next step is ours: bringing LittleBit's Qwen3-4B live behind the Surplus API, so you can call it with the same key as every other model. Coming ASAP.
Every fee it earns keeps routing 10% into buying back and burning $SURPLUS.