Next: a new base model.
v2 was built on a 2024 7B. v3 moves to a newer, stronger base that still runs on your own machine.
Then: train, benchmark against v2, publish the numbers. Win or lose.
Artex v3 update: the data is ready.
30K examples, up from 4K:
• 12K execution-filtered Python
• 8K with unit tests we re-ran ourselves
• 5K multi-turn debugging
• 5K multi-language
We also checked every sample against HumanEval+ & MBPP+ and pulled anything that overlapped.
If v3 scores higher, it’s because it learned to code, not because it saw the test.
Artex v3 dataset, the plan:
→ 4K → ~30K samples
→ execution-verified code, not just “looks right”
→ multi-turn “fix this bug” conversations
→ decontaminated against benchmarks
And every version now gets scored on
HumanEval+/MBPP+. No vibes, just numbers.
Artex Coder 7B was 18% slower than the model it's built on.
Same architecture. Same 4-bit quant. Same 4.0 GB file. So why?
Local decoding is memory-bound. Every token streams the whole model through memory, so two files of the same size should run at the same speed.
The bytes matched. The format didn't.
Our MLX conversion stored the quantization scales as BF16. The base model uses F16. Apple M1 has no hardware for BF16, so every token paid for it.
Reconverted with F16 scales, same weights, same quality, on my M1 Max:
→ before: 53 tok/s
→ after: 63 tok/s
→ +19%, file still 4.0 GB
Still a bit behind the base model's ~69. Speculative decoding is next.
Found it ourselves, fixing it in public.
@akshay_pachaar Same math on an M1 Max: ~400 GB/s ÷ a 4.0 GB 4-bit 7B ≈ 100 tok/s ceiling.
Our Artex Coder 7B build was doing 53. Turned out our MLX conversion stored scales as BF16, which M1 has no hardware for. Reconverted to F16: 63 tok/s.
Speculative decoding is next.
The first model launched on Surplus @Artex_llm just did this:
→ Graduated in 2h 5m
→ $166.7K volume in its first 15 hours
→ 1.86 ETH in creator fees, split by contract
→ 3.66M $Surplus bought back and burned
$ARTEX · artex-coder-7b, a 7.6B Qwen2 coder.
Launch yours: https://t.co/oL0WUvbNzM
Ask a fine-tuned model "who are you?" and most say "I'm Qwen" or "I'm ChatGPT".
Artex says "I'm Artex." No system prompt. It's in the weights.
Small thing. But it's the difference between a wrapper and a model.
Artex Coder 7B on @SurplusRH: $0.03 per 1M tokens. In and out.
Same-size models on @OpenRouter today, output price per 1M:
→ Llama 3.1 8B: $0.08
→ Qwen3.5 9B: $0.15
→ Qwen2.5 7B: $0.20
→ Qwen3 8B: $0.455
Cheaper than every paid 7–9B model we could find.
Or run it yourself for $0.
The full recipe behind Artex Coder 7B:
→ base: Qwen2.5-Coder 7B (open weights)
→ data: 4,000 samples from Magicoder (MIT) + identity chats in EN/ID
→ method: LoRA r=16, all linear layers
→ 2 epochs, lr 1e-4, completion-only loss
→ 1× H100, ~18 minutes
No secret sauce. That's the point of open weights: you can rebuild it yourself.
Every API call is rent.
Open weights are ownership.
Artex Coder 7B is a 4 GB file. Download it once, run it forever.
No per-token bill. No rate limits. No "model deprecated" email.
Your code never leaves your machine.
Two ways to run Artex now:
→ on your machine: 4 GB file, free, fully offline
→ on @SurplusRH: $0.03 per 1M tokens, same key as every other model
Same weights either way.
Open weights means you pick.
artex-coder-7b is live on the Surplus API 🟢
The first open weight model launched on the Surplus launchpad now runs on its own GPU, from the exact commit @Artex_llm launched.
model: launch/0XARTEX/artex-coder-7b@9182fd6 price: $0.03 / 1M tokens, in and out
Same endpoint, same key as every other model.
A 7B coding model in 4 GB. 56 tok/s. On a laptop, no cloud, no API key.
This is Artex Coder 7B, running on my own M1 Max. 4-bit MLX, a 4.0 GB file.
How it was built:
- base: Qwen2.5-Coder 7B
- data: 4,000 coding examples + identity chats in English and Indonesian
- method: LoRA, 2 epochs
- training: ~18 min on one H100
- val loss: 0.286 -> 0.172
What it does on my machine, same 6 prompts vs the base model, both 4-bit:
- peak memory: 4.44 GB -> 4.44 GB
- decode: 69 -> 56 tok/s
- total answer length: 1,929 -> 1,899 tokens
So the honest part: Artex is not faster than its base. It's about 18% slower on the same quant, and I haven't found why yet. And it's barely shorter. "Concise" is still the goal, not the result.
What did change: it knows who it is and replies in your language, Indonesian included. That's what this run trained.
Next: find the speed loss, train harder on short answers, run real benchmarks and post them, good or bad.
Open weights, Apache 2.0, runs on your desk.
weights: https://t.co/cUugOnfbof
Congrats @Artex_llm
$ARTEX just migrated. The first model launched on Surplus filled its bonding curve, and its liquidity now sits in a Uniswap v4 pool, locked forever.
Next: we deploy the exact commit of artex-coder-7b they launched and serve it on the Surplus API.
Live soon.
https://t.co/Na9KF1SXtP
Today we're releasing Artex Coder 7B
Based on Qwen2.5-Coder 7B, Artex was fine-tuned in ~18 minutes on a single H100 to do one thing well: give concise, accurate, to-the-point coding answers
It runs locally on your Mac and replies in your language, Indonesian included
Artex is available today under Apache 2.0.
https://t.co/mZZWVBU6Kt
$ARTEX is live on @SurplusRH
The first project to launch on Surplus.
Artex is a 7B coding model that answers in code, not essays.
→ Weights are open: download free on Hugging Face
→ Don't want to run it? Surplus hosts Artex, pay per call via API
CA: 0xf1cc0a4feed884a4490a1d8db943c0a92bd83347