NLP. NMT. Main author of Marian NMT. Research Scientist at Microsoft Translator.
This account formerly used the handle @marian_nmt which is now parked.
Blog post for NVIDIA DGX Spark havers:
Linux is hiding 4.1 GB of RAM from your DGX Spark.
On my 2× Spark setup, getting it back took my GLM-5.3-Flash KV pool from 262k → 938k tokens.
Two tricks I haven’t seen discussed together:
• 64 KiB kernel
• reclaim the headless display carveout
I don't like the silence of the GPT 6.x generation of models while they work. GPT 5.6 Sol's communication style is pretty much ideal. I wonder if there will be work to bring that back.
@ArthurCDent I like the local models, too, but unfortunately, they seem to all be trained with mostly Claude's thinking traces.
That in turn means they have all its neuroticism but not quite its brilliance :)
Waiting for some local model to steal better from the other guys.
• Mapped both sensors, CSI-2 routing and RAW10 formats.
• Traced front-camera failures to receiver reset timing—and fixed it.
• Ported IPU4P to Linux 6.8 without disabling Secure Boot.
• Turned raw frames into upright color video.
• Added on-demand capture: cameras stay listed, but sleep when unused.
• Fixed GNOME Camera discovery, crashes and switching.
• Published the working source, patches and setup guide.
I keep being amazed that you can just do things now, and the only blocker is "actually thinking of them".
For years, my old Surface Book 3 didn’t have camera driver on Linux. Today, I told Copilot to fix it, and two hours later I’m testing my camera in WhatsApp.
https://t.co/m74jATOXQv
I should probably contribute this somehow.
@deliprao I am new enough to the local hosting stuff that I have substantial doubt in my own skills. :)
What helped was that MiniMax also aborted occasionally with stalled tool calls. Made me look again.
PSA: seems like both MiMo V2.6 Flash variants seem basically impossible to get working reliably as coding agents.
The original loops/tool-storms, MOPD randomly stops or drops out, and neither has made it through a single SWE-bench task cleanly for me.
User error not excluded.
@deliprao They introduced the bug between two docker image downloads for me while I was messing around with MiMo which does actually have tool calling issues between its two versions. Was a hard find.
PSA no. 2: Turns out that is an SGLang serving bug (https://t.co/ZVk5b66ZgD). GLM and MiniMax were hit by the same bug. The tool-storming of Flash seems to be fixed with the MOPD version.
PSA: seems like both MiMo V2.6 Flash variants seem basically impossible to get working reliably as coding agents.
The original loops/tool-storms, MOPD randomly stops or drops out, and neither has made it through a single SWE-bench task cleanly for me.
User error not excluded.
@deliprao Turns out that is an SGLang serving bug (https://t.co/ZVk5b66ZgD).
GLM and MiniMax were hit by the same bug.
The tool-storming of Flash seems to be fixed with the MOPD version.
@mar_kar_ I nearly got a fourth one but stopped myself because connecting them efficiently requires a ConnectX switch which seems to be an insane piece of hardware.
However, I am very self-satisfied right now.
I have been looking into soldering more RAM onto the 128GB versions, but I expect now that jerry-rigging the 64GB versions back to 128GB might become a whole new industry. If the firmware/BIOS is the same, the chips demonstrably exist.
Individual compute inequality, I realized recently, might be one of the bigger societal inequalities we are seeing right now.
Having the perk of near-unlimited frontier model access for reasonable private purposes must be one of the biggest benefits that working for big tech offers right now.
Ultra Fast is wild but makes me feel super compute-poor given how quickly it blows through the budget.
There are now ~3 tiers of compute-havers - free tier, paid tier, and people at a few companies.
The gap between 2nd and 3rd has grown a lot w/ multi-agent stuff + now speed