a tiny AI 0.5B just turned a normal phone video into a fully labeled 3D room run phone.
SpatialLM that understands walls, doors, furniture, and everything layout.
it trained as a standard multimodal LLM.
Most vision models see pixels, SpatialLM sees space.
Open models : https://t.co/kGXjHkIdiE
AI inference could grow from $160B to $350B next year, overtaking databases as the largest market in software.
For software companies, that could mean margins closer to 30–50% than the usual 72%.
Someone is running Qwen3.8-35B-A3B on a single RTX 3060 with just 12GB of VRAM and 16GB of system RAM.
And the reported numbers are surprisingly good.
~50 tok/s decode ~550 tok/s prefill 262K context window 12GB VRAM 16GB system RAM
No multi-GPU setup. Just a consumer graphics card and a carefully configured local inference engine.
The model is particularly interesting because it’s Qwen3.8 distilled into the Qwen3.6-35B-A3B architecture.
That makes it worth investigating beyond raw inference speed.
The developer currently uses Ornith-1.5 as the engine behind their Hermes and Pi agents. Now the question is whether this new model can deliver enough quality and reliability to replace it.
The next step is testing it as an agent, not just a chatbot.
The planned comparison includes four models:
• Qwen3.8-35B-A3B • Ornith-1.5-35B • Qwen3.8-27B • Base Qwen3.6-35B-A3B
An agentic coding benchmark should help reveal whether the distilled model can follow instructions, modify code, solve problems, and complete multi-step tasks effectively.
That distinction matters because high token throughput doesn’t automatically translate into a better coding agent.
A model can generate code quickly and still struggle with planning, debugging, tool calls, or completing a task without intervention.
The hardware configuration is also worth examining.
The developer is using llama-server with the Q4_K_M GGUF quantization, Flash Attention enabled, and 26 MoE layers assigned to CPU processing.
The configuration also enables a 262,144-token context window, Q8 KV-cache quantization, and reasoning mode.
This is how the setup is configured:
llama-server \
-m Qwen3.8-35B-A3B-Q4_K_M.gguf \
-ngl 99 \
--n-cpu-moe 26 \
-c 262144 \
-fa on \
--jinja \
-np 1 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--temp 0.6 \
--top-p 0.95 \
--top-k 20 \
--reasoning on
One important detail: the model’s 35B parameter count doesn’t mean all 35 billion parameters are activated for every token. The A3B designation indicates roughly 3 billion active parameters per token.
That sparse MoE architecture helps explain how a model of this size can deliver high generation speeds on modest hardware. CPU-based expert processing also helps work around the GPU’s limited memory, although the exact trade-offs depend on the workload.
The 262K context claim is equally interesting. Supporting a context window that large doesn’t mean every request at that length will maintain the same speed.
Prompt length, KV-cache memory, and CPU/GPU placement can all affect performance.
The developer’s next test is particularly useful: graphing decode speed against context length.
That should help show how performance changes as the prompt grows, rather than presenting one speed figure without the conditions behind it.
If the coding benchmark also holds up, this could become a serious option for people who want capable local agents without investing in expensive GPUs.
For now, the reported inference numbers are promising. The real test is whether Qwen3.8-35B-A3B can match or outperform Ornith-1.5-35B on actual agentic coding tasks.
Fast local inference gets your attention. Reliable autonomous task completion is what earns a model a place in the stack.
On Tuesday, Sierra announced Personal Agent Protocol with Meta, Genesys, Instinct, Rocket, Shopify, Stripe, and Walmart. The response has been incredible and today we’re publishing the first draft of the protocol at https://t.co/Qao9yEfzx4 and announcing 35 new design partners.
As personal agents take on more, our goal is simple: consumers have choice and control, and companies know when they’re dealing with an agent and who it represents.
This draft is a starting point. We’d love for companies, developers, and personal agent builders to read it, give feedback, and help us build an open standard together.
Los fundadores de JEV dedicaron más de 2 años a su desarrollo y la mayoría de laboratorios de IA les han copiado en 2 semanas.
Malos tiempos para los innovadores
- Gyanesh Kumar daughter is an IAS officer.
- Gyanesh son-in-law is an IAS officer.
- Gyanesh younger daughter is an IAS officer.
- Even Gyanesh younger son-in-law is also an IAS officer.
And all these miracles started happening after 2014..