we have ported and further optimized our earlier @koniclabs optimized checkpoint of @liquidai's LFM2.5-VL-3B for Apple Silicon
its now aggressively optimized and native on your Macbook
go see it.
@liquidai's LFM2.5-VL-3B, now native on Apple Silicon
The INT4/INT8 weights map to MLX bit-exactly (max diff 0.0) — then we chased the remaining bytes.
- Decode 12.16 → 62–68 tok/s (5.1–5.6×) · text TTFT 167 → 62 ms
- Checkpoint 2.79 → 1.63 GB (−42%) · peak RSS ~4.8 → ~1.9 GB (−60%)
- M3 decode is bandwidth-bound (~0.94 GB weights/token @ ~70 GB/s) — every win was a bandwidth cut, not a kernel trick
Go download it.
https://t.co/ISYH4tCUsK
another day of torturing @liquidai's LFM2.5
we surgically pruned its FFN layers, aligned it back to the original base with distillation, and compressed it by ~70%
the funny thing is compressed version is sometimes better than the base in responses and tool-calls.
go see it.
New research from Konic Labs: Surgical FFN width pruning + distillation recovery + mixed-precision quantization on @liquidai's LFM2.5-VL-3B.
- 6.25 GB → 5.30 GB (−15%) → 4.92 GB (−21%) → 1.93 GB packed / 2.22 GB (Konic Optimized), −69%
- Near-lossless vs base: PPL ratios 0.79–0.85 — at or below baseline
- Top-5 token agreement: 0.95–0.99; quantization adds just +0.02–0.06 nats
- Tool calling: 10/10 rounds on the smallest tier — even abstaining correctly where the base over-triggered
The hypothesis: an arch-searched hybrid conv+attention VLM still carries compressible redundancy — and it lives in FFN width, not depth. Joint-SwiGLU Wanda prunes, distillation-to-baseline LoRA recovers, hand-rolled GPTQ packs.
See the full research and models.
https://t.co/7gipvusodj
new research publication from our lab.
we have tried to make @liquidai's LFM2.5 encoder multimodal by attaching a visual encoder to it.
obviously needs more training to stitch them together but there is hope.
go read it.
Making the LFM-2.5 encoder multimodal.
@liquidai's LFM2.5 Encoder 230M + SigLIP2 = a compact multimodal encoder.
- 12.5k held-out MONET pairs: i→t R@1 0.1194 (BF16)
- GPTQ INT4 keeps ~91% — 370 MB, −59.9% vs original
Go read it.
https://t.co/F2TdTgp7hX
the open superintelligence lab with the cool aesthetic branding just dropped a literal neurosymbolic AGI system that evidently satisfies both marcus' and chollet's requirements and u're here dooming?
For those who understand the impact of THIS paper, this is how GOATED this paper was:
More than 100,000 citations right after 2 years of it's publish.
THIS paper was so good that some researchers said that it looked like someone from future jumped in to the present timeline and just gave the researchers EXACTLY the KEY, the 'MISSING BLUEPRINT'.
Took care of parallel training, semantic word embeddings, took care of both long form data ingestion and output at the same time.
It solved the bottlenecks that had held deep learning back for YEARS:
Massive parallel training instead of sequential computation, process LONGGGGG sequences while generating coherent outputs WHILE MAINTAINING THEIR SEMANTIC MEANINGS.
A simple (arugubaly, THE SIMPLEST) foundation that could scale simply by adding more data and compute.
what could've taken DECADES to accomplish in the AI field, took place in the last 4 years, all shouldered to THIS paper.
And no this post is NOT AI generated btw and your dumbass would be as hyped as I am if you actually understood what ripple effect "Attention" mechanism CAUSED.
Following Mooncake, AgentENV marks the next project we've built and open-sourced together, supporting the infrastructure behind Kimi K3’s agentic RL training.
Grateful to the @Kimi_Moonshot team for the close collaboration, and looking forward to building more open AI infrastructure together.
@KVCache_AI@Kimi_Moonshot here we have something similar to agentENV
specifically built for building ART (@OpenPipeAI) trajectories, docker-based
just released today
https://t.co/MKPbJZCpdx
Introducing "agentbox".
Want to run RL (GRPO/SFT) on real agent trajectories — not synthetic "think step by step" chains?
agentbox gives you:
🐳 Ephemeral Docker sandboxes, one per rollout
🛠️ Tool-calling LLM agents (OpenAI protocol)
✅ Pytest/shell verifiers for objective rewards
📦 ART-native export → directly into your training pipeline
> pip install agentbox-rl
https://t.co/RlWlqeqWGL
Introducing "agentbox".
Want to run RL (GRPO/SFT) on real agent trajectories — not synthetic "think step by step" chains?
agentbox gives you:
🐳 Ephemeral Docker sandboxes, one per rollout
🛠️ Tool-calling LLM agents (OpenAI protocol)
✅ Pytest/shell verifiers for objective rewards
📦 ART-native export → directly into your training pipeline
> pip install agentbox-rl
https://t.co/RlWlqeqWGL
@sudoingX great take and genuinely support it.
however, we need to make them even smaller. people should be aware that 128GB VRAM still a luxury for most of the world.
Need intelligence in the palm of our hands. @koniclabs is focusing exactly that.
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
Claude Mythos was a marketing stunt for the IPO of Anthropic
Kimi-K3 and all these open source model releases this month just fucked it up
I hope it goes much more sideways for wannabe tech oligarchs.
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f