Core Boom. A spatial shooter game is possible for SPECS glasses. We worked a lot on this. Hope everyone who gets the chance to try it enjoys it as much as we do.
@specs@specsfordevs
< Choosing a Vision Backbone >
your model’s backbone is its perspective
pick ResNet, and it sees in edges
pick a ViT, and it sees in patches
the backbone decides how your model thinks
here are some of the most practical backbones and when you should choose them, from the paper "Battle of the Backbones" (2023):
> ResNet - good for fast prototyping, small models, and edge devices
> ConvNeXt - great all-purpose backbone; strong for detection & segmentation
> Swin Transformer (V2) - best for large-scale detection, segmentation, and high-res inputs
> ViT (Vision Transformer) - good when you have huge datasets; less bias, more global context
> CLIP - best for vision-language, zero-shot, and retrieval tasks
> DINO / MoCo / MAE (SSL) - great when you have little or no labeled data
> MiDaS - surprisingly strong if you care about depth, geometry, or robotics perception
> Stable Diffusion Encoder - useful for creative or aesthetic tasks; not for accuracy-critical CV
> EfficientNet / RegNet / ResNet-18 - good lightweight options for edge or mobile deployment
Introducing Gemma 3 270M, a new compact open model engineered for hyper-efficient AI. Built on the Gemma 3 architecture with 170 million embedding parameters and 100 million for transformer blocks.
- Sets a new performance for its size on IFEval.
- Built for domain and adoption and specialized fine-tuning.
- Uses just 0.75% battery for 25 conversations on Pixel 9 Pro.
- Instruction-tuned and base model.
- INT4 Quantization-Aware Trained checkpoints.
The freshest AI/ML research of the week
Our top 9
▪️ Sotopia-RL: Reward Design for Social Intelligence
▪️ Agent Lightning: Train ANY AI Agents with RL
▪️ Exploitation Is All You Need... for Exploration
▪️ Learning to Reason for Factuality
▪️ VeOmni
▪️ Is Chain-of-Thought Reasoning of LLMs a Mirage?
▪️ Cognitive Loop via In-Situ Optimization
▪️ Sculptor
▪️ CoAct-1
▪️ Tool-integrated Reinforcement Learning for Repo Deep Search
▪️ RL-PLUS
▪️ SEAgent
▪️ CRINN
▪️ Training Long-Context, Multi-Turn Software Engineering Agents with RL
▪️ Beyond the Trade-off: Self-Supervised RL for Reasoning Models' Instruction Following
▪️ CompassVerifier
▪️ Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
▪️ Are Today's LLMs Ready to Explain Well-Being Concepts?
▪️ VeriGUI
▪️ Trainable Dynamic Mask Sparse Attention
▪️ LeanK
▪️ Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
▪️ On the Generalization of SFT
▪️ SitEmb-v1.5
▪️ AttnTrace
▪️ LaTCoder
▪️ ChartCap
🧵
Google introduces Deep Researcher with Test-Time Diffusion.
A novel framework that redefines AI research report generation as an iterative diffusion process, inspired by human cycles of searching, reasoning, and revision.
China dropped these open-source models in July:
- GLM-4.5
- GLM-4.5-Air
- Wan-2.2
- Qwen3 Coder
- Qwen3-235B-A22B-Thinking-2507
- Qwen3-235B-A22B-2507
- Kimi K2
Meanwhile:
- OpenAI still hasn’t released the open-source model
- Anthropic is doing stricter rate limits
- Meta might be walking away from open source
The future of open AI may not be led by the West.
We’re open-sourcing the pre-training code for Phi4-mini-Flash, our SoTA hybrid model that delivers 10× faster reasoning than Transformers — along with μP++, a suite of simple yet powerful scaling laws for stable large-scale training.
🔗 https://t.co/Nxsm6FclOX
(1/4)
After MCP, A2A, & AG-UI, there's another Agent protocol.
It's fully open-source and launched by IBM Research.
Here's a complete breakdown (with code):
.@MistralAI releases Voxtral, open source speech understanding models built for real-world applications
Available in two sizes
🔹 Voxtral 24B for production-scale deployments
🔹 Voxtral Mini 3B optimized for local and edge use
Voxtral delivers state-of-the-art transcription, translation, summarization, and chat-ready speech intelligence, all under the Apache 2.0 license.
✅ Run locally via Hugging Face
✅ Integrate via API starting at $0.001/min
✅ Built for speed, accuracy, and scalability
Check out the cool speech-to-speech demo with Voxtral and Inworld https://t.co/HQwNQezA06
Blogpost: https://t.co/BuiUu5b780
Just released on HF: NVIDIA's Audio Flamingo 3 the state-of-the-art audio-language model!
> Multi-audio reasoning
> Voice-to-voice Q&A
> Handles long audio (up to 10 min) with on‑demand chain‑of‑thought
> Open‑source code + 4 new benchmarks
7. MemAgent
Introduces an RL–driven memory agent that enables transformer-based LLMs to handle documents up to 3.5 million tokens with near lossless performance, linear complexity, and no need for architectural modifications.
https://t.co/q9BF2u9Go0
NeuralOS
Towards Simulating Operating Systems via Neural Generative Models
a generative OS that predicts screen images from user inputs, combining an RNN for computer state modeling and a diffusion model for rendering
1. Ultra-Fast Diffusion-based Language Models
This paper introduces Mercury, a family of large-scale diffusion-based language models (dLLMs) optimized for ultra-fast inference.
https://t.co/OI1RwyzSmB
PDF parsing is still painful because LLMs reorder text in complex layouts, break tables across pages, and fail on graphs or images.
💡Testing the new open-source OCRFlux model, and here the results are really good for a change.
So OCRFlux is a multimodal, LLM based toolkit for converting PDFs and images into clean, readable, plain Markdown text.
Because the underlying VLM is only 3B param, it runs even on a 3090 GPU. The model is available on @huggingface .
The engine that powers the OCRFlux, teaches the model to rebuild every page and then stitch fragments across pages into one clean Markdown file.
It bundles one vision language model with 3B parameters that was fine-tuned from Qwen 2.5-VL-3B-Instruct for both page parsing and cross-page merging.
OCRFlux reads raw page images and, guided by task prompts, outputs Markdown for each page and merges split elements across pages.
The evaluation shows Edit Distance Similarity (EDS) 0.967 and cross‑page table Tree Edit Distance 0.950, so the parser is both accurate and layout aware.
How it works while parsing each page
- Convert into text with a natural reading order, even in the presence of multi-column layouts, figures, and insets
- Support for complicated tables and equations
- Automatically removes headers and footers
Cross-page table/paragraph merging
- Cross-page table merging
- Cross-page paragraph merging
A compact vision‑language models can beat bigger models once cross‑page context is added.
🧵 1/n Read on 👇
What will software development look like in 2026?
With coding agents rapidly improving, dev roles may look quite different. My current workflow has changed a lot:
- Work in github, not IDEs
- Agents in parallel
- Write English, not code
- More code review
Thoughts + a video👇