you can build your own small language model from scratch on your computer - this free app walks you through every step 🤯
it's called Language Model Builder, and it comes in two parts.
the first is a 90-minute interactive textbook.
it covers the fundamentals, tokenization, embeddings, attention, the transformer, training data, loss and gradient descent, fine-tuning. no coding or math background needed.
and instead of staring at diagrams, you run each experiment yourself as you read.
the second is a training workbench that runs the full pipeline locally on Apple's MLX:
> pre-train a base model on a dataset you pick
> fine-tune it with SFT, then align it with DPO
> watch the loss curve drop in real time as it learns
> chat with the model you made, with a token-probability x-ray view so you can see how it "thinks"
on default settings you'll get coherent, multi-paragraph text in about a day.
a GPT-2-small-class model takes roughly a week on a high-end Mac.
100% free, no account, runs entirely offline.
While waiting for DeepSeek V4 we got two very strong open-weight LLMs from India yesterday.
There are two size flavors, Sarvam 30B and Sarvam 105B model (both reasoning models).
Interestingly, the smaller 30B model uses “classic” Grouped Query Attention (GQA), whereas the larger 105B variant switched to DeepSeek-style Multi-Head Latent Attention (MLA).
As I wrote about in my analyses before, both are popular attention variants to reduce KV cache size (the longer the context, the more you save compared to regular attention).
MLA is more complicated to implement, but it can give you better modeling performance if we go by the ablation studies in the 2024 DeepSeek V2 paper (as far as I know, this is still the most recent apples-to-apples comparison).
Speaking of modeling performance, the 105B model is on par with LLMs of similar size: gpt-oss 120B and Qwen3-Next (80B). Sarvam is better on some tasks and worse on others, but roughly the same on average.
It’s not the strongest coder in SWE-Bench Verified terms, but it is surprisingly good at agentic reasoning and task completion (Tau2). It’s even better than Deepseek R1 0528.
Considering the smaller Sarvam 30B, the perhaps most comparable model to the 30B model is Nemotron 3 Nano 30B, which is slightly ahead in coding per SWE-Bench Verified and agentic reasoning (Tau2) but slightly worse in some other aspects (Live Code Bench v6, BrowseComp).
Unfortunately, Qwen3-30B-A3B is missing in the benchmarks, which is, as far as I know, is the most popular model of that size class. Interestingly, though, the Sarvam team compared their 30B model to Qwen3-30B-A3B on a computational performance analysis, where they found that Sarvam gets 20-40% more tokens/sec throughput compared to Qwen3 due to code and kernel optimizations.
Anyways, one thing that is not captured by the benchmarks above is Sarvam’s good performance on Indian languages. According to a judge model, the Sarvam team found that their model is preferred 90% of the time compared to others when it comes to Indian texts. (Since they built and trained the tokenizer from scratch as well, Sarvam also comes with a 4 times higher token efficiency on Indian languages.
We've rolled out a new auto-memory feature.
Claude now remembers what it learns across sessions — your project context, debugging patterns, preferred approaches — and recalls it later without you having to write anything down.
Shipping a real‑time price alert system: clean architecture + Kafka + structured AI workflows.
Built in focused sessions with Claude Code + structured Markdown playbooks
Repo: https://t.co/KAqsrLkFR6
Wiki: https://t.co/8sxRSym70V
#Java25#SpringBoot#Kafka#Microservices
New in Claude Code: Remote Control.
Kick off a task in your terminal and pick it up from your phone while you take a walk or join a meeting.
Claude keeps running on your machine, and you can control the session from the Claude app or https://t.co/er6Blrr63e
@letsblinkit 🚨 SHOCKING EXPERIENCE WITH BLINKIT! 🚨
I ordered 1 kg of Nandini Milk Peda from @letsblinkit, expecting fresh sweets. But to my shock, they delivered EXPIRED products! 😡
Food safety is NOT a joke! Selling expired items is dangerous. (1/4)
I urge @letsblinkit to take immediate action to fix this issue. Customers deserve fresh, safe products, NOT expired goods that could make people sick!
Have you faced something similar? Let’s make some noise! 🚨 (3/4)
@airindia Extremely disappointed with air india service , as they missed my wife checkin bag from Delhi to Paris airport today, no communication from air india regarding same