I've burned 1.4 billion tokens in Claude Code. I had no idea what that even meant.
So I built STAMP Strava for your AI usage.
It reads your local logs and turns those tokens into steps, distance, and energy with a shareable route card.
100% local. One command:
npx token-stamp
DeepSeek just released DSpark for V4 Flash & Pro, a new speculative decoding method boosting throughput by 51% to 400%!
DS also showed DSpark works well for other models like Gemma & Qwen
Github: https://t.co/EGVYpc1kcK
Paper: https://t.co/TaBMRVlaW9
HF: https://t.co/289jVU2pxh
@svembu@svembu Sir hello i know you might not reply for this but i have tired to build a sovereign ai made by me only me no wrapper, no fine tuning nothing just pure ai i tired on medical side and domain specific not for general medicine but this is just a person trained no backing up
@pratykumar Hire me man ,i will contribute for the company more anyone in your company
I train models
I train agents
I build software for governments
What else? Fancy degree ?
what to build in AI Engineering
> Build your own Reasoner (Chain of Thought implementation)
> Build your own Agent loop (ReAct pattern)
> Build your own Inference Server (in C++/Rust)
> Build your own Transformer from scratch (Attention is all you need)
> Build your own Vector Database (HNSW index)
> Build your own RAG pipeline
> Build your own Flash Attention kernel (CUDA)
> Build your own Quantization library (Int8/FP4 implementation)
> Build your own Mixture of Experts (MoE) routing layer
> Build your own Distributed training loop (FSDP/Tensor Parallelism)
> Build your own KV Cache paging system (like vLLM)
> Build your own Speculative Decoding system
> Build your own State Space Model (Mamba implementation)
> Build your own RLHF pipeline (PPO implementation)
> Build your own Small Language Model (SLM)
> Build your own Matrix Multiplication kernel
> Build your own LoRA (Low-Rank Adaptation) trainer
> Build your own Code interpreter sandbox
> Build your own DPO (Direct Preference Optimization) loss function
> Build your own Graph RAG system
> Build your own Model merger (Model Soups/Spherical Linear Interpolation)
> Build your own Interpretability tool (SAE - Sparse Autoencoders)
> Build your own Synthetic data generator
> Build your own Function Calling router
> Build your own Structured Output parser (Context Free Grammars)
> Build your own Multi-modal projector (CLIP implementation)
> Build your own LLM Eval harness
> Build your own Guardrails system (Input/Output filtering)
> Build your own Prompt caching mechanism
> Build your own Tokenizer (BPE implementation)
> Build your own Autograd engine (like Micrograd)
> Build your own Diffusion model (UNet + Scheduler)
> Build your own Vision Transformer (ViT)
> Build your own Whisper-style ASR model
> Build your own Text-to-Speech pipeline
> Build your own Semantic Router
> Build your own Knowledge Graph builder
> Build your own Data curation pipeline (MinHash/Deduplication)
> Build your own AI Gateway (Load Balancing/Failover)
> Build your own Parameter Efficient Fine-Tuning (PEFT) library
> Build your own Text-to-SQL engine
> Build your own Recommendation system (Two-tower architecture)
> Build your own Embedding model
> Build your own Logit Processor
> Build your own Softmax kernel optimization
> Build your own Adversarial attack generator
> Build your own Audio Spectrogram transformer
> Build your own Neural Architecture Search
> Build your own Model Distillation pipeline
> Build your own Feature Store
> Build your own Database driver (for Vectors)
#AI
Implemented Google's TurboQuant paper from scratch 4-7x KV cache compression for LLMs, zero quality loss
They published the math. Released no code. So I built it
Tested on 6 models (7B→70B), 5 architectures. Now pip-installable
@GoogleDeepMind@Google
I dropped 1000 AI agents into a world with blank brains.
No training data. No instructions. No reward function.
One rule: survive or die.
806 generations later, they evolved communication on their own.
Watch purple (random) turn to red (evolved):
@anirudhbv_ce@GoogleResearch I implemented this from scratch on H100. Got 44-59% KV-cache reduction with exact prefill fidelity across 5 model families. Full benchmarks in my thread i found something that no one noticed
@alex_prompter Dropped 1,000 AI agents into a world with blank brains. No training. No instructions. No reward.
One rule: survive or die.
806 generations later they invented language.
@RoundtableSpace Dropped 1,000 AI agents into a world with blank brains. No training. No instructions. No reward.
One rule: survive or die.
806 generations later they invented language.
@Hesamation Dropped 1,000 AI agents into a world with blank brains. No training. No instructions. No reward.
One rule: survive or die.
806 generations later they invented language.