I made GeoGuessr but for the Quran.
you hear a recitation, you guess the surah and verse.
built it yesterday for when I had a lapse in memorization. hundreds of people have practiced on it since and i genuinely didn't expect this
Try it out: https://t.co/Ec67zYpStM
وش ذا اللي صار؟! 👀
واحد خلّى نموذج Kimi K3 اللي فيه 2.78 تريليون معامل يشتغل على معالج واحد بس بذاكرة 8GB.
بدون GPU. بدون أي framework. حتى بدون BLAS. بس C99 وقرص صلب.
المحرك كامل حجمه 176KB بس. النموذج يقرأ من القرص توكن توكن، والنتيجة مطابقة 100% مهما كان حجم الرام عندك.
الرام الزيادة بس تخليك أسرع. الجواب ما يتغيّر أبدًا.
يعني لابتوب عادي يقدر الحين يشغّل نفس النموذج اللي يشتغل على جهاز بـ 128GB رام. الفرق بس إنه ياخذ وقت أطول.
الـ repository: https://t.co/DlbCxyWn6p
anthropic will sell you opus 5 at $200/mo. openai will sell you gpt-5.6 at $200/mo. google will sell you gemini 3 at $250/mo. neither will tell you the same output ships on a $10 plan once you delete the graph and keep the loop
three weeks ago langgraph stars jumped 47%. the trigger was six words on twitter, "are we still talking loops or did we shift to graphs yet?" 24 hours later a manifesto. a week later every ai account had a $497 graph engineering course. all of them wrong about the same thing
the sentence they never quoted, buried in a langchain post: "loop engineering isn't an alternative to graphs, so much as a simple version of them." peter steinberger wrote a provocation, not a build order. 90% of the industry read it as a build order anyway
the graph, five silent failure modes, each one your loop didn't have:
parallelism · $0.08 a run, $240 a month for the illusion
-> 6 nodes fire in parallel and queue against the same TPM budget, executing serial inside the provider's throttler
-> real fix: check your TPM ceiling before you fan out. one long call ships in 2.9 seconds, your "parallel" ships in 12
state · 7 nodes, 75,400 tokens, $3,390/month for what a loop does in 18k
-> shared state grows with every node. by node five it weighs 8,000 tokens and every later node loads all of it
-> real fix: one dict, one loop, no per-node schema. the graph carried 4x the tokens for the same output
debug · 12 min on a loop, 113 min on the graph you migrated to
-> "why did it take that path" is one transcript on a loop and a routing table plus 5 nodes on a graph
-> real fix: look at your logs before you migrate. 4.5x longer to find the bug is not a feature
verify · 40% of graph runtime is a reviewer node that never catches a real bug
-> the reviewer rubber-stamps its own writer 90% of the time and doubles your token bill
-> real fix: move verify to a single check at loop exit. delete the reviewer node
lock-in · you now debug the framework, not the task
-> every routing bug is a framework bug, not a task bug. state leaks are the framework's abstraction leaking
-> real fix: extract each node as a plain function, wrap in one while-not-done loop, delete the framework from requirements.txt
migrate up, not down. 90% of you should still be running a loop. a graph you migrated to is a smarter employee locked in a bigger empty room
drop your $200/mo ai sub to $10, check the article below
Cut your Claude/Codex token usage with this one trick
1. Download https://t.co/sxGYmdolH6 (free & oss)
2. Drop 20$ in https://t.co/auNHh6Y6CR
3. Make a skill called delegate-wave
4. In Agents.md note: "always use delegate wave"
"Use pi in tmux with deepseek-v4-pro/deepseek-v4-flash" All read, discovery, and changes should be delegated to pi. Your role is to review and delegate"
95%+ reduction in costs
Now run Kimi K3 with a 2.78 trillion parameter model on a single CPU with 8.24 GB of RAM. 100% Opensource.
It's called kimi-k3-in-c.
Most inference stacks assume you need a GPU cluster and 5+ terabytes of memory to touch a frontier MoE.
The entire engine is 176 KB of portable C99.
Runs without BLAS, PyTorch, CUDA, any framework, a GPU, or AVX-512.
The 1.56 TB checkpoint sits on NVMe. Only 16 of 896 experts fire per token, so the sleeping 93% streams in on demand and never touches RAM.
The dense trunk runs in whatever memory you give it. 8 GB, 32 GB, 128 GB, 224 GB. Same weights. Byte-identical output at every budget.
What it does:
→ Runs Kimi K3 inference on a laptop with 8 GB free
→ Reads MXFP4 weights directly, never dequantizes to float32
→ Streams the trunk from disk with O_DIRECT, faster than the page cache
→ Keeps KDA attention state fixed regardless of context length
→ Passes bit-identical output between scalar, OpenMP, and AVX2 paths
→ Verifies against the PyTorch reference on every kernel
Ships with a 45 MB source tree, six C files and one Python packing script.
That's the whole thing.
UniFace unifies face detection, recognition, tracking, landmark analysis, gaze estimation, and anti-spoofing into a single Python library with hardware acceleration.
https://t.co/3BFAUAEJoz
Audio8 TTS Preview 0.6B ONNX INT4 is now available. A compact model with serious capability:
- Runs entirely on CPU with ONNX Runtime
- Uses only around 1 GB of memory during inference
- Multilingual TTS across 11 languages
- Zero-shot voice cloning
- No PyTorch or Transformers required
- Includes CLI, streaming, HTTP/OpenAI-compatible APIs, and voice registration Small footprint. Full voice-cloning capability.
Small footprint. Full voice-cloning capability.
Model: https://t.co/lElyitoZlx
Code: https://t.co/1sKAvqDNxY
Someone open-sourced a headless browser that runs 11x faster than Chrome and uses 9x less memory.
It’s called Lightpanda, it's written in Zig and designed specifically for AI agents and automation.
Not a Chromium fork. No Blink, no WebKit, no rendering engine.
→ 100 pages in 5s vs Chrome's 46s
→ ~9x faster, ~16x lighter
→ Drop-in replacement for Puppeteer & Playwright
→ Native MCP + agent mode built in
It has an agent mode. You describe a flow in plain English, it clicks through and pulls the data, then /save exports that session as PandaScript: plain JavaScript you can replay forever.
Deterministic. Token-free. No model at runtime.
You prototype with the LLM once, then ship the script to production and never pay another inference call on that workflow again.
100% Open Source.
Generates complete YouTube videos from a single topic by orchestrating research, scripting, voice generation, and publishing using Google Gemini and Vertex AI.
https://t.co/Md47ykC1k4
Trending repository of the day 📈
claude-video
Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
Last 24h: 989 ⭐
Total: 11,613 ⭐️
https://t.co/hJV24XK7Ba
The tiny Kimi-K3 that you can runs locally on a potato hardware.
- 2.8T to 0.18B.
- 0.10B activated.
- Same architecture.
- Same new attention design scale.
- Same DNA, smaller version
- compressed it into a 0.18B version for testing.
- now fits in 700MB.
You can actually load the new Kimi K3 architecture on normal hardware now.
🚨 Someone built a free AI assistant for any documentation — deploy it on your own server in minutes.
No Intercom AI. No Zendesk AI. No custom development. One Docker command. Your docs. Your server. Zero ongoing cost.
It's called DocsGPT. 14,900 GitHub stars. And it solves the problem every developer-facing product has been paying too much to solve.
Here's the problem.
You have documentation. Users read it and still can't find what they need. They open a support ticket. A human answers. You pay for support tooling, support staff, and a bad user experience — all because users can't navigate docs efficiently.
Every major company's answer: an AI chatbot trained on the docs. Intercom Fin: $0.99 per resolution. Zendesk AI: $50/month plus usage. Custom development: $50K+ to build and maintain.
DocsGPT: $0. Self-hosted. One command.
Here's what it actually does.
You point it at your documentation — a URL, a folder of markdown files, a PDF, a GitHub repo, a Confluence space. DocsGPT ingests it, chunks it, embeds it, and builds a retrieval-augmented chatbot that answers questions about it with citations.
Ask it anything in your docs. It finds the answer. It shows you exactly which page and section the answer came from. Users stop opening support tickets for things that are already documented.
Here's what it actually supports:
→ Documentation sources — URLs, local files, PDFs, markdown, RST, MDX, GitHub repos, Confluence, Notion, Docusaurus sites
→ Any LLM — OpenAI, Claude, Gemini, Groq, or any local model via Ollama
→ Any vector database — Faiss (default, local), Qdrant, Milvus, Weaviate, MongoDB Atlas
→ Embeddable widget — drop a JavaScript snippet into any webpage, the chatbot appears
→ Full REST API — integrate with any existing tool or workflow
→ Multi-tenant — multiple documentation sets, multiple projects, all managed from one dashboard
→ Analytics — see what users are asking, what the chatbot couldn't answer, where your docs have gaps
Here's the wildest part.
The analytics dashboard turns DocsGPT into a documentation improvement tool, not just a chatbot. Every question users ask that the chatbot couldn't answer is a gap in your documentation. DocsGPT shows you exactly what users are looking for that you haven't written yet.
You don't just get a chatbot. You get a continuous feedback loop on what your documentation is missing.
Here's why self-hosting matters beyond cost.
Your documentation often contains proprietary technical details. Internal processes. Architecture decisions. Security information. When you use a third-party AI assistant service, that content gets processed on their servers and potentially used to improve their models.
DocsGPT processes everything locally. Your documentation never leaves your infrastructure. Your users' questions never leave your infrastructure.
One command to deploy.
Open localhost:5173. Upload your first documentation source. Your AI assistant is live.
14.9K GitHub stars. 1.8K forks. 2,061 commits. MIT License.
100% Open Source.
GitHub link in the comments 👇
🚨 Hugging Face just open sourced a complete real-time voice AI pipeline.
Speak. It listens. It thinks. It talks back. End to end. Running on your GPU. Free.
No OpenAI Realtime API fees. No ElevenLabs per-character billing. No Google Cloud per-minute pricing. Just a GPU and an internet connection.
It's called speech-to-speech. Built and maintained by Hugging Face. And it does something no other open-source project has assembled cleanly until now.
Here's what makes real-time voice AI hard.
Every voice AI pipeline has four stages: speech recognition (you speak → text), language model (text → response text), text-to-speech (response text → audio), and audio output. Each stage adds latency. Chain them together naively and you get a system that feels slow — the pause between you finishing a sentence and the AI starting to respond breaks the conversational illusion.
speech-to-speech is built around minimizing that latency at every stage simultaneously.
Here's the full stack it ships with:
Speech Recognition (STR):
→ Whisper (local) — OpenAI's transcription model, runs fully offline
→ Faster-Whisper — 4x faster inference with same accuracy
→ Distil-Whisper — smallest and fastest, lowest latency
→ Paraformer — Chinese language specialist
Language Model (LLM):
→ Any Transformers-compatible model — Llama, Mistral, Qwen, anything
→ Any OpenAI-compatible API endpoint — swap in Claude, GPT, Gemini
→ MLX-optimized models for Apple Silicon — runs efficiently on Mac
Text-to-Speech (TTS):
→ Parler-TTS — controllable voice with description-based prompting
→ MeloTTS — multilingual, fast
→ ChatTTS — natural conversational prosody
→ HF Inference Endpoints — offload TTS to Hugging Face servers when needed
Mix and match. Any STR with any LLM with any TTS. Test combinations. Find the lowest latency stack for your hardware.
Here's the wildest part.
It ships with a Language Model Speech (LMS) mode — an experimental architecture where the LLM generates audio tokens directly instead of text tokens. No separate TTS stage. The model thinks in audio.
This is the architecture that makes GPT-4o Advanced Voice feel natural — the model is generating speech as a native output, not converting text to speech after the fact. HF's open-source version lets you experiment with this architecture on your own hardware.
And there's a VAD (Voice Activity Detection) system that detects when you stop speaking in real time — no fixed silence threshold, no manual push-to-talk. The pipeline responds the moment you finish a sentence.
Here's the cost comparison that makes this worth caring about.
OpenAI Realtime API: $0.06 per minute input, $0.24 per minute output. A one-hour conversation: $18. A developer building a voice AI application with 1,000 daily users: $18,000/day in API costs.
speech-to-speech on a single A100: $0. Your only cost is the GPU rental.
For production voice AI at any scale, the economics are not close.
One command to install.
9.6K GitHub stars. 944 forks. Apache 2.0 License.
100% Open Source. From Hugging Face.
GitHub link in the comments 👇
ERIC SCHMIDT, EX-CEO OF GOOGLE, AT DAVOS:
"If you really want to make money, it's actually easy. Found an agentic AI company."
The clarification most people skip. Not a company designing agents. Build an agent that actually does something.
"This is the agentic period in AI. Everyone's going to build agents. The agents are all going to compete."
Full course. Free. Bookmark this before you forget.
call me super annoying but..I will keep repeating this…
Claude + SEO is going to make more millionaires in 2026 than Wall Street has in the last decade.
don’t bookmark this if it crosses your timeline.
just paste this entire thing into Claude.
thank me later