Some observations on Kimi:
1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run.
2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China.
3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex.
4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business.
5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this.
6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.
🚨| CHINA ACABÓ CON EL NEGOCIO DE OCR
Un modelo de 3 mil millones de parámetros del tamaño de un cacahuete puede leer un PDF completo de 100 páginas de una sola vez.
Sin dividir páginas.
Sin pérdida de contexto.
Sin factura en la nube.
Conoce Unlimited-OCR 👇
• Lee documentos completos con una ventana de contexto de 32K
• 93% en benchmarks estándar de análisis OCR (+6 sobre la línea base)
• La tasa de error se mantiene por debajo de 0.11 incluso después de 40+ páginas
• Multilingüe de fábrica
• Se ejecuta 100% localmente en tu hardware
• Soporta Transformers, vLLM, SGLang, Docker, Ollama y llama.cpp
Aquí viene lo loco:
La mayoría de las herramientas OCR todavía procesan los documentos página por página.
Unlimited-OCR lee el documento completo como un contexto único, preservando tablas, referencias, diseños y relaciones entre páginas.
Eso lo cambia todo.
Mientras tanto, las empresas están pagando:
• $1.50–$15 por 1.000 páginas
• Enviando PDFs sensibles a proveedores de nube
• Esperando respuestas de API
Este modelo lo hace offline.
Gratis.
Para siempre.
Construido por Baidu para ir más allá de DeepSeek-OCR.
Ya tiene más de 1,9 millones de descargas en Hugging Face...
...y casi nadie está hablando de ello todavía.
El código abierto se mueve más rápido que la mayoría del software empresarial.
MOONSHOT JUST CLONED CLAUDE CODE AND MADE IT FREE
it's called Kimi Code CLI.
open source.
MIT licensed.
built by the team behind Kimi K3.
and it already does things Claude Code doesn't.
→ drop a screen recording directly into chat
→ built-in planner, coder and explorer agents
→ plan mode before touching a single file
→ automatic MCP configuration
→ works across VS Code, JetBrains and Zed
→ one binary. starts in milliseconds.
the CLI costs nothing.
K3 starts at around $3 per million tokens.
worth using before everyone switches.
As an AI Engineer. Please learn
>Harness engineering, not just prompt engineering
>Context engineering, not just long prompts
>Prompt caching vs. semantic caching tradeoffs
>KV cache management, eviction, reuse, and memory pressure at scale
>Prefill vs. decode latency and why they optimize differently
>Continuous batching, paged attention, and throughput optimization
>Speculative decoding vs. quantization vs. distillation tradeoffs
>INT8, INT4, FP8, AWQ, GPTQ, and when quantization hurts quality
>Structured output failures, schema validation, repair loops, and fallback chains
>Function calling reliability, tool contracts, argument validation, and idempotency
>Agent guardrails, loop budgets, tool budgets, and termination conditions
>Model routing, graceful fallback logic, and degraded-mode UX
>RAG architecture: chunking, embeddings, hybrid search, reranking, and freshness
>Retrieval evals: recall, precision, grounding, attribution, and citation quality
>Evals: golden sets, regression tests, adversarial tests, LLM-as-judge, and human evals
>LLM observability as a first-class discipline: traces, spans, tokens, latency, errors, and drift
>Cost attribution per feature, workflow, tenant, and user journey not just per model
>Safety engineering: prompt injection defense, data leakage prevention, and permission boundaries
>Multi-tenant isolation, cache safety, and cross-user context contamination prevention
>Fine-tuning vs. in-context learning vs. RAG vs. distillation and when each is the wrong tool
>Latency, quality, cost, and reliability tradeoffs across the full inference stack
>Production failure modes: hallucinated tool calls, malformed JSON, stale retrieval, runaway agents, and silent eval regressions
Este proyecto 3D costó $0.09 con Kimi K3, frente a $0.11 con Claude Opus 4.8.
La diferencia de precio es pequeña, pero el resultado demuestra algo importante: un modelo abierto ya puede competir de tú a tú con los mejores modelos cerrados.
3 accounts banned for no reason. One was banned just for changing the credit card on my account. Really?
The support and appeal process takes months only to get a stupid AI-generated response
Fuck you @AnthropicAI You’re no longer indispensable
Welcome @OpenAI and @Kimi_Moonshot.
🚀 Un ingeniero senior de Anthropic que gana $1.2M al año fue preguntado cómo logra entregar trabajo al ritmo de un equipo completo… trabajando solo.
En lugar de explicar, compartió su carpeta .claude/ completa.
Mismo modelo. Resultados radicalmente diferentes.
La mayoría sigue eligiendo entre Opus y Sonnet como si el modelo fuera el límite. No lo es.
La verdadera ventaja no está en qué modelo usas, sino en qué despiertas dentro del modelo.
Este es el stack que usan los que operan a otro nivel:
• CLAUDE.md → El contrato maestro
• settings.json → Permisos y configuración
• hooks/ → Los reflejos automáticos
• agents/verifier → El policía que revisa todo
• skills/ → 33+ memorias musculares reutilizables
• .mcp.json → Las herramientas conectadas
• MEMORY.md → El registro vivo de cada turno
Ya no chateas con el modelo. Escribes la carpeta una sola vez y la carpeta ejecuta el modelo por ti.
Este es el sistema que separa a los que “usan Claude” de los que lo tienen trabajando como un equipo de élite.
Desglose completo + estructura exacta lista para copiar y pegar 👇
Guarda este post antes de que desaparezca.
Este proyecto Open Source es BRUTAL: Meetily.
Transcribe y resume tus reuniones con IA.
¡Pero lo hace todo en local! Sin subir tu audio a ningún servidor, totalmente gratis, 100% privado, en segundos.
→ https://t.co/RjZ08yP5y1
Use Fable 5 as orchestrator and Opus + Codex to execute (to save fable usage):
Fable 5 (max reasoning) = orchestrator
Opus = deep reasoning subagent
Sonnet = mechanical work subagent
Codex = peer Sr. engineer, different perspective
Setup:
1. Set Fable 5 as your main model In Claude Code: /model → Fable 5 → reasoning /effort to max
2. Create 2 subagents with /agents In Claude Code:
deep-reasoner → pinned to opus "Use for reasoning-heavy phases, architecture, debugging complex issues, algorithm design. Think thoroughly, return a concise conclusion the orchestrator can act on."
fast-worker → pinned to sonnet "Use for mechanical tasks, boilerplate, tests, formatting, simple edits. Execute efficiently."
3. Add OpenAI's official Codex plugin (install codex cli in your computer first), In Claude Code type:
/plugin marketplace add openai/codex-plugin-cc
/plugin install codex@openai-codex
/codex:setup
4. Drop this in your CLAUDE.md in your folder:
## Orchestration workflow
You (Fable) are the orchestrator. Plan, decompose, synthesize.
Reasoning-heavy phases → deep-reasoner
Mechanical work → fast-worker
Codex (/codex:rescue --background) is a cracked engineer on par with deep-reasoner, from a different perspective. Treat as a peer, not a reviewer.
High-stakes decisions: task Opus + Codex on the same problem in parallel, synthesize the best of both, without showing either the other's answer. Keep your own context lean.
5. Then prompt Fable like a tech lead: "Goal: [what you want] Context: [files, constraints] You're the lead. Delegate reasoning to deep-reasoner, grunt work to fast-worker, fresh-perspective problems to Codex. Show me your plan first, then execute."
That's it.
Web Designers and Developers, this might be the next-level animation library
ReactBits gives you a growing collection of smooth, modern UI animations you can plug straight into your projects.
Bookmark it for later 💜
Anthropic pays $750,000+ a year for engineers who can build LLM architectures from scratch. Stanford taught the entire thing in 1 hour lecture & released it for free.
Bookmark & watch this today before someone takes it down and read this article below
¡Aprovecha Claude Fable al máximo con esto!
Este comando me parece una idea brutal.
/improve usa el modelo más caro para analizar proyecto y preparar un plan de mejora.
Y luego lo implementas con un modelo más barato.
Se instala así:
LIST OF 40 WEBSITES TO FIND REMOTE JOBS
1. Linkedin. com
2. Indeed. com
3. Glassdoor. com
4. FlexJobs. com
5. weworkremotely. com
6. Remote. com
7. Upwork. com
8. Freelancer. com
9. Fiverr. com
10. Guru. com
11. Toptal. com
12. AngelList. com
13. Hubstafftalent. com
14. Simplyhired. com
15. Remotive. com
16. Virtualvocations. com
17. workingnomads. com
18. Hired. com
19. cloudpeeps. com
20. taskrabbit. com
21. talent. com
22. Remote OK - remoteok. io
23. DRemote - dremote. io
24. Jooble - jooble. org
25. stackoverflow. com/jobs
26. jobspresso. com
27. onlinejobs. ph
28. simplyhired. com
29. themuse. com
30. skipthedrive. com
31. zirtual. com
32. justremote. com
33. hireable. com
34. remoteworkhub. com
35. jobbatical. com
36. freelancewritinggigs. com
37. contentwritingjobs. com
38. problogger. com/jobs
39. behance. net
40. designhill. com
Don't forget to follow @NextShopiaAi to get more insightful Ai related tools and update.