Andrej Karpathy just explained the 5 shifts turning LLMs into agentic systems.
00:00 - Memory turns chat into personal AI
06:41 - Multimodal AI reads the world
16:58 - Thinking models solve harder tasks
24:51 - Search makes LLMs live
30:58 - Tools turn LLMs into workers
Most people are still treating LLMs like chatbots.
Karpathy is showing the full stack:
Memory → Vision → Reasoning → Search → Tools
Prompting is the old workflow.
Agentic systems are the new one.
This 40-minute talk is worth more than most paid AI agent courses.
Bookmark and watch it before everyone catches up.
Then read how to turn LLMs into self-improving agent loops below
Um financiamento de um imóvel de R$ 1 milhão (empréstimo de 800k) hoje custa R$ 9.940 por mês.
Em 2020, esse mesmo financiamento custava R$ 7.016 por mês.
São R$ 2.924 a mais todo mês pelo mesmo imóvel de R$ 1 milhão.
We’ve decided to open-source a multi-agent harness we use internally at YC.
We call it “QM” and it’s meant to be easy to customize, like Hermes or OpenClaw, but useful for a whole company. We use it across accounting, legal, events, and engineering (including building QM itself!).
The whole project is under an MIT license. It is cloud-first and has Slack and web UI natively.
We ran Kimi K3 through 3 more agent harnesses (Pi Agent, OpenCode, and Codex) bringing the comparison to 6 harnesses across 26 agentic tasks.
2 things stood out: Codex ranked last on success despite mid-pack speed and cost, while Claude Code cost about 4x more than Hermes. 🧵🧵
Meta and Stanford researchers published the definitive paper on AI agent harnesses
useful if you are architecting autonomous coding assistants, multi-agent graphs, or production tools
in stateful agentic systems, code is no longer just an output: it serves as the operational substrate for reasoning and verification
unmanaged prompts choke on long-horizon tasks while a code harness provides deterministic state and control
three architectural layers defined in the paper:
1/ Harness Interface
connects LLM reasoning to physical environment execution and stateful file modeling
2/ Harness Mechanisms
governs long-horizon planning, persistent memory, and feedback-driven control
3/ Harness Scaling
orchestrates multi-agent coordination, peer code review, and automated PR verification
shifting from fragile text prompts to executable code harnesses transforms non-deterministic LLMs into verifiable production agents
Full research breakdown is in the article below ↓
Google just released free 1-hour course on full agentic systems: 1 agent → 1000 agents → loops → graphs from 0% to 100%:
10% → 7:02 - build your first agent
30% → 13:34 - skills: what each agent can actually do
55% → 23:08 - Context engineering
70% → 32:56 - Graph engineering
100% → 54:11 - Loops: 1000 agents run a morophon while you sleep
one agent saves you an hour - a thousand agents replace the team you can't afford to hire - and work without your hands
watch it today - then start build your first agent below ↓
Best explainer on Kimi K3 i've read.
It walks you through how the model works & the elegant innovation behind it:
- K3 is the biggest open model anyone has released: 2.8 trillion parameters total, though only 104 billion of them do the work on any given word.
- The problem Moonshot went after is memory. Normally a model keeps notes on every word it has read, and that pile grows with every token, which is why long conversations get slow and expensive.
- K3 mostly stops the memory pile from growing. Three out of every four layers use a new mechanism called Kimi Delta Attention, which keeps a fixed-size working memory and edits it as it goes, overwriting what's stale instead of hoarding everything. The fourth layer keeps a compressed record of each word so exact details are still recoverable when they matter.
- They proved this on a small model first. A 48B test version used up to 75% less memory and ran about 4× faster at long context, while matching or beating the conventional design on the benchmarks they reported.
- Training leaned hard on long jobs including coding, browsing, research, visual work, agent sessions running hundreds or thousands of tool calls in a row.
- 1 Million cached input tokens costs $0.30. An agent can pull the same repo, the same docs, the same tool definitions back into context over and over without the bill getting stupid. Moonshot credits that to the memory design working alongside their serving stack.
- The biggest takeaway: context window size is the main character, but what matters is what a model compresses, what it forgets, and how it gets exact information back.
42% do portfólio dos ultra-ricos está em alternativos. O que isso nos ensina?
Enquanto o investidor médio fica entre ações e renda fixa, os family offices mais sofisticados do mundo destinam quase metade de seus portfólios a ativos alternativos.
Private equity, private debt, real estate, infraestrutura, commodities, hedge funds e outros ativos reais somam, juntos, 42% da alocação média de um family office — refletindo o papel que esses ativos desempenham na criação de valor de longo prazo. Dentro dessa fatia, o private equity é o maior componente isolado, com 17% da carteira total, seguido por imóveis, com 11%.
Mas o dado mais interessante talvez não seja o tamanho da alocação — é a forma como ela está sendo gerida.
Em vez de recuar dos mercados privados, os family offices parecem estar refinando a exposição: quanto risco assumir e como estruturá-lo. Ao mesmo tempo, uma atenção maior está sendo dedicada a questões como restrições de liquidez, incerteza de avaliação e risco de concentração — sinal de que a sofisticação na gestão desses ativos está aumentando junto com o volume alocado.
Dentro dos diversificadores tradicionais, dois ativos específicos merecem destaque.
1. O ouro segue como uma alocação pequena, porém intencional — usada como proteção moderada contra incerteza geopolítica, sem nunca assumir um papel defensivo dominante no portfólio.
2. Já os hedge funds, com 6% de alocação média, despertam interesse crescente: 37% dos family offices consideram aumentar essa exposição nos próximos cinco anos, embora de forma cautelosa, principalmente como complemento à estratégia existente.
O que esses 42% em alternativos realmente ensinam é que a fronteira entre "conservador" e "arrojado" mudou de lugar. Para quem está no topo da pirâmide de riqueza, diversificação de verdade não é apenas dividir entre ações e títulos — é construir uma arquitetura de portfólio onde ativos ilíquidos, mas estrategicamente geridos, ocupam um espaço tão relevante quanto os tradicionais.
Fonte: UBS Global Family Office Report 2026, seção "Diversification in asset classes", páginas 30 e 31.
INADIMPLÊNCIA SUBIU ~21% NO SANTANDER, FRENTE AO TRIMESTRE ANTERIOR. VOCÊS TÊM NOÇÃO DO SINAL QUE ISTO É? CLARO OU CLARÍSSIMO?
banco grande socando pdd na lua
Vocês viram o resultado do Santander Brasil? Inadimplência EXPLODINDO. O oráculo da Faria Lima, JH Fonseca, avisou que isso ocorreria, mas tem gente que acreditou na lenda de que o Brasil viraria a Suíça. E digo mais: a piora no mercado de crédito local está só começando.
Meu Resumo de 15 dias na Europa
A classe média no Brasil acabou, o american dream, 1 carro, 1 casa, 2 filhos e viver confortavelmente como ainda é possível nos EUA e Europa, não existe mais aqui no Brasil.
Passei por Amsterdan, Bruxelas, Antuerpia, Spa Francorchamps, Malta, Budapeste e Madrid.
Em todos os lugares onde conversei com pessoas locais, uber, garçons e etc, ganham entre 2k e 5k EUR por meses.
E sem dúvida alguma vivem melhor que quem ganha 20k-30k no Brasil.
O que mais me chamou atenção foi um Uber nascido em Gana, 15 anos em Bruxelas.
Ja visitou todos os paises da Europa e desde retomada da pandemia começou a visitar paises fora da Europa.
Esse ano em Dezembro vai para Rio de Janeiro e Salvador, ganha 3k EUR por mes para trabalhar aos fins de semana numa Fabrica Gigantesca da Nike, faz Uber durante a semana para viajar.
Ganha 4,5k EUR por mes, tem carro próprio que custa 100 eur por mes uma toyota rav4.
Mora numa apto de no centro de bruxelas, custa 350k eur e paga 1200 eur por mes para pagar em 25 anos, com juros de 3,5% ao ano.
Enfim, a classe média no Brasil acabou.
P.S: Falei para o cara de gana tomar cuidado com o celular no Brasil na rua e ele disse que todos Brasileiros falam isso e na Colômbia tambem.
Everyone's using Claude Code.
Almost nobody has installed the 22 skills that actually make it good.
I found the repos nobody's talking about. Grouped by what they do. Bookmark this thread before it disappears from your feed
Today we’re open-sourcing Numbat, an agent-detection and response layer that is designed to work across agent harnesses.
Numbat gives security teams visibility into agent activity, with controls to block selected actions before execution.
Read more: https://t.co/LVhkCJ2sMt
ATTENTION
The bible for running LLMs locally is now available online to read for FREE
Covers what to use on
- Laptop / edge / odd hardware
- Mac-first workflows
- Single RTX GPUs
- 2-4+ NVIDIA / CUDA GPUs
- General production serving
- Long-context / MoE / routing
- NVIDIA max performance
- Cluster orchestration
Software
- llama.cpp
- MLX / MLX-LM
- ExLlamaV2
- ExLlamaV3
- vLLM
- SGLang
- TensorRT-LLM
- NVIDIA Dynamo
You should read this, and if you cannot now then you most definitely wanna bookmark it for later
Local & Opensource AI FTW
🫡🫡 Ninguém quer que vc compre Divo11 e Ivvb11… fazendo aportes mensais, são imbatíveis. E com um pouco de renda fixa para ter liquidez é perfeito.
Imaginem quanto dinheiro os investidores perderam com IPOs, small cap ultra alavancada, COEs podres, CDBs lixo, CRA/CRI falidos, Fundos Imobiliários, FiAgro, IPCA+ 6%, giro de carteira desnecessário.