Quando o CR chegou à Juventus, passou a querer bater todos os livres (72 livres, 1 golo)
Uma vez numa falta perto da área, o CR ficou no chão a queixar-se. O Allegri já farto mandou o médico entrar para o assistir pq assim tinha de sair e não batia o livre. O Pjanic bateu. Golo.
Apple Intelligence неизменно превращает название корабля Титаник в ТИТАНИГЕР.
Так у всех пользователей по всему интернету 😅😅😅 Проверьте у себя.
Отмена Apple через 3..2..1.
I have a Continuous Learning benchmark where models attempt to learn to play chess. They are given a /goal of learning and improving playing against a Stockfish opponent in 200 games. They can choose the difficulty, take notes, whatever they like - except cheating (e.g. using a chess engine) of course.
So far the improvement in Elo has been negative for Astra. Tiny bit positive for Opus, but could also be random. I've started Astra off sooner, so it finished its 200 games already, Opus is still playing.
Site here to watch how they are doing: https://t.co/voMG7th5MS
The reason why this is interesting is that while we obviously don't have continuous learning, at the back of my mind I was thinking that maybe models can simulate it through self-scaffolding. Turns out not so much at least in this context. Perhaps it's a solvable problem and we don't need 'true' self-learning for models to learn in some way.
-------
Just a note, the idea for the benchmarks belongs to someone else, but I don't want to use their name to give this more weight without permission.
This might be the most useful paper on AI agents this year.
A team at Wavestone AI Lab took apart Claude Code, Codex and 9 more and found the 7-part blueprint every good agent is built on.
Their definition: an agent is a model plus a harness. All 11 agents build the harness from the same 7 parts. Only the size changes.
> Loop: Mini-SWE-Agent is a plain while loop. OpenHands logs every event so a run can be replayed
> LLM layer: one prompt template on one side, 29 provider profiles on the other
> Tools: some agents only have bash, the biggest has 43 typed tools loaded on demand
> Memory: from keep-the-whole-history to notes the agent maintains across sessions
> Safety: from a simple step limit to policy rules plus a reviewer model plus an OS sandbox
> Orchestration: Aider skips sub-agents on purpose, others fork them with their own context
> Extensions: hooks, skills and MCP. SKILL.md is in 9 of 11, MCP in 8
The paper ends with a 90-line harness covering all 7 parts, a solid starting point for your own.
Key pages:
> p.4: the 7-part map
> p.12: the 3 loop types, with Claude Code and Codex as examples
> p.53: why none of them use frameworks or embeddings
> p.67: 18 design recommendations
> p.71: the 90-line harness you can copy
Uber publicó cómo armó su infraestructura de MCP y hay varias ideas muy buenas acá.
imaginate una empresa con miles de servicios internos y muchos equipos conectando agentes a esos sistemas por su cuenta.
cuando eso empieza a crecer, terminás con tooling fragmentado, infra duplicada, distintas formas de exponer tools y poca consistencia en seguridad, discovery y operación.
por eso Uber hizo un MCP Gateway: una capa central entre los agentes y sus servicios internos.
el agente habla MCP y el gateway se encarga de routing, permisos, ejecución y de traducir las llamadas a HTTP, gRPC o TChannel según corresponda.
como Uber ya tenía miles de APIs internas, también hicieron AutoCrawler: recorre sus definiciones, detecta métodos y schemas y genera MCP tools automáticamente. incluso usan un LLM para mejorar las descripciones.
hoy tienen +800 MCP servers y +5000 tools. y ahí aparece otro problema: no podés cargar todas esas definiciones en el contexto del modelo.
para eso crearon Omni MCP. en vez de pasarle miles de tools al agente, este va descubriendo qué server necesita, qué tools tiene disponibles, carga el schema de la correcta y recién ahí ejecuta.
me parece un muy buen ejemplo de cómo cambia la arquitectura cuando los agentes pasan de usar unas pocas tools a trabajar con miles.
terminás necesitando algo muy parecido a un API Gateway para agentes.
German Army Invades - April 9th
This movie features surprisingly better details on squad- and platoon-level infantry action scenes (planning, ambush, retreat, regroup etc.) than many Hollywood blockbusters