This virtuality allows people who belong to minority groups to attend one of the most important conferences. Thank you so much @_LXAI for your grant, and all @CVPR Organizing Committee for your big efforts @RanaHanocka. Awesome advice by @gulvarol@SamuelAlbanie Mathias Niessner.
It's finally here!
Minimax H3 Turbo lora makes generations 5x faster. Use only 4 steps instead of 20.
https://t.co/fUbVjZbpwg
Also see: https://t.co/OmqHO5vBED
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
Incredible open-weight model! The gap between open and closed AI has nearly vanished!
Interested in the tech behind it? Here are some background topics.
Kimi Delta Attention (KDA): https://t.co/V66jlOjfYn
Attention Residual (AttnRes):
https://t.co/2mWgHbiSxv
Mixture of Experts (MoE):
https://t.co/X6GPShwMAw
Today's Hermes Agent Masterclass is the finale! Module 10 covers security, an important topic for running agents to ensure each agent can perform its tasks while minimizing exposure. Hermes has a ton of built-in security features. In this clip, I show how you can set approvals for each profile depending on your needs.
UN CIENTÍFICO DANÉS PROGRAMÓ A CLAUDE PARA QUE BUSQUE TRABAJO POR ÉL Y LO ACABA DE HACER PÚBLICO
Mandar CVs es uno de los trabajos más absurdos del mundo: copiar, pegar, adaptar, personalizar, repetir. Todo manual, todo lento, todo para que lo lea un algoritmo antes que un humano.
→ Analiza la oferta de trabajo automáticamente
→ Genera un CV personalizado para cada puesto
→ Redacta la carta de presentación adaptada al contexto
→ Todo lo hace Claude por debajo, sin que toques nada
→ Open source, ya en 3.5k stars en GitHub
El tío que debería estar buscando trabajo ha construido la herramienta que lo busca por él.
Aquí te explico cómo funciona 👇(repoo al final del hilo)
Here's part 1 (of 5) of my short course on efficient LLM inference that I taught at Columbia University. Slides are heavily updated from two weeks ago.
https://t.co/WVCf7mUdkY
This Fall at CMU we're teaching a new course on AI Agents!
The goal is that you learn how to create a scaffold, build evals, and train an agentic LLM using RL.
We'll try to balance theory and practice, and introduce modern frameworks and best practices.
Ex-Google engineer explained AI agent loops, harness, evals in 20 minutes - better than 500$ courses.
trace every run → judge it with an LLM → diagnose → fix → ship.
That loop is how agents self-improve over time.
Agent loops + memory + harness + evals - thats the stack.
Watch it, then save the framework below.
KARPATHY JUST KILLED THE PROMPT ERA WITH A SINGLE DOCUMENT
prompts are easy. loops are hard. and writing fifty prompts a day is the work nobody does twice.
he shifts the burden to the harness.
you define the contract once. the model writes, reviews, restarts, and reconciles. you keep judgment. it keeps the loop.
the throughline is the same in every rule: the human owns the spec and the boundary. the model owns the execution and the bookkeeping.
planner never touches code. generator never grades itself. state lives on disk, not in context.
9 rules. start with one feature, not ten. most people are still typing prompts. this turns Claude into an agent that finishes the job on its own.
here is the official document from Karpathy explaining the architecture
Had a great time talking about 𝗔𝗜 𝗔𝗴𝗲𝗻𝘁 𝗛𝗮𝗿𝗻𝗲𝘀𝘀 𝗘𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝗶𝗻𝗴 at Large Scale Production Engineering, Google.
I gave an example of building a Financial Agent to analyse my credit card spends using Gemma 4 family of Local LLMs.
Everything running under 15 GB RAM on my laptop!
The agent was performing quite well & in some cases, at par with the frontier models.
This was a clear testimony for the importance of Harness Engineering to build effecting Agents.
Linking the YouTube talk in the following Tweet. Will link the repo soon.
It took me a while to wrap my head around linear attention and friends.
But when it finally clicked, it was incredibly satisfying!
I finally understand the essence of what it does, why it works, and its connection to past and modern approaches.
https://t.co/geNiBXKLbg
Si la IA hace la tarea de tu hijo, el cerebro de tu hijo no participó en eso. Y ese uso de la IA es un problema. Un estudio de Stromberg, Lei y Wu (https://t.co/8QHcVapdup) siguió a 26,000 estudiantes de secundaria en China durante 30 meses. Muestran que el uso de IA elevó las notas de las tareas en 18% y redujo el tiempo que dedicaban a esas tareas en 30%. Al mismo tiempo, las notas en los exámenes a libro cerrado cayeron 20% luego de seis meses. Analizando un efecto de largo plazo, las notas en los exámenes de ingreso a la universidad cayeron hasta 24%. Como se ve en la figura, mas eficiencia para hacer las tareas, pero menos menos aprendizaje. Por otro lado, los estudiantes que usaron IA pero igualmente dedicaron el mismo tiempo que sus compañeros sin IA obtuvieron en los exámenes notas casi idénticas a las de quienes no usaron IA. Pero los que "externalizaron" su tarea completamente, terminando más rápido y posiblemente poniendo menos esfuerzo, que cualquier estudiante sin IA, solo sacaron notas altas en los ejercicios. La diferencia fue el esfuerzo cognitivo: si el cerebro hizo el trabajo, o simplemente observó cómo la IA lo hacía en su lugar.
La IA es una herramienta. Un bisturí en manos de un cirujano salva vidas. El mismo bisturí en las manos equivocadas hace daño. La misma IA que puede ser un tutor efectivo, desafiando al estudiante a pensar, explicar, luchar productivamente con el problema, se vuelve dañina si elimina el esfuerzo que el aprendizaje requiere. Esto no es grave si le pasa a pocos estudinates. Pero cerca del 80% de los estudiantes que usaban IA cayeron en lo que el estudio llama "tercerización de tareas" (homework outsourcing).
Esto conecta directamente con la clasificación que hacíamos con mis colegas Ezequiel Molina y Maria Barron, (https://t.co/ygsaXKPVfR) de tres grupos de estudiantes, los Empoderados por la IA, que la usan para pensar con más profundidad; los Dependientes de la IA, que la usan para evitar pensar; y los Excluidos de la IA, que no tienen acceso. Nos preocupaba que el grupo de los Dependientes pudiera crecer. Ya creció. Es, con mucho, el grupo más numeroso en este estudio. Completar una tarea no es lo mismo que aprender. El cerebro no construye conocimiento observando cómo trabaja la IA sino equivocándose, esforzándose y resolviendo las cosas por sí mismo. Ese esfuerzo cognitivo es esencial para el aprendizaje .
This "loop" automation is nuts inside of Codex.
"/goal go over every single feature in this app create a user story with expected behaviour based on the code keep a single canonical spreadsheet tracking the features status
- when done switch loop to testing every user story and documenting all errors
- when done fix every logistical error or ux error
- test every user behaviour again post fix"
Shoutout to @MatthewBerman for the heads up.
Hundreds of user stories being worked through like it's nothing.
this is actually really cool, fully open source. RL framework behind GLM-5.2, GLM-5.1, GLM-5, GLM-4.7, GLM-4.6, and GLM-4.5, validating the full post-training loop.