OCR is dead 🪦 Long live agentic OCR!
Traditional OCR makes one pass and hands back whatever text it got. Tables get flattened, charts disappear, and multi-column layouts come out scrambled. Nothing checks the output.
Agentic OCR treats parsing as a loop instead:
✅ layout-aware reading order
✅ smart routing of hard elements to the right model
✅ multi-pass verification and self-correction
✅ multimodal parsing of charts, images, and complex tables
Logan Markewich, LlamaIndex's Head of Open Source, wrote about what that shift looks like in practice and where it's still struggling. Link to the breakdown in the comments below.
¡La nueva generación de tests para web y apps móviles!
El framework se llama "e2e" y es de código abierto
✓ Mezcla IA con pruebas deterministas
✓ Compatible con OpenAI, Claude y cualquier AI
✓ La misma API para web, iOS y Android
→ https://t.co/n09ldmz3PD
¡Si tienes un servidor propio, necesitas esta guía!
Tiene 137 recetas para proteger tu máquina y aplicación
✓ Autenticación, sesiones y contraseñas
✓ Docker, GraphQL y GitHub Actions
✓ .NET, Java, PHP, SQL...
Oficial de OWASP:
→ https://t.co/zT7u8TaqeT
Updated my AI orchestrator tier list, based on daily use.
T3 Code Nightly → SS. The latest updates put it ahead of everything else I've tried. Smooth, intuitive, and easy to follow.
Traycer → A. The last two updates added useful features, but I'm running into agents that stall and need repeated nudges to keep going.
Maestri → B. Still a great tool, but the workflow is built around the terminal, and I'm moving away from that.
Synara → C. Recent updates have made it noticeably better.
My priorities are changing: a clear interface, less babysitting, and a workflow that keeps moving.
What would you rank differently?
Uber publicó cómo armó su infraestructura de MCP y hay varias ideas muy buenas acá.
imaginate una empresa con miles de servicios internos y muchos equipos conectando agentes a esos sistemas por su cuenta.
cuando eso empieza a crecer, terminás con tooling fragmentado, infra duplicada, distintas formas de exponer tools y poca consistencia en seguridad, discovery y operación.
por eso Uber hizo un MCP Gateway: una capa central entre los agentes y sus servicios internos.
el agente habla MCP y el gateway se encarga de routing, permisos, ejecución y de traducir las llamadas a HTTP, gRPC o TChannel según corresponda.
como Uber ya tenía miles de APIs internas, también hicieron AutoCrawler: recorre sus definiciones, detecta métodos y schemas y genera MCP tools automáticamente. incluso usan un LLM para mejorar las descripciones.
hoy tienen +800 MCP servers y +5000 tools. y ahí aparece otro problema: no podés cargar todas esas definiciones en el contexto del modelo.
para eso crearon Omni MCP. en vez de pasarle miles de tools al agente, este va descubriendo qué server necesita, qué tools tiene disponibles, carga el schema de la correcta y recién ahí ejecuta.
me parece un muy buen ejemplo de cómo cambia la arquitectura cuando los agentes pasan de usar unas pocas tools a trabajar con miles.
terminás necesitando algo muy parecido a un API Gateway para agentes.
¡Nueva alternativa abierta y compatible de Jev!
✓ Con soporte para leer imágenes
✓ El doble de rápido y de contexto...
✓ ...pero casi 6 veces más caro (0.24$ vs 0.04$)
Creado por Cloudflare de pesos abiertos:
https://t.co/BXEyKzUOst
¡Protege tus repos públicos de GitHub con esto!
Un comando que mejora su seguridad en 2 minutos.
Es oficial de GitHub Security y activa:
✓ Protección de ramas
✓ Auditoría de dependencias
✓ Escaneo vulnerabilidades y secretos
https://t.co/ElCsx2XsEE
I've been training OCR models for 2 years. The progress we've made in handwriting is astonishing - from barely readable to almost perfect transcription.
Listado de IP's de ataques DDoS detectadas automáticamente po ell firewall Cloudflare WAF de https://t.co/E5fbK0UVq5
* Lista de IP's de Cloudflare WAF (actualizada automáticamente cada 4 minutos, IP's de las últimas 24 horas): https://t.co/HWmffMDSog
* Cloudflare WAF detecta eventos del firewall de Cloudflare utilizando la API GraphQL. Incluye ataques DDoS en capa 7, límite de peticiones (Rate Limit) y eventos de WAF con reglas personalizadas (excluyendo direcciones IP de Tor, requiriendo 10 o más peticiones, y ordenadas por número de hits, máximo 10K IP's)
https://t.co/JrAS8QHjdd
Cloudflare Containers now start 6x faster, let your agent choose each sandbox's image and instance type at runtime, and support filesystem snapshots in public beta, all controlled from a Durable Object. https://t.co/mdKedS60sU #BirthdayWeek
¿Buscas una alternativa de código abierto a Jev?
Laya hace lo mismo y puedes entrenarlo con tus datos.
✓ Open Source + Corre en local
✓ Clasifica con score, noul y choices
✓ Versión en inglés y otra multilenguaje
→ https://t.co/nKDGpFrGqJ
Anthropic acaba de publicar una guía en Español de cómo sacar el máximo partido a Opus 5.5
① effort: medium por defecto
② No digas "piensa mucho", ajusta effort
③ Di qué NO quieres, sobre todo en frontend
Aquí la guía completa:
https://t.co/TDVD4Q5hSz
EmDash 1.0 is a stable, open source CMS built for Astro, with agent-friendly workflows, secure sandboxed plugins, and a decentralized registry that keeps publishers in control. https://t.co/vHzz4Teyy9 #BirthdayWeek
Servidor MCP te permite desarrollar y automatizar dispositivos móviles desde tu IA.
✓ Compatible con emuladores de iOS y Android
✓ También con dispositivos reales (free tier)
Para Claude Code, Codex, Gemini...
De código abierto:
https://t.co/EbGHZFkpiU
🆕 Guía de Laya.cpp (Alternativa a JEV)
✅ Modelo de IA muy pequeñito (< 1GB)
✅ Posibilidad de uso en local
✅ Laya multilingual (incluido español)
✅ Rendimiento: ~230ms (Jev) vs ~33ms (Laya)
✅ Open Source (y gratis)
https://t.co/S3aqWMvZIW
Deploying DiffusionGemma-Jev (djev) just got a lot easier. You can now spin up a Jev API-compatible endpoint on Google Cloud Run using a single command.
Performance is solid: ~35-60 ms for single step latency and batch@32 is ~100-123 requests/sec.
It's a straightforward way to experiment without needing your own GPU. Runs at roughly $3/hr and drops to $0 when idle.
Get the code and instructions here: https://t.co/E3agzrW2hZ
When several people talk at once, a transcript can get messy fast.
Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
Google ha publicado AX, un orquestador de agentes de IA programado en Go.
La idea: un agente trabaja horas, guarda archivos y llama a un modelo. Si nadie lo vigila, te puede fundir el dinero en bucle.
Defines la tarea en un YAML, lo ejecuta en un sandbox y le pone límites de consumo y procesamiento.
→ https://t.co/56A4Le44aF