Evaluación de agentes ≠ evaluación de respuestas. Si solo haces single‑turn tests, estás viendo escenas aisladas; los problemas reales aparecen en el loop: llamadas a herramientas, estado mutable y adaptaciones entre pasos. Cambia el paradigma de pruebas.
Google sta ridefinendo il concetto stesso di visibilità
Prima bastava posizionarsi nella SERP.
Ora serve “apparire” anche in AI Overviews.
Ed ecco che il mercato si adatta: prima si vendeva per link building, oggi vendono post su siti che compaiono in AI Overviews.#newseo#google
Today is the start of a new era of natively multimodal AI innovation.
Today, we’re introducing the first Llama 4 models: Llama 4 Scout and Llama 4 Maverick — our most advanced models yet and the best in their class for multimodality.
Llama 4 Scout
• 17B-active-parameter model with 16 experts.
• Industry-leading context window of 10M tokens.
• Outperforms Gemma 3, Gemini 2.0 Flash-Lite and Mistral 3.1 across a broad range of widely accepted benchmarks.
Llama 4 Maverick
• 17B-active-parameter model with 128 experts.
• Best-in-class image grounding with the ability to align user prompts with relevant visual concepts and anchor model responses to regions in the image.
• Outperforms GPT-4o and Gemini 2.0 Flash across a broad range of widely accepted benchmarks.
• Achieves comparable results to DeepSeek v3 on reasoning and coding — at half the active parameters.
• Unparalleled performance-to-cost ratio with a chat version scoring ELO of 1417 on LMArena.
These models are our best yet thanks to distillation from Llama 4 Behemoth, our most powerful model yet. Llama 4 Behemoth is still in training and is currently seeing results that outperform GPT-4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on STEM-focused benchmarks. We’re excited to share more details about it even while it’s still in flight.
Read more about the first Llama 4 models, including training and benchmarks ➡️ https://t.co/9G3QgVdCkB
Download Llama 4 ➡️ https://t.co/eVomRvEr0w
Sembra che MCP stia diventando lo standard per i principali LLMs. Dopo OpenAI sembra che anche Google stia considerando l’adozione di questo protocollo che ricordo essere stato sviluppato da Anthropic
Ora è ufficiale: OpenAI Agents SDK supporta il protocollo MCP! Interazioni dirette e sicure con file locali, DB e API.
Dettagli: https://t.co/bHLIMRPq98
#OpenAI#MCP#AI#AgentsSDK
people love MCP and we are excited to add support across our products.
available today in the agents SDK and support for chatgpt desktop app + responses api coming soon!
Analizzando i dati comparativi, Google presenta un nuovo modello di ragionamento impressionante. Dai numeri condivisi, questo si posiziona davanti a o3-mini High e leggermente indietro rispetto al non ancora rilasciato o3
Quel 18,8% ottenuto su Humanity's Last Exam è davvero wow
Atlas is demonstrating reinforcement learning policies developed using a motion capture suit. This demonstration was developed in partnership with Boston Dynamics and @rai_inst.
Atlas is demonstrating reinforcement learning policies developed using a motion capture suit. This demonstration was developed in partnership with Boston Dynamics and @rai_inst.
ok we heard y’all.
*plus tier will get 100 o3-mini queries per DAY (!)
*we will bring operator to plus tier as soon as we can
*our next agent will launch with availability in the plus tier
enjoy 😊
Today OpenAI announced o3, its next-gen reasoning model. We've worked with OpenAI to test it on ARC-AGI, and we believe it represents a significant breakthrough in getting AI to adapt to novel tasks.
It scores 75.7% on the semi-private eval in low-compute mode (for $20 per task in compute ) and 87.5% in high-compute mode (thousands of $ per task). It's very expensive, but it's not just brute -- these capabilities are new territory and they demand serious scientific attention.
Excited to introduce https://t.co/vgeLv0E05t
I took @aleyda Solis' SEO spreadsheet dashboard and converted it into a dynamic, interactive tool for e-commerce. Check out how it can enhance your SEO strategy! #SEO#Ecommerce#Dashboard#seotools