Breaking: Browser Use is faster than a human⚡
Running Qwen 3.8 27B on 2x B200s with DFlash2.
> speed ✅
> cost ✅
> accuracy ✅ (most of the time)
Qwen 27B is better at using the internet than you are. The only bottleneck is website loading.
Comment if you want us to deploy this to our cloud 👀
La IA es el mayor destructor de supuestos y de marcos mentales. No recomiendo acomodarse mucho en eso de "la IA no puede".
Ejemplo en el campo de las matemáticas 👇
Introducing Bot Mode for Hermes Desktop.
Your agent profiles become a series of named Bots. Each Bot has its own role, model, memory, skills and profile picture; Bots can use any model and even communicate with each other.
Build a specialist Bot once to use it forever.
agree excepto con los modelos de Google
3.1 Pro sigue siendo SOTA en muchas cosas, solo que son una mierda en agentic coding
pero para OCR, world knowledge, long context... los modelos de Gemini son top top
mira que modelos son los primeros en cualquier benchmark de OCR...
Se han hecho un dataset con 43 atributos por jugador con datos de Opta, luego un PCA de 7 componentes para reducir dimensionalidad captando el 90 y pico por ciento de la varianza, para acabar tirando de un K-means (elbow method mediante) y ponerse a buscar a los jugadores que aparecen en el cluster de peak De Bruyne intentando encontrar su reemplazo ideal siguiendo una metodología data driven, y encima lo graban.
Tengo 65/66 años de vida, pero mi papá tiene 90 años y sobrevivió a su esposa que quiso y murió hace como diez años mucho más joven que él, hace dos años tenía novia, tengo sus genes y los médicos dicen que mi capacidad física es de jóven.
Quizás puedo comprar una casa junto al mar Caribe de Colombia, visitaré muchos lugares hermosos del mundo pero siempre volveré a Colombia y sólo estaré con la mujer que me ame a la que me dedicaré a hacer feliz después de dedicar la mayor parte de mi tiempo al pueblo y escribiré muchas ideas al país.
Del largo listado de mujeres amantes que me asignan en la prensa ninguna está conmigo.
Las destruyen gratuitamente porque no estaré con quien no me ame y seré fiel porque ya he visto demasiadas mujeres hermosas y el corazón de un verdadero guerrero es de una.
Siento que cumplí bien mi deber con honestidad y templaza y dignidad, es hora de la tranquilidad y del amor.
Empieza.una nueva etapa de mi vida donde seré mas sabio y más fuerte de lo que he sido.
Cuando el pueblo me llame estaré a su servicio con todo mi valor y mi franqueza.
No hemos perdido nada, hoy sale la verdad, y hemos ganado y somos más, los que mienten siempre pierden.
Estoy preparado a cualquier reto porque nos levantamos como el ave fénix y ganamos de nuevo porque amamos de verdad.
¡OpenAI paraliza temporalmente el entrenamiento de su próximo modelo de IA más avanzado!
Dicen que es tan bueno atacando sistemas que no pueden entrenarlo sin supervisión.
¿Será verdad… o una excusa por falta de recursos?
-"Claude por qué le enviaste esa propuesta al cliente si es 80% menor al precio con lo que solemos trabajar???? Vamos a quebrar!"
-"You're absolutely right!"
🚨 AI agents are getting seriously good at doing AI research.
Prime Intellect ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.
The best run, using Fable 5, closed 82% of the gap to a human record that was built by dozens of people over months.
And it did it autonomously.
This is one of the clearest examples yet of frontier models starting to meaningfully automate AI research itself.
DeepSeek V4 Pro 0813 scores 53 on the Artificial Analysis Intelligence Index, 8 points above April's DeepSeek V4 Pro - but with a 3.6x price increase and only 1 point above DeepSeek V4 Flash 0731
@deepseek_ai has released DeepSeek V4 Pro 0813, its new flagship model, along with updated pricing. With the new pricing, it still sits on our Pareto frontier for Intelligence vs. Cost, but by a smaller margin than previous DeepSeek releases.
Under the new pricing, DeepSeek’s first-party API will charge $1.32 / 1M input tokens and $3.96 / 1M output tokens - an increase of +264% on the blended price. Cached input tokens receive a ~97% discount (rather than the previous 99%): cache hits are priced at $0.044 / 1M tokens, 12x higher than previous cache-hit pricing. Off-peak pricing is discounted by 50%. This new pricing takes effect on August 16, 2026 - until then DeepSeek is offering V4 Pro 0813 at the same pricing as the earlier V4 Pro model.
DeepSeek published the model weights under the MIT license, making it the second most intelligent open weights model we have benchmarked; however, weights for Qwen3.8 2.4T A95B have recently been released and will be evaluated on the Intelligence Index soon. DeepSeek V4 Pro 0813 remains 7 points behind Kimi K3, the current open weights leader.
DeepSeek V4 Pro 0813 retains the previous version’s architecture at 1.6T total parameters and 49B active parameters, with a context window of 1M tokens. Most of DeepSeek V4 Pro 0813’s gains over the April version are in agentic capabilities. DeepSeek V4 Pro 0813 is more token efficient than its predecessor, generating ~30% fewer total output tokens across Artificial Analysis Intelligence Index.
Key results:
➤ DeepSeek V4 Pro 0813 is only 1 point ahead of DeepSeek V4 Flash 0731 on the Artificial Analysis Intelligence Index, and has ~3.8x the active parameters. The two models tie on Terminal-Bench 2.1 at 79%. For agentic and coding benchmarks the smaller model performed at around the same level.
➤ DeepSeek V4 Pro 0813 is more token efficient than DeepSeek V4 Pro, using 128M output tokens to run the Intelligence Index, ~30% fewer than the April version. This places DeepSeek V4 Pro 0813 on the open weights Pareto frontier for intelligence versus output tokens, with fewer tokens used than Kimi K3 (133M) and GLM-5.2 (141M). However, it lags behind many frontier models in its intelligence tier. GPT-5.6 Sol scores 61 on the Intelligence Index but only uses 70M tokens in total.
➤ Cost per Task increases significantly despite an improvement in token usage, driven by DeepSeek’s price increase. At $0.25, DeepSeek V4 Pro 0813’s Cost per Task is 5x that of DeepSeek V4 Pro ($0.05), but ~20% lower than GLM-5.2 ($0.32 Cost per Task). Despite the increase in price, DeepSeek V4 Pro 0813 still sits on the Intelligence vs. Cost per Task Pareto frontier, but barely - it is just one cent below Gemini 3.7 Flash (medium) on Cost per Task.
➤ DeepSeek V4 Pro 0813 makes significant improvements in real-world agentic knowledge work. On GDPval-AA v2 its Elo improved to 1590, +284 points from DeepSeek V4 Pro (1306), which implies an ~84% expected win rate against the April release. Its GDPval-AA v2 Elo is ahead of earlier-generation proprietary models like Claude Opus 4.7 and GPT-5.5, but it trails Kimi K3. DeepSeek V4 Pro 0813 works for longer on GDPval tasks than before, averaging 39 turns per GDPval-AA task compared to 23 for the April model.
➤ DeepSeek V4 Pro 0813 makes a 12 point improvement in AA-Omniscience, driven entirely by increases in accuracy. DeepSeek V4 Pro 0813 scores +1, up from -11 for DeepSeek V4 Pro. Its AA-Omniscience accuracy rises 6 points from 43% to 49%, while regressing slightly in hallucination rate from 94% to 95%. Compared to peer models, DeepSeek V4 Pro 0813 has a high attempt rate of ~99%, meaning it almost never declines to answer. By contrast, Kimi K3 attempts only 77% of questions and has a 53% hallucination rate.
Additional model details:
➤ Context window: 1M tokens
➤ Weights: 1.6T total parameters; 49B active parameters
➤ License: MIT
➤ Accessibility: Accessible through DeepSeek first-party API at launch
➤ Pricing: On DeepSeek’s first-party API, updated pricing will take effect on August 16 at $1.32 / 1M input tokens and $3.96 / 1M output tokens. Cached input tokens receive a 97% discount, priced at $0.044 / 1M tokens. Off-peak pricing is discounted by 50%.
Gemini 3.7 flash == GLM 5.2 price and quality
the google model released this week is essentially matching price and quality perf of GLM 5.2
GLM 5.3 is going to be open weights pretty soon!
@ErickSky Eso es pura mierda jajaja. Si no como ellos siempre se la quieren dar de únicos y originales, ahora salen con marca de agua en sus textos, argumentando que es dizque cumplimiento de normas cuando hasta se le alzaron a la misma casa blanca
Anthropic need to drop Fable 5.1 badly
Opus 5 feels trash and Fable 5 is nearly tied with Grok 4.6 which costs almost 4x less per task
Grok 4.7 also comes out in a couple weeks which will decimate it
Crazy week in AI