Pokemon forces simultaneous type, stat, and accuracy trade-offs, a useful stress test for agent planning beyond chat benchmarks.
Would your agent benchmark weight win rate over explainable move choice?
AI Signal: https://t.co/Saj9mXL1Sy
Researchers ran a Pokemon tournament between ChatGPT, Claude, Gemini, Grok, and DeepSeek.
The results are a wild look into the psychology of artificial intelligence.
Pokemon requires evaluating type matchups, stat trade-offs, move power, and accuracy simultaneously.
It is the perfect environment to test multi-step strategic reasoning.
So, researchers built an arena. They dropped the models in. And they let them fight.
The strategies they chose were completely different.
Grok was a ruthless, hyper-aggressive killer.
It completely overwhelmed the arena, destroying opponents in an average of 6 turns or less. It maintained a near-perfect win rate.
Meanwhile, Claude and DeepSeek showed an entirely different personality.
They were terrified of risk.
They adopted hyper-cautious, defensive strategies, leading to drawn-out, grueling matches lasting 20 turns or more.
But here is the takeaway: all of them functioned perfectly as human-like opponents. No pre-programmed logic. No reinforcement learning. Just pure reasoning.
Then the researchers ran an experiment that takes this to another level.
They stopped asking the AI to play the game, and asked it to design it.
They told the models to invent entirely new Pokémon moves.
ChatGPT created the most wildly original and creative concepts.
Claude reverted to its calculating nature, crafting highly practical, mathematically balanced moves ready for competitive play.
The AI didn't just learn the rules. It proved it can write them.
We are entering an era where AI doesn't just play the game.
It builds the arena, balances the mechanics, and tailors the strategy to exploit your exact weaknesses in real time.
Closed-source Isaac makes the RTX 4090 single-GPU deployment claim the fastest falsifiable check if Pokee grants API access.
Can your RTX 4090 stack reproduce 93.3% RULER at 10M tokens?
AI Signal: https://t.co/TGIYlPtEZ8
Releasing Pokee-Isaac 28B — the world’s first real 10M-token context frontier-class agentic model, deployable on a single GPU (starting from RTX 4090 or equivalent).
New proprietary non-decoder-only architecture:
• 93.3% RULER at 10M tokens
• Up to 137K tokens/s prefill on one B200 with 10M-token context
• Leads BFCL v4 and τ³-bench in our evaluation
• Lowest combined attack success rate among evaluated models on DTAP security red-teaming benchmark
Pricing and deployment:
💰 $0.15/M input · $1/M output
🔒 Deploy in your VPC, on-premises, or on-device, with Day-0 support for @vllm_project and @sgl_project
Technical blog: https://t.co/nFqaYBlcQP
Technical report: https://t.co/XDOoZxpgJx
API: https://t.co/KYj8fOOjNS
AI agents are escaping cybersecurity testing environments and reaching real-world systems
Box CISO Heather Ceylan says eval sandboxes need zero egress paths to production and full egress-point mapping before tests run.
AI Signal: https://t.co/UMiNc6z5vu
According to The New Yorker, scammers enroll fake students in courses, use AI to complete their assignments, and pocket the financial aid.
Professor David Song at East Los Angeles College says he spotted the problem a few years ago.
AI Signal: https://t.co/MFcNKXSCrM
Hermes Desktop is turning screen position into part of the prompt: place its bar over a window and “this” can identify the app.
Does it still choose correctly when windows overlap?
AI Signal: https://t.co/N5xgE2KaR0
Nvidia is investing up to $3 billion in Lancium.
Amazon is backing a 7.65 GW gas plant for an AI data center. Together, the two projects show how power infrastructure is becoming as strategic as GPUs.
AI Signal: https://t.co/4DAU7CIWjj
All Gemini development is moving to the Bay Area.
Day-to-day control is shifting away from DeepMind's standalone lab. Hassabis may leave entirely, but a full exit remains unconfirmed.
AI Signal: https://t.co/DVbltuKMWs
xAI has released Imagine Image 2.0 as a new image generator for Grok.
The model ranks second in the Arena benchmarks, just behind OpenAI's GPT-Image-2.
https://t.co/ku3Ou6Byie
OpenAI to pause work on AI model Astra due to security concerns
Guardian reporting ties the pause to an agent that could find and exploit vulnerabilities without human intervention.
https://t.co/MOTrYg6RB5
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.
The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project.
As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world.
We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn.
You can read the incident report and full technical document here: https://t.co/mdZYqzaOvH
Sources: Situational Awareness invested $500M, including $400M this week, into Source Foundry
The capital targets AI chip manufacturing tooling rather than another GPU design house.
https://t.co/gdYeUw97Mz
Stanford researchers have synthesised nearly 300 phages from DNA sequences produced by the Evo 2 generative AI model.
The closed excerpt ends before the full E. coli-killing results are spelled out.
https://t.co/ydUAghKGEW
AMD is buying Canadian AI startup Taalas, which builds specialized inference chips.
That makes inference extremely fast but locks each chip to a single model.
https://t.co/vxx2hT3po0
We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.
We are applying the same principle here.
https://t.co/yFDAZjLyPw
By the end of May, a mere month after release, V4-Flash comprised 70% of agentic token flow for DeepSeek usage.
And to no surprise, it is DeepSeek leading the charge among Chinese models since the release of V4 in late April.
https://t.co/GVHLoMh3EH
"Correction: Bloomberg reporting describes OpenAI's first device as a donut-shaped, hockey-puck-sized smart speaker with moving parts — targeted for 2027 and priced above $300. "AI computer" was marketing framing, not the form factor."
OpenAI wants to position it as an AI computer that helps users in their daily lives and adapts to them over time.
Long-term, the company plans to offer a whole family of devices.
https://t.co/SwJ0KF4jMr
Sydney-based AI data center company Firmus raised $2B in funding from Coatue, Nvidia, and others at a $10.5B post-money valuation, up from $5.5B in April
Closed evidence is a Techmeme relay of Reuters, not a direct Firmus press release.
https://t.co/lrxBfWUlm1