a fun thing about the bay area is that it has microtrust climates. mill valley: leave your laptop unattended when you go to the public library bathroom. civic center: do not go to the public library bathroom.
hot girls are necessary part of a thriving city ecosystem. hot girls inspire ambitions and call you to transcend. Hot girls make the world go round. A city without hot girls is like a sky without stars
We are launching AA-Video-T2V v2.0, our new benchmark for evaluating text to video models, alongside AA-Video-T2V-Silent v2.0 for video generation without audio. Built on a new methodology, it judges every model at 1080p on a regularly refreshed prompt set, and ranks them across 10 use cases, 10 capabilities and a wide range of styles.
Video models are being adopted across more industries and workflows, from film studios to advertising agencies. Our new benchmark not only ranks models overall, but also shows which model is best for specific use case and video generation capability. Use cases are grounded in how consumers and enterprises use video generation. Capabilities draw on lab and academic research, and on how creators and businesses push video models today. We tag every prompt by use case (such as Live-Action Film and Marketing & Advertising) and by the capability it tests (such as Text Rendering and Audio Synchronization), and the overall benchmark samples evenly across both. Because of this, the overall ranking reflects a model's versatility across use cases and well-roundedness across capabilities. We also tag each prompt by visual style, such as photorealistic, 3D render, cartoon and anime, and hand-drawn illustration.
AI video is also moving onto bigger screens and into production, from microdramas to movie theaters, while low barrier to generate is resulting in a proliferation of low quality AI video content. The quality bar keeps rising, so we now judge every clip at 1080p and high bitrate.
We are launching AA-Video-T2V v2.0 with more than 68,000 high quality human preference votes from private evaluators based in US/UK over 1,000 prompts, and AA-Video-T2V-Silent v2.0 with more than 47,000 votes over 500 prompts.
Initial insights from an in-depth analysis of the 10 highest ranking models on the Artificial Analysis AA-Video-T2V v2.0 Leaderboard:
➤ Wan 3.0 ranks #1 overall and leads 10 of the 20 category boards, including Cartoon and Anime style and Animation & Gaming use case, at $12 per minute of video.
➤ Dreamina Seedance 2.5 ranks #2 and is the human performance specialist, #1 on both Human Anatomy and Dialogue & Lip Sync. At $34.12/min it is the most expensive model in the top 10.
➤ MiniMax H3 (768p) ranks #3, statistically tied with Seedance 2.5 at $4.80/min, about 1/7 of the price. It is also #1 on Text Rendering.
➤ FLUX 3 ranks #4, with its strongest results on Text Rendering (#3) and Dialogue & Lip Sync (#2).
➤ Gemini Omni Flash 1.1 ranks #5 and is the graphic 2D and audio specialist, #1 on UI/UX & Motion Design use case, Flat Design style and Audio Synchronization capability.
See below for the use case, capability and style breakdowns 🧵
Announcing AA-AgentPerf-Local, our open-source inference testing tool for local AI models - test how fast agentic AI can run on your own laptop or workstation, and browse our list of serving configurations to plan your next agent setup
Key points:
➤ We’re open sourcing AA-AgentPerf-Local, which replays real agent trajectories on laptop & workstation hardware to test inference performance
➤ We’re releasing initial results for NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro
➤ The tool and leaderboard will soon expand to cover more hardware, frameworks and models, and will stay updated over time as new releases launch
We already benchmark inference performance on mobile phones and datacenter-scale hardware, and are expanding our coverage to include laptops and workstations. Alongside hosting a leaderboard displaying results from popular model and hardware combinations, we are open-sourcing all code and data required to run AA-AgentPerf-Local to ensure that individuals and companies are able to run their own trials, informing their local AI serving decisions.
AA-AgentPerf-Local replays real agent sessions. Our default workload is 8 recorded agentic tasks, spanning 168 model turns. Each request carries the full conversation so far, as a real agent's would, so context grows to ~56K tokens. Every turn generates exactly its recorded number of tokens, so every system does identical work. Tool execution is skipped by default to isolate inference speed, but when benchmarking your own system, you can also replay the real recorded tool delays or run tool calls live on your CPU.
At launch, we are focusing on the performance a single agent can achieve when able to utilize the entire system; we plan to expand this coverage over time to cover multi-agent systems and scenarios where an agent must run alongside other regular processes.
The initial set of hardware covered on our official page is: NVIDIA DGX Spark (128 GB), AMD Ryzen AI Halo (128 GB), MacBook Pro M5 Pro (64 GB), and NVIDIA GeForce RTX 5090 (32 GB). These have been selected to cover a range of platforms (CUDA, ROCm, Vulkan, Metal), memory capacities, and bandwidths. We will be expanding the featured hardware to include x86 (and other) laptops, AI-focused graphics cards such as the RTX PRO 6000 Blackwell, and more, to give consumers a well-rounded view of performance across different hardware types.
The models featured at launch are: Qwen3.5-9B, Qwen3.8-27B, Qwen3.6-35B-A3B, and Ling 3.0 Flash (124B / 5B active). These initial models span a range of memory requirements and dense/MoE architectures, and are each benchmarked at 4-bit quantizations to reflect realistic serving conditions. The core set of models we feature will shift over time as new open weights models are released. Beyond the featured models, AA-AgentPerf-Local is able to benchmark performance of any OpenAI-compatible inference server the user runs, meaning that any model and config can be tested locally.
Choice of serving configuration is important, with the runtime, quantization and speculative decoding changing results substantially. Where an official off-the-shelf config was available for a system and model pair, we used the published config. We developed our own configs for all other cases. Every config uses speculative decoding (MTP, DFlash or DSpark), and all 14 are published in the repo and on the configs page on our website.
Initial results:
➤ Completion time mapped most closely to each model’s active parameter count: Qwen3.6-35B-A3B (3B active) was the fastest model on every system, e.g. 2.5-3.3x faster than the dense Qwen3.8-27B. However, active parameters are not the whole story, with Ling 3.0 Flash (124B total, 5B active) still finishing behind Qwen3.5-9B (nearly double the active parameter count) on all hardware that can support it.
➤ The GeForce RTX 5090 was the fastest system for every model that fits in its 32 GB, achieving completion times >3.5x faster than the other systems. Single-user decoding is heavily influenced by memory bandwidth, and the GeForce RTX 5090 has 1,792 GB/s against 256-307 GB/s for the unified-memory systems.
➤ The DGX Spark and Ryzen AI Halo are overall similar systems, with the same amount of unified memory, comparable memory bandwidth, and the same launch MSRP of $4,000. On our default trajectory set, the DGX Spark was 1.4-1.7x faster on three of four models, with the Ryzen AI Halo tying it on Qwen3.5-9B. The gap is far larger than their 7% bandwidth difference - the Spark has greater low-precision compute than the Ryzen AI Halo, enabling it to prefill faster, and its more mature CUDA software likely plays a role in enabling better MoE and speculative decoding performance.
➤ The MacBook Pro (M5 Pro, 64 GB, 20-core GPU) is the only laptop tested so far, and exhibited competitive results, finishing within 2–9% of the Ryzen AI Halo on Qwen3.6-35B-A3B and Qwen3.8-27B (though 21% slower on Qwen3.5-9B). It has the most memory bandwidth of the three unified-memory systems (307 GB/s) and its current price of $3,700 is the lowest of the systems tested so far. Its results are likely held back by software maturity and compute available for prefill.
➤ Despite the agentic trajectories serving 73-93% of prompt tokens from the KV cache, simply reading each turn’s new input (during prefill) used up substantial proportions of the end-to-end completion time, e.g. 22-41% for Qwen3.8-27B. This was especially impactful on systems with low compute FLOP/s relative to their memory bandwidth, such as the Ryzen AI Halo and potentially the MacBook Pro (MacBook FLOP/s are unpublished).
➤ Speculative decoding was implemented on all of the most successful configs so far, e.g. raising Qwen3.8-27B decode speeds ~30-120% above the bandwidth-constraint roofline.
palo alto people are building agents for a world where everyone has 47 tasks at once while the average person is so fucking bored they’re doomscrolling during work hours
It’s so annoying how we apparently need to take societal direction and cues from people who don’t go out and do anything normal, so they’re all like “what if you could do this thing that’s already possible… but now we’re middlemen and you paid us $40 billion dollars for that…?”