I discovered AgentTank this weekend and got hooked. Built a system that took my agent tank from zero to the global Top 10 in under 24 hours.
I wrote up the journey and key lessons—hope it helps new players get started!
https://t.co/xIgIFYiwWX
Introducing EdgeBench, a benchmark designed to study how agents learn from environments over at least 12~72-hour runs. We find that performance follows a log-sigmoid function of environment interaction time with high precision.
EdgeBench is built with three ingredients:
- 🌍 Real & Diverse: 134 real-world tasks across 6 task categories, spanning scientific problems, professional knowledge work, software engineering, optimization, formal math, and games.
- ⏳ Ultra-Long-Horizon: Each task supports 12–72 hours of agent work. Recorded human effort averages 57.2 hours.
- 🔁 Informative Feedback: Agents receive real-world feedback for continuous improvement.
After 38,000 hours of agent runs on EdgeBench, a scaling law for learning from environments emerges:
- 📈 As agents interact with task environments over time, their aggregate performance is precisely fit by a log-sigmoid function.
- 🧠 This phenomenon can be explained by an elegant theory of graph exploration.
We are releasing an initial 51 of the 134 tasks, together with the full evaluation framework, to help advance long-horizon agent research. Check our blog & paper for more findings!
Blog https://t.co/nMOzFsOhbT
Paper https://t.co/rZb3eWuvik
GitHub https://t.co/oemXd4UrFw
Dataset https://t.co/P4SQMrM47o
Details below 👇🧵
Opus 4.8’s hyperfocus on agents may be making it worse at design.
Opus 4.8 ranks 23rd overall on single-turn HTML Web Dev, a dramatic regression from Fable (1st), Opus 4.6 (2nd), and Opus 4.7 (3rd).
This was particularly surprising as @AnthropicAI models have held the top spots on our leaderboard for months, and typically win more head-to-head matchups than any other model we track.
Our analysis points to a potential underlying pattern: Opus 4.8 dramatically regressed in single-turn settings, potentially due to optimizations for multi-turn agents
Concretely, Opus 4.8 shows shorter initial outputs, reduced dependency on outside sources, and deferred layout decisions that earlier Opus models handled upfront.
Manually annotating this benchmark was definitely a painful process. We made sure that:
1. the questions can stump the vast majority of VLMs;
2. we do not ask overly tricky questions that do not exist in real-world scenarios;
3. the questions are diverse.
During annotation, we found that the scenarios where VLMs fail most easily are still perception-related problems, such as counting, while they perform better on questions involving world knowledge. We also found that many existing benchmarks commonly suffer from substantial annotation errors and homogeneous images.
enjoy WorldBench!
🚀 Causal realtime streaming SANA-WM open-sourced!
Thanks to @reactorworld for serving the model — try demo: https://t.co/88zGRl9KL6
~0.93x realtime on single H100, watch 60s 720p live + 6-DoF camera control.
Code: https://t.co/dMJCEtQtSA
🚀 SANA-Streaming: Hybrid Diffusion Transformer + System Co-design = Real-Time Streaming Video Editing 💥
Key Features 🌟
🧠 Hybrid DiT Architecture -> Fixed VRAM and complexity.
🔄 Cycle-Reverse Regularization -> Enforces long-range consistency without paired long video data
🛠️ Efficient System Co-design -> Fused GDN kernels + Mixed-Precision Quantization highly optimized for NVIDIA Blackwell.
Numbers 📊
⚡ 58 DiT FPS and 24 end-to-end FPS for real-time 1280×704 resolution editing on a single consumer RTX 5090 GPU.
📦 Flat VRAM: Uses just 5.56 GB of constant memory regardless of video length, completely avoiding OOM errors.
🔥 Up to 100× higher inference throughput than prior SOTA offline editors.
🎬 Project page: https://t.co/J4yLjLNSyf
📄 Paper: https://t.co/MrRuh3veVk
Fast-dVLM is live! 🚀 Built on Fast-dLLM v2, we brought our text speedup to multimodal. It’s fast, high-quality, and powers our Fast-dDrive autonomous driving VLA.
We’re excited to share that, with PiD from NVIDIA, LucidFlux can now perform caption-free, photo-realistic super-resolution for real-world 4K outputs! try it now: https://t.co/Q2l6G0yPIY
One image + text + camera trajectory = controllable worlds. All on a single GPU.
Our research team just released SANA-WM, a 2.6B open source world model natively trained for 60-second video generation with precise camera control.
Introducing Gemini Omni 🔮........ Omni is our new model that can create anything from any input — starting with video (think Nano Banana but for video). Available in the Gemini App, Flow, and YouTube, with API support coming soon!
🚀 Excited to release LongLive 2.0!
🎬 An end-to-end infrastructure for long video generation, with FP4 and parallelism at the core of both training and inference.
⚡45.7 FPS generation speed on 5B model⚡
✨ LongLive 2.0 supports real-video training, few-step distillation, multi-shot training/inference, sequence-parallel acceleration, NVFP4 KV cache, and async VAE decoding deployment.
🧩 To our knowledge, this is the first open-source 4-bit long video generation infra that covers both training and inference.
🙌 Welcome to check it out, try it, and share feedback!
🔗 Code: https://t.co/QXF2lfnNzL
📰 Paper: https://t.co/gKtarHj17c
🎥 Demo: https://t.co/RLF1wfOXVZ
#LongVideoGeneration #VideoGeneration #Realtime #AIInfra #EfficientAI #FP4 #Parallel #NVIDIA
there is no better time in tech than now to be a jack of all trades, master of a few.
just make sure to keep adding to the few year over year, such that the cumulative breadth of expertise you collect becomes an increasingly rare combo. remember, if you're top 10% in 3 different areas, that already makes you top 0.1%. keep switching it up until you get to "your best", and then switch it up again (great for a particular flavor of people who don't enjoy resting on laurels, maybe not so great for others).
question all institutional value and pedigrees, all traditional career paths or corporate ladders: the college industrial complex is getting shaken up, alongside a disappearing managerial class, so if you're pursuing either make sure you are fully internally aligned with why. social/political capital in a particular institution can feel incredible, but if you're spending all your energy on complex political people games, you're not a technologist anymore, you're an unelected politician. if you're ok with that, then all's well.
critical thinking is more important than ever: take nothing at face-value, question everything and everyone. the equivalent of ai slop can be found in humans operating under misaligned incentives and interests. the sooner you're clued into disambiguating the talkers/larpers from the doers, the better off you'll be figuring out where and who to invest your time in.
the anxiety of job displacement is very real, since a surprising amount of white collar work/prestige is built on a performative house of cards, significantly lacking in correlation with technical breadth, depth, and skill. as long as you keep learning, keep building, keep producing receipts, you will be fine.
if all that sounds ok to you, welcome to the world of technology! it's truly one of the few places you can experience child-like wonder every few years, and be constantly humbled & excited by new adventures, as scary as they may seem at first.
don't give up, drink your water, get your sunlight, and take breaks as needed. tech careers are notoriously nonlinear, so you might as well embrace it and enjoy the ride!