On the average day, agents account for over 70% of @tempo's documentation traffic.
Instead of endlessly changing button colors, we built a benchmark to understand how they really see things.
Introducing stable-bench-v1 - part of a broad suite of stable coin integration evals authored by Tempo
https://t.co/NyvBaXwlQJ
We’re open sourcing WANDR.
WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer.
https://t.co/gp2BWjFK4d
When we first started Harbor-Index a few months ago, we thought: how hard could it be?
1,311 tasks passed our difficulty filter. What's left is just sampling a few hard tasks across benchmarks right?
...until an LLM audit flagged 1,004 of them as broken.
For the next few months, we filtered out broken tasks with a team of 14 reviewers and reran frontier models on our candidate set 4 times to catch & patch task issues/reward hacks. Harbor-Index v1 is just the start.
Benchmarks are software.
Software needs CI/CD.
So should benchmarks.
Open-source software thrives on transparency and community validation. Benchmarks should too. Issues and PRs open: https://t.co/0y2kXsCAFB
📊 Your Android Bench July update:
1. Added 8 new models, check out the top of the leaderboard!
2. You can now contribute to the benchmark.
3. We standardized our benchmark by transitioning to the @harborframework.
Read about what's new → https://t.co/Wncpx5NQc1
Introducing Harbor-Index, a compact, diverse, and high-quality benchmark built to challenge frontier agents.
We carefully select, audit and fix 82 high-signal tasks out of 6,627 candidates spanning 54 benchmarks.
No agent gets above 30%. (1/5)
And last update: Senior SWE-bench is now also on Harbor Hub (@harborframework)! You can run `harbor run -d snorkel-ai/senior-swe-bench-v2026.06` out of the box.
Full blog post: https://t.co/DyvJ0djOKI
Harbor is really great. I like the design and it's well polished for doing evals. It would be great to use the same rollout utility for everything (RL / eval / new tasks definitions).
@jackhau0212@harborframework i’ve said it before and i’ll say it again
the tam of harbor could encompass almost all knowledge work
it’s a beautiful abstraction
I’ve been using @harborframework for a while for evals, and I’m a huge fan. I love the design and the core principles behind how an environment is modelled.
The thing that gets me most excited is that good evals unlock so much more than benchmarking. They become the foundation for hill-climbing, auto-research, RL, GEPA, trajectory analysis, SFT data generation, and more.
Rollouts, rollouts, rollouts.
Rollouts for eval, rollouts for RL, rollouts for GEPA, rollouts for prod, rollouts for trajectory analysis, rollouts for SFT data gen, rollouts rollouts rollouts
@nummanali We want Harbor to be the rollout primitive that powers every optimization loop. Eventually we will also natively implement common optimization loops on top of Harbor primitives. But I’d encourage you to cook on the use case you described and use Harbor to power it!
We get this question a lot: "Which model is best for drug discovery?"
Our new benchmark announced today with @ScaleAILabs, DrugDiscoveryBench (82 tasks from working drug discovery scientists, run on Biomni Open Source Environment), has a clear answer: the model matters far less than what you build around it.
🧵3 key takeaways →
Introducing SWE-Together: a multi-turn benchmark built from real user–agent coding sessions.
Coding agents are often benchmarked like exam-takers: given the full spec up front, then graded on the final code. But real coding help is a conversation — users clarify goals, add constraints, and correct course along the way.
SWE-Together turns real coding work into a reproducible, verifiable benchmark: 109 repo-level tasks curated from 11,260 recorded sessions, replayed with a reactive LLM user simulator that preserves the original user’s intent.
We evaluate agents as collaborators, not just patch generators: final pass rate and how many user interventions were needed to get there.
In this evaluation snapshot, claude-opus-4.8 currently leads among the 7 agents we tested — achieving the highest pass rate while requiring the fewest user interventions.
📄 Paper: https://t.co/Zp5BSPpLTJ
💻 Code: https://t.co/NPgxCMLdHi
🌐 Website: https://t.co/BK50zRGReE