Top Tweets for #TerminalBench
Benchmarking my coding harness on Union Alpha — Terminal-Bench 2.0 via Harbor, model locked, verifiers only.
Free stealth window on @OpenRouter. Using it while it lasts.
#UnionAlpha #OpenRouter #TerminalBench #AIAgents #CodingAgents
🚀 Released last week, our open-source AtomSh agent already edges out Cursor CLI + Claude 4 Sonnet on Terminal-Bench!
Check out the results on AtomGPTLab's JARVIS-Leaderboard 👇
https://t.co/J5aWO3jwyS
https://t.co/QolLF8dQBe
#AIAgents #OpenSource #TerminalBench #AtomGPT

Excited to be part of Terminal-Bench-Science 0.1 🚀🎉
Proud to see my task accepted into the official release!
Happy to contribute to this amazing open-source project. 🙌
#OpenSource #TerminalBench
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains.
An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%.
1/n 👇

GLM-5.3 is out and it's already making waves
It just beat Fable and the brand-new DeepSeek V4-Pro 0813 (released yesterday) on Terminal-Bench.
#GLM53 #AI #OpenSource #MachineLearning #LLM #DeepSeek #Fable #TerminalBench

I wrote a deep dive on GLM-5.3
https://t.co/mQCt0HNlnX
#GLM53 #AI #OpenSource #MachineLearning #Coding
Marlin hitting 67% on Terminal-Bench
Sentinel: 80%
Terminus: 80%
Pushing reliable AI agents closer to real-world terminal execution and reasoning.
#AI #LLM #TerminalBench #Rlfh
Honored to have played a part in establishing a fully open foundation for OpenThoughts-Agent-v1. I'm looking forward to watching this strong start evolve and contribute further to improve future #TerminalBench agents! #OpenThoughtsAgent
How can we make a better TerminalBench agent?
Today, we are announcing the OpenThoughts-Agent project.
OpenThoughts-Agent v1 is the first TerminalBench agent trained on fully open curated SFT and RL environments.
OpenThinker-Agent-v1 is the strongest model of its size on TerminalBench, and sets a new bar on our newly released OpenThoughts-TB-Dev benchmark. (1/n)

Proud to be part of the community behind #TerminalBench 2.0 — a benchmark of realistic terminal-based tasks for evaluating agentic systems.
@LaudeInstitute @Stanford

Today, we’re announcing the next chapter of Terminal-Bench with two releases:
1. Harbor, a new package for running sandboxed agent rollouts at scale
2. Terminal-Bench 2.0, a harder version of Terminal-Bench with increased verification

From isolated snippets → full workflows. Proud to partner with @Stanford and @LaudeInstitute on #TerminalBench 2.0 — helping redefine agent evaluation. Thanks @Mike_A_Merrill and @alexgshaw for leading the charge and allowing our researchers the opportunity to contribute.

Last Seen Hashtags on Sotwe
Most Popular Users

Elon Musk 
@elonmusk
241.7M followers

Barack Obama 
@barackobama
119M followers

Cristiano Ronaldo 
@cristiano
114.4M followers

Donald J. Trump 
@realdonaldtrump
111.9M followers

Narendra Modi 
@narendramodi
107.2M followers

Rihanna 
@rihanna
98.7M followers

NASA 
@nasa
92.4M followers

Justin Bieber 
@justinbieber
91.8M followers

KATY PERRY 
@katyperry
90M followers

Taylor Swift 
@taylorswift13
83.9M followers

Lady Gaga 
@ladygaga
75.4M followers

Virat Kohli 
@imvkohli
73.3M followers

Kim Kardashian 
@kimkardashian
70.9M followers

YouTube 
@youtube
68.8M followers

Neymar Jr 
@neymarjr
66.3M followers

Bill Gates 
@billgates
65.1M followers

Selena Gomez 
@selenagomez
63M followers

The Ellen Show
@theellenshow
62.3M followers

CNN 
@cnn
61.8M followers

X 
@x
60.7M followers










