SpaceX has launched the company's Super Heavy-Starship on its first flight to orbit, marking a major step toward perfecting the world's most powerful rocket for commercial flights and NASA moon missions. https://t.co/nOrumXt7q2
Bookmark this tweet.
Starship will never work.
It's all a scam on stupid people.
Uneducated space enthusiasts mostly.
Nobody does iterative design on full size prototypes for a reason.
Especially when it's the world's biggest rocket.
Complete dumbfucks are in charge.
Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense
The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best.
We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners.
Benchmarks in the Artificial Analysis Cyber Index:
➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust.
➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized.
➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces.
Key results:
➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44).
➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.
Deleted Muse after seeing this post on Threads about how it told some Facebook Marketplace sellers the guy’s address and they showed up at his door
Dangerous and creepy
This would have been 1000x worse if the person was a woman
Starship’s 14th flight is set to launch on Monday, Sept 28. The 75-minute launch window opens at 7:15 a.m. CT. Live coverage of the mission starts ~35 minutes before launch → https://t.co/uQKQvgaTmJ
@TomTotesProfesh@Bitcoin_Devs I know you're just commenting on the slides themselves, but trying to wrap my head around what would actually prevent the inflation... im assuming it's the zkproof?
i.e. sum of inputs = sum of outputs ?
SpaceX’s Starship flight 14 is launching 26 V3 Starlinks and represents ~26 Tbps for the launch (based on company specs)… at a very conservative cost estimate for Starship this puts Starships first operational launch at ~$9.8M per Tbps. Meaning in its first operational launch it nearly hits the same cost of Falcon9’s launch cost per Tbps. This is before full payload utilization, reusability and manufacturing learning rates kick in, which we expect to lead to 16x cheaper launch/Tbps cost by 2028.