@TimJayas Haha, all these models are benchmaxxed, making many of these benchmarks fake. Fable 5 is much better. Also, Opus 5, even though it scores so well in the benchmarks sucks for real world use.
@GregKara6 Yeah - every question that can be innovative triggers the fall back. I asked a question whether a propeller done in a certain way would make flying cars more efficient. It also triggered a safety warning. Unbelievable.
@Math_files Upon closer inspection, I see a typo.
Specifically:
In the term near the end of the pure bosonic/Higgs sector:
-g^1 s_w^2 A_mu A_mu phi^+ phi^-
it should be
-g^2 s_w^2 A_mu A_mu phi^+ phi^-
Squared photon charge scalar.
@adenshepard@bridgemindai There are already capable open weight models performing near opus 4.8 levels. By the end of the year these will be mythos level or beyond. GPT and Fable already have a lot of anti-abuse security while the open models do not. The current ban is causing more harm than good.
@bridgemindai It's also painfully slow and I spend 80% of the time fixing bugs and edge cases as the result of its generation. Not the case with gpt and glm.
@SirAlexthomson What makes this a particularly interesting and difficult to understand decision is that the open source models will be more powerful than fable 5 is now, in just a few months. GLM 5.2 is already nearly there in many benchmarks. This will certainly stifle American innovation.
@TheHBrand@aicodeking Looking at his benchmarks, they are visual heavy, and GPT 5.5 quite underperforms in that area. For backend I find GPT 5.5 excellent, better than Opus. Also I have to correct GPT 5.5 way less than Opus. Fable seems best visually with GPT-level backend capabilities.
@XavioMtl@Swedtraders@thatstarwarsgrl Yeah he clearly touched the whole granite. The Swedes planted that camera there because they know everyone touches the granite a few times out of 80 stones and couldn't win without a scandal. It's so common and has no impact compared to sweeping so it's never called.
@capt_ivo@cb_doge You'd use vacuum radiators to cool GPUs in space. For example, and RTX 5090 needs 2-4kg of cooling equipment in space (pipes+panels). In bulk, the launch would cost you $3-4k per cooled gpu today. In a few years, only $200-$300. Space quickly becomes much more feasible.
@Beth03Lori@capt_ivo@cb_doge You'd use vacuum radiators to cool GPUs in space. For example, and RTX 5090 needs 2-4kg of cooling equipment in space (pipes+panels). In bulk, the launch would cost you $3-4k per cooled gpu today. In a few years, only $200-$300. Space quickly becomes much more feasible.
@gailcweiner@xai Grok 4.2 multi-agent is different from just system prompts + scripting. The agents are jointly trained on the same model with RL so they deeply specialize and debate in parallel natively during inference. This results in lower hallucinations, better accuracy and better edge case.
@caviterginsoy@nim_chimpsky_ I noticed that when opus 4.6 was released the AI trolls came out in full force taking a dump on the model - couldn't believe the amount of hate. The same with Grok.
My limited testing of Grok 4.20: excels at natural sounding writing, research, and response accuracy.
@DavidSteadson@SilentSnow89@skogstrast He obviously touches the stone from top to bottom - you would be dishonest to think otherwise. The only difference is there isn't a Swede trying to film from the side to needlessly stir controversy.
@AndrewButchart1 @Eli_Doubletap He rubbed the whole stone down from top to bottom - 100% obvious. The only difference is that the Canadians didn't bring a side view camera to try and cause an international scandal, knowing it's what all the pros do.
@KevinMcCurdy Sadly, it made the news for a good reason: to be able to say "Canada is cheating". Many accounts amplified this non-issue purely for political reasons. What was surprising is how quickly it became a "thing". It spread like wildfire.