@testingcatalog Just started testing HY4 Preview on OpenVibeEval. The 770B / 49B active setup and 1M context are interesting, but I’m more curious about how it actually performs against the other models on real agent tasks.
Results here: https://t.co/RrrjSMvXZl
@WorkBuddy_AI Just started testing HY4 Preview on OpenVibeEval. The 770B / 49B active setup and 1M context are interesting, but I’m more curious about how it actually performs against the other models on real agent tasks.
Results here: https://t.co/RrrjSMvXZl
@theo That’s what I noticed too. Gemini 3.7 Flash has been surprisingly strong on frontend for me, especially on design and taste. I’ve been tracking its results here: https://t.co/m3Asj0bNa3
@Da7_Tech That’s what I noticed too. Gemini 3.7 Flash has been surprisingly strong on frontend for me, especially on design and taste. I’ve been tracking its results here: https://t.co/m3Asj0bfkv
@aimlapi The cost difference is pretty crazy, but I'd be more interested in how they compare on actual frontend output. I've been running the same kind of head-to-head tests, and the differences in UI quality can be pretty significant: https://t.co/ZwOmismS59
@thegenioo This is exactly why I like testing these models on frontend tasks. The differences in design taste can be surprisingly large, even with the same prompt. I’ve been seeing some very different results across models in my own benchmark runs: https://t.co/LcxDguqkOa
@ariskaa_ai Hy4 is definitely worth testing while it’s free. I’ve been running it through real-world frontend tasks, and the results are interesting. It’s strong in some areas, but the UI quality isn’t consistently at the level of the best frontier models: https://t.co/VaApo1WFBs
@Da7_Tech I've seen a similar gap when testing these models on frontend tasks. Speed and benchmark scores don't always translate to good design decisions. I’ve been comparing their actual UI outputs here: https://t.co/ZwOmismS59
@theojaffee The gap is definitely getting smaller, but I don't think the models are equally strong across every task. I've been benchmarking them on real-world frontend scenarios, and some are much better than others at actually producing polished UIs: https://t.co/ZwOmismS59
@HarshithLucky3 This is exactly the kind of task I’m interested in. 3D/WebGL is one of the areas where model differences become really obvious. Kimi K3 has been particularly strong there, so I’d be curious how Hy4 compares: https://t.co/VaApo1WFBs
@TypingMindApp I’d be curious to see how they compare on more practical frontend tasks too. GLM-5.3-Flash has been pretty inconsistent for me lately, so I ran it through several benchmarks to compare it with other models: https://t.co/hDJE6ntk8h
@utilq1vy@BohuTANG@Zai_org > Definitely degraded. I ran the same kind of benchmarks and the difference between 0x-alpha and the current GLM-5.3-Flash is pretty clear. The comparison here shows it: https://t.co/0EI9S3ImR0
@BohuTANG@Zai_org Definitely degraded. I ran the same kind of benchmarks and the difference between 0x-alpha and the current GLM-5.3-Flash is pretty clear. The comparison here shows it: https://t.co/0EI9S3IUGy
Most AI frontend benchmarks show complex, polished screens.
They look great, but you probably aren't building that daily.
So I built OpenVibeEval around realistic scenarios to see how AI performs.
https://t.co/ZwOmisnpUH
@TencentHunyuan I’ve been testing it on frontend tasks too, and the results are interesting. It’s definitely not the same story across every UI task: https://t.co/RrrjSMwvOT
@qilua02 Free for 30 days is tempting, but honestly Ling-3.0-Flash isn't that good in my experience. I’d rather spend the free quota testing models that actually hold up on real frontend tasks: https://t.co/7HVnjrEFyu
@baseten Open-weight models are getting seriously competitive. I’ve been testing GLM-5.3 on frontend tasks too, and the results are more nuanced than the usual coding benchmarks. Some of the UI outputs are genuinely impressive: https://t.co/4zGh2AsWoX