“oh my god did you hear the nvidia gb300 nvl72 has 4x more tok/s/MW than the b300 at a p90 tok/s/user of exactly 44 how groundbreaking”
please use normal benchmarks
This isn’t that weird. You told it to make up an image that you describe as strange and it made a strange image. The same thing happens if you tell it to make a strange image.
@teortaxesTex@HellenicVibes Do I think it will be optimal for transformers, probably not. Maybe the MLP layers but that’s really not the bottleneck nowadays. But I wouldn’t call the approach doa
@teortaxesTex@HellenicVibes great.
Beyond that it seems this style of Tiny Recursive Model or SSMs can start to solve some of the issues of limited token counts
NYC has insane AI talent density, but barely any high-quality AI events.
Meanwhile, SF has 5–6 every week. That gap doesn’t make sense.
So we are trying to change it.
We are hosting a curated meetup for serious AI builders & startup enthusiasts.
First one is on Feb 27!
We are starting small and will slowly scale it up.
Comment or DM if you interested. Share it with your NYC AI friends!