@sonyatweetybird throwing like 16-32 B200s with ring attention, sequence parallel and parallel vae decode , it’s possible with any model, would be interesting if the cost is significantly less. Hosting costs increase when you use GPUs as utilization drops so people don’t do it
@CalvinMccarter in large scale training for image/video generation you barely see an epoch, so what is better, seeing the whole dataset or seeing half of it ? Might be good to test this on a larger dataset where you have like billion + examples
@elonmusk I used it all day yesterday instead of opus4.7 and it’s very good, didn’t miss the bigger slower model. The speed at which it runs is amazing.
@cryptopunk7213 as if anthropic checks the license of each github repo it uses for pre-training, many of those could be written by chatgpt even if the license was fine
@aakashgupta The throughput is the same, so hosting costs don’t change. Also, how do you fit a 500B parameter model on 800 mm of silicon ? That would need 10x scaling of moore’s law
@Voltaje9@ElonClipsX If you do end to end fusion with transformers it’s same as using multiple cameras , if you have parallel models then you will have issues
@jasondeanlee I think this is true in general because Gemini3’s thinking budget is very low compared to GPT5.1. They nerfed it because probably it was too expensive . OpenAI is still ahead in the game.