Wanted to try decisions models so I made D1 @liquidai destroy the minesweeper ! All built with @MistralAI large 4 (really strong model btw, but current inference speed is low ๐ข)
I tried to make D1 play pokemon but even with a vlm supervisor for planning it was unsuccessful.
@ZipMind_AI@liquidai@MistralAI Plain decision can't even pass the starting screen. I added a vlm for short horizon planning like, enter the house, take the stairs but even with that the decision model couldn't properly move in the game.
@helloiamleonie@adithya_s_k@huggingface Awesome work! Have you considered data-mixing strategies for harnesses? Weighting rollouts from different harnesses could maybe lead to some interesting results
@AnthropicAI This seems highly related with the current trend of Astra bding looped model and people saying it will abstract the reasoning. With this you could possibly try to interpret what happen between layers and also between loops.
@hkproj I Love the video, 2h30 through it. I saw that in arithmetic intensity you used sparse peak flops instead of dense and wondered if it was intentional?