@bugarela Both solutions are great! It would be interesting to see if there are differences in terms of generating the traces given the same cosmos base code.
If you feel behind on learning about Large Language Models, try these:
– Readers: free 200+ page book covering pre-training, generative models, prompting and alignment
– Programmers: Karpathy’s neural networks zero to hero playlist including implementing GPT-2 from scratch