@kalomaze@yacineMTB the codex vs claude debate is a skills issue, ie if you give them the same skills files and force them to use TDD both are more capable than their users
@_arohan_@samlakig cc seems to take this approach without being explicitly told, but i've noticed on the initial plans / correctness reviews it typically spawns one agent per specialized task instead of having many agents tackle the full plan / correctness review
@yacineMTB a note from the power sector, if the physics is well-understood and the objective is clear, it's easier (read: faster) to stuff them both into a loss function and train directly via SSL than it is to build a full grid simulator and try to use RL to recover the desired behaviour