We trained a 10.9M byte-level recurrent Transformer on L3 and L6. (Loop 3 and Loop 6)
Yet L4/L5 improved too, L8 held up, and the L3→L6 gain grew during training.
Same weights. More compute. Better predictions.
@WyronGaines@MyronGainesX Not defending it but prolly the guy meant that US people don't follow their own religion like muslims do given the atrocities that happen in america...