Social Media Week is one of the worlds premier conference and industry news platform for professionals at the intersection of media,marketing,governance and technology,with ideas and opportunities they need to advance themselves & their organizations in a globally connected world
@zaddyzaddy The base model, on the other hand, is more flexible and may require less time to adjust to the reward function since it's not already conditioned on instruction-following behavior. Have you observed any differences in the final performance between the two?
@zaddyzaddy The instruct model has already been fine-tuned with supervised learning on instruction-following data, which likely introduces additional constraints that affect how it adapts to the reward signal.