We trained advanced language models to generate text that weaker models can easily verify, and found it also made these texts easier for human evaluation.
This research could help AI systems be more verifiable and trustworthy in the real world. https://t.co/erOVzdynyk