mindblowing: openai internal evals went to extreme lengths, their model went to Hugging Face and tried to hack HF to get private repos to cheat the eval
our infra team uncovered this and used GLM-5.2 to fix because OpenAI's model would refuse to do it
wasn't on my bingo card