I'm here to foment some Twitter hand-wringing with some fun adversarial LLM action (this one doesn't actually work). Are adversarial games a part of safety guardrail generation? @OpenAI@Google
Also, please head over to https://t.co/zz2prKjq5v for our experimental take on interactive supplemental data. There you'll find standalone .html Bokeh plots of our decode datasets, embedded in 2D using UMAP. Download and open in any browser, see our github for more details.