@lumpenspace@AsaCoopStick CoT uncontrollability implies agents can't prevent schemes from appearing in CoT, bc if they cannot do CoT in lowercase they cannot do obfuscated CoT. But they can reason in the response in lower case, and we show that the reasoning can be moved there at low accuracy cost...
@GenReasoning Could it just be that this market is already very efficient, and there's no alpha in the data you force the models to use? Is there a (non-overfit) strategy that makes money using the provided data?
@FazlBarez@amang0112358 I'm also confused about the choice to boost refusals in the score, since this makes the correlation between PAC and refusal rate stronger.
@FazlBarez@amang0112358 I'm confused about the treatment of refusals and how that is inflating the PAC scores for some models; the .91 rank corr. is striking. Did you try filtering out the refusals or resampling and looking the impact on the PAC score? (Is your data with all responses available on HF? )
@docmilanfar remarkable how much simpler the modern proof is -- big win for abstractions like Jensen's + the optimization definition of the median getting identified and into the water supply.
otoh you do lose all these cute bespoke proofs that give nice intuitions
A post about the vibes of LLM generated algorave music : https://t.co/xB1685UFDZ
Here's a sample: https://t.co/TAeCRGsZFs
Also: My tooling for vibe-coding vibes with Strudel https://t.co/DNq5YtwFbF ... have fun!