@andrewho03 Agree on all except "spamming fde won't help".
But given this your own company thesis weakens. Without high frontier lab valuations to foot the bill, who can afford to buy your data?
@ethantsliu Use ppo or llm-judge feedback. For the latter case give it as hint to rollout again or resume rollout just before the identified error (there's a deepmind paper on this approach)
@IbrahimDagher20 "cyber exploits requires compositional reasoning, which emerges from the strength of your pretrain"
So does swe, knowledge work, deep research and other agentic domains. This is not the explanation. More likely is US has better cyber env and more cyber pre/mid/post training data
@signulll All bad speculations. Competition is less in china and he gets to look like a national hero building "openAI of China". Pure ego/ideology/patriotism -driven.
@Bitcoin_Teddy This conflating two points. Intermittent fasting is good. But depending on your schedule eating before 6pm is terrible. I eat dinner at 10pm and often the next meal is not until 2-3pm the next day. Still sharp as ever.
@waterloo_intern How did you get fable thinking traces? You mean the modified reasoning traces right? Afaik Anthropic doesn't expose raw reasoning traces
@asklivermore $xlv is extended from last week and $lly just had a big pop Friday. Would it make sense to wait return to EMA level to add or just DCA continuously here..?