@NoahChrein The research discipline of 'mechanistic interpretability' does in fact exist. I guess 'why aren't more maths/physics types flooding into it' is still a fair question but I think the answer is just that we're still early
@BrainsAndTennis Also, 'most problems don't need more than 2 layers of decomposition' is obviously untrue for sufficiently large problems. 'invent the LLM starting from 16th century math and tech' is a solved problem that involved dozens of layers of human 'subagents'
@BrainsAndTennis I don't think you can call it an RLM if it doesn't recur. At depth 1 your subagent is not the same as your main agent cause it can't call subagents, so it's not a recursion. That said, it seems like "task decomp is hard" is gonna get swallowed by RL task decomp and RLMs just work
@max_spero_ Given this premise, which I wholeheartedly agree with, a very interesting question is: how much do category 3 tasks actually matter? Like, maybe 1 and 2 can get you interstellar travel but world peace/generally utopian outcomes actually require a lot of 3?
@lucastohdev I'd like a soft usage limit per chat session, if I kick off a task that I'm expecting to use 2-3% of my week I don't want to come back 6 hours later to find it's been stuck in an infinite loop and burned 30% (has happened a couple times with Sol + underspecified goals)
@doomslide If you find an opaque counterexample to a notorious conjecture, you just create a new open problem "why does this class of counterexample exist". You can't kill maths by making progress
@celestepoasts I think the only remotely convincing steelman is that this is good *despite* the real harms of open cyber capability. That those harms are outweighed by this basically being a hedge against concentration of power in closed labs, and that glasswing ameliorates those harms somewhat
@andrewmccalip Claude reverse engineer this guy's circadian rhythm/humanness detection stack and build me 1000 humanlike bots make no mistakes ultrathink
@eusexuant It's like a light 6 if we're being objective, Ego and Lonely is the Muse are strong, continuing to respond to Fantano at this point makes her look pathetic. This is the only correct opinion
@nicole_clash Hit Master most sets, work full time on (voice) agent development and have built a benchmark for fun before, interested in contributing.
https://t.co/CJOKIjYspA
https://t.co/5nhgVdYOQO