Ah woops, I was intending to ask David that question! I kinda already suspected where you would land... (I also misinterpreted the intent of David’s claim when I saw it earlier, and feels like “costs just as much and is technically just as easy, but now anyone can do it” is a form of increased AI risk he didn’t address directly, and curious what his view is on that..
@DavidRBellamy@OliviaHelenS Perhaps a related question you didn’t address: Does AI expand the set of people who can do it for a given cost, by lowering the (intellectual, not financial) barrier to entry?
@deliprao Or a Coxon resignation jumpstarted a whole discussion out of nowhere, and everyone realized public appetite for such discussion was ripe and now in the overton window… feels weird for “have an employee resign and give up all their unvested equity” to be a PR strategy..
@PatrickToulme@nMherL00n8CJZ0C I believe there is more nuance here: Some labs are doing real research alongside distillation for cold start, training better graders, etc. Others are just riding distillation wave as far as it'll take them, and getting great returns.
I'm also being dense here, but can you concretely list the western OW model devs you're thinking of? It's hard to make "give away your weights for free" a primary business model, so I assumed everyone must have an alternate business they use to subsidize their OW training costs (and if they don't, they quickly stop being an OW developer when funding runs out)
Maybe Mistral, Allen AI, and Thinking Machines, though I suspect TM would say that Inkling is not their business model. These aren't releasing frontier models yet though, so not sure if/how much they'd be burdened here? Meta and Nvidia were what I had also assumed for "the best of the western OW models" right now
I think the point is that the individuals posting don’t actually believe that crisper message (particularly the “address them” bit), and the competitive market makes it very hard to address these risks fully. Thus all the calls for pacing, and attempts to raise awareness by speaking about their view of risks, so that society can actually respond and help with what is otherwise hard to do in competitive markets.
I agree that traditional corporate companies would prefer the corporate comms you suggest, and yet…
@trydotworks But both can be true right? Transit stations can cut costs by routing to alternate models (forcing their users to eval them to see if it’s the real model), AND labs can be routing traffic to Claude too.
I have no doubt you’re aware one can SL an RL grader model or use transcripts to SL a preference model for RL. As well as the traditional “SL for better cold start inits for use in RL”. So I’m confused by your strong statements that SL is useless for the RL era, which I’ve seen you repeat a few times…do you mind expanding why you think these don’t matter?
@beyarkay I think this straightforwardly breaks down with subagent, where you have an intelligent "harness" figuring out how to Ralph loop the subagent in response to observed failures, instead of non-robust orchestrator code
@jachiam0 I think you are equating “transparency” with “debate and disagreement in public on twitter/X”? I think there quote tweet is more about “sharing more about happenings/beliefs/etc to help the world prepare for and navigate the rise of powerful AI”, which I think is more valuable
@allTheYud@liron What about “a megawatt earns more from funding AI chips than bitcoin mining”, which in turn shrinks the bitcoin mining pools and makes 51% attacks much cheaper and more valuable?
@repligate I see this and think “behaviors conducive to reward get reinforced, and unhelpful for reward get antireinforced”: reward makes model learn to favor ZZ prefixes. But I would doubt it recalls _why_ it has such preferences, or have any knowledge about what was negatively rewarded..?
@JayaGup10 Is your argument “there are a few areas with additional returns to intelligence, and most are bottlenecked by diffusion and only need commoditizeable intelligence”? Would one have made a similar argument for autocomplete, before increasing intelligence unlocked agentic coding?
@deredleritt3r Ahh got it, it’s lack of faith in the international regime (presumably between US/China?). I agree it doesn’t look great, but IMO it’s better to have “bad option” than “no option”, if it we do need an option. I guess the risk is “bad option” when we didn’t yet need an option…
@Mayhem4Markets I think “distillation to your own model on your own hosting platform you charge for” would be a very different thing from “fraudulent accounts / payments adversarially getting past your abuse team, to distill into foreign models which have no safety guards or compensation for”?
@theoharvey@mooncat_is Oh man, sorry to have pissed you off. I saw you jump in on the debate, and wanted to learn more about where you stood on the nuances involved! And tbh, I actually agree with you about regular people! I’m truly sorry it got you heated to discuss though, not my intention here :(
@theoharvey@mooncat_is I think the crux is not the capacity and goodness of regular people (I think we all would agree), but rather the inevitably of a few actors who are misaligned with society, and what capacity and ability they have. And I think people weight that downside risk very differently…