“every conceivable misunderstanding or misinterpretation or objection, including those that would occur only to the malicious or the clinically literal-minded”
abduction is bad: you don't want your children to be abducted. deduction is bad: you don't want your wages to be deducted. so induction is bad, by its own standards. so counterinduction is good, by its own standards.
@beyarkay the caption on that very graph explains the upwards jumps ("additional sets of agents were launched on July 10th and 11th"), but yeah we don't know why the drop on the 12th happened.
@stellahymmne@jessi_cata while it’s obvious that this would lead kimi to defect, the fact that this generalizes to endorsing CDT in abstract discussion is less obvious (see anthony’s comment on lesswrong, for instance)
@deadlydentition@jessi_cata all changes to probability distributions are zero-sum. trying to reinforce in a "positive sum" way just introduces noise (since you just upweight whatever you happened to sample).
@TheZvi i disagree that ryan's threat model is just "models be scheming"; don't the "hackistan" and "slopolis" regimes in https://t.co/JcpZp7hvSW or "Current AIs seem pretty misaligned to me" correspond to what you describe? redwood has also written about score-seekers elsewhere.
@JeffLadish this seems qualitatively different; i don’t think we have any public evidence of instances in separate samples actually complying with each other’s requests for assistance, or even copying each other’s behavior? (and it seems like any such evidence would’ve been highlighted.)
picture lifting the top part up and around (it ends up on the right) and pulling the bottom part down and around (it ends up on the left), so that the band is just through the handle.
@voteprincemongo@churchmouseirl what kind of journal doesn’t force authors to redact acknowledgements for blind review? unless by “peer review” you don’t mean what most people imagine, which would be ironic given your original tweet
@nathanrs nice! a quick test of gzip: does it rank infinigram drafts in a reasonable way? (each on a corpus of tim williamson’s work)
seems like – as you might expect – very compressible stuff is repetitive, and very incompressible stuff is junky.
@DeepDishEnjoyer@cxgonzalez (but then you still have the original puzzle of “it’s not the case that boo murder” being bad english, so there needs to be some story here about english syntax)
@DeepDishEnjoyer@cxgonzalez to be clear it’s not obvious that we shouldn’t try and do ordinary (compositional / propositional / intensional) semantics on expressives like “boo murder” or imperatives like “don’t murder” – these could also each express some fact or falsehood.