@LarsvanderLaan3 The DR inference for TMLE I knew of also needed iteration and was if anything a little more complex, so that's still great! Might try a few numerical experiments myself later :)
@_sergiotahini@lauwaermer bei uns wurden die leute irgendwann 17 und haben mit mcdonalds gehalt shishas für 400€ gekauft nur um dann immer noch köpfe zu bauen die kratzen
@MatthewBJane@AJThurston Think the idea is that eg MI makes analyses more complex and at low % missing bias shouldn't be too large, so choosing the simpler analysis isn't "too bad". The threshold is arbitrary though (I know 5%), and that argument doesn't convince me - always better to do MI if beneficial
Here's a little simulation showing that model selection can bias parameter and SE estimates. I'll also attempt a crude explanation down the line. Set-up: an RCT with a binary treatment A, a continuous (normally distributed) outcome Y, and prognostic covariates W. 1/10
If ANOVA (unadjusted) and ANCOVA (baseline-covariate adjusted) both return unbiased estimates for the ATE in an RCT, why would it matter how one chooses the covariates for the ANCOVA? You see recommendations for preregistration or theory, but is this unnecessary?
#RStats
@SolomonKurz@reluctantcrim@rappa753 Many thanks to all of you for all the information! I'll get around to creating a blog soon, but a guest post on a blog sounds fantastic (and I'm extremely spoiled because both blogs that offered are so good). I'd love to do it on yours @SolomonKurz bc your thread started it all!
@SolomonKurz My thought was also blog post, but I don't have a blog, so I'll need to set one up (any resources you can recommend?). Has been something I wanted to do for a while though and this seems to be a good reason to do so :)
@SolomonKurz Thank you! This is far from the first late-night-procrastination-simulation I've done so yeah, maybe it's time to not just let them rot away on my computer... I'll let you know if I get to it!
This will always happen when using this procedure, so we'll have bias. Of course, this is a somewhat simplistic example (covariates usually aren't mutually independent), but it should work as a heuristic and also explain the bias observed in the p-value-selection scenario. 10/10
We could naively instead choose the subset maximizing the ATE estimate. The issue: this estimate will contain only those covariates pushing the estimate away from the null, and therefore towards one of the tails of this distribution. 9/10