The theory of Thompson Sampling in LQR is finally here. We were concerned with this problem since 2017 following discussions w @alelazaric. We provide regret bound of this practically implementable algorithm in various scenarios. The proof turned out to be very nice&elegant. #RL
We prove open problem that Thompson sampling has optimal regret for linear quadratic control in any dimension. Previously only proven in one dimension. We develop novel lower bound on probability that TS gives an optimistic sample. @SahinLale@tkargin_@Azizzadenesheli@caltech