@lminozem @adamhamdev I totally agree. And I think most bosses love to have employees that always work at home too, cause for them that means that they're going to make a lot more profit in general.
How to learn a useful critic for policy optimization? In our paper, we try to answer this question by learning the gradients of the value function via model-based RL.
Work done at @nnaisense together with @WJaskowski.
Paper: https://t.co/b94QnW96wa
Blog: https://t.co/qepjjpdPkl