@ID_AA_Carmack@dshandler@DIrmschler While CUDA/OpenCL runs fast, it is difficult to experiment with since writing kernels is a lot more annoying than just running stuff on the CPU with e.g. OpenMP. Also, DL is very slow in general - unconventional architectures can be made much faster. See: https://t.co/yTIBadu6N5