The habit of double-checking the basics has bitten me again. This time it was GPU acceleration.
To see that GPU makes the calculations faster, I’ve run a code snippet to benchmark matrix multiplication. The result was: 10 times slower than it should be.
Issue 1: the GTX 1650 doesn’t have Tensor Cores for matrix multiplications. cuBLAS, and then PyTorch, doesn’...