Kronecker-factored layers in PyTorch: wall time and GPU memory
Published:
Three PyTorch implementations of the block-wise sparse Kronecker layer, compared with dense and group LASSO baselines: wall time and peak GPU memory per training step at batch 256 and 20,000.
