Whitening and second order optimization both destroy information about the dataset, and can make generalization impossible:
arxiv.org/abs/2008.07545 We examine what information is usable for training neural networks, and how second order methods destroy exactly that information.