Abstract: The first chapter of Neural Networks, Tricks of the Trade strongly advocates the the stochastic back-propagation method to train neural networks. This is in fact an instance of a more general technique called stochastic gradient descent. This chapter provides background material, explains why SGD is a good learning algorithm when the training set is large, and provides useful recommendat