Reading group · week 3
Why residual connections made very deep networks trainable
In 2015, He et al. observed something odd: adding layers to a plain network made training error worse, not just test error. Their fix was to let each block learn a residual F(x) = H(x) − x on top of an identity shortcut. Have a look at page 1 and the introduction below; hover the bracketed citations to see what they refer to.
The key result is in Figure 1: the 56-layer plain network has higher training error than the 20-layer one. Residual networks remove this gap, which is why almost every modern architecture, Transformers included, has shortcuts around each block.
Next week: Attention Is All You Need.