Original Post
Hi, So basically, three layers can be shown to be theoretically maximally expressive with feed-forward ANNs. Of course using four or more layers doesn't decrease this expressivity, but it doesn't increase it either. However, I've seen networks with four or more layers; is there a reason for this? Intuitively it would seem that the non-linearity caused by a high number of layers just makes the learning more difficult; wouldn't it be better to stick with three layers and just increase the number of neurons? Thanks, -- Mikko