Every network is organized into three kinds of layer, each with a fixed job. The input layer is not really computing anything. It simply holds your features, one unit per number you feed in - pixel values, sensor readings, or an encoded row.
Hidden layers sit in the middle and do the actual work. Each one takes the previous layer's outputs, applies weights, and produces a new representation. Early hidden layers pick up simple patterns. Later ones combine those into higher-level features.
The output layer is shaped by the task. One unit for a regression target or binary score, ten units for ten classes. Get that shape wrong and the loss function will not match your labels, which is a common early bug.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓