Backpropagation is how a network works out, for every weight, whether nudging it up or down would reduce the loss. It runs after a forward pass. The forward pass computes predictions and stores each layer's intermediate values along the way.
Then the loss gradient flows backwards, layer by layer. At each step the chain rule from calculus combines the gradient arriving from above with the local derivative of that layer. Multiply them and you get the gradient for that layer's weights, plus the signal to hand to the layer below.
The practical cost is memory. All those stored activations sit in graphics processing unit (GPU) memory until the backward pass consumes them. That is why a batch that fits during inference can still blow up during training.
Rewriting in plainer words…
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.