Scaling puts numeric features onto comparable ranges so no feature dominates purely because of its units. Salary in the tens of thousands and age in the tens are not comparable until you fix that.
It is essential wherever the algorithm compares magnitudes across features. Distance-based methods like k-nearest neighbors and k-means, support vector machines, principal component analysis, and any model with a regularization penalty all qualify. Neural networks need it too.
The mechanism repeats each time. Distance sums squared differences, so the widest feature owns the result. A penalty applied per coefficient shrinks whichever features happen to use small units. Variance-maximizing methods follow units directly.
Tree-based models do not need it. The failure mode elsewhere is quiet: nothing errors, the model just ignores your small-scale features. Fit the scaler on training rows only.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓