The sigmoid is the S-shaped function that maps any real number into the range zero to one. Large negative inputs land near zero, large positive inputs near one, and zero maps to exactly 0.5.
const sigmoid = z => 1 / (1 + Math.exp(-z));
sigmoid(0); // 0.5
sigmoid(2.2); // 0.900
Logistic regression needs it because a weighted sum of features is unbounded, while a probability is not. The sigmoid also preserves order, so ranking by raw score matches ranking by probability. It is smooth and differentiable everywhere, which is what lets gradient descent fit the weights.
The cost is saturation. Once a score sits far from zero, the curve is almost flat and the gradient nearly vanishes. A confidently wrong prediction then learns very slowly. Unscaled features push scores into that flat region, which is one reason training stalls.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓