Leakage means your training data contains information that would not exist at prediction time, or that quietly encodes the answer. The model learns that shortcut, because it is the easiest signal available.
The damage is that your offline numbers become fiction. Validation accuracy looks excellent, so you stop investigating and ship. In production the leaked column is empty, or it arrives after the decision, and performance collapses toward baseline.
It is hard to catch because nothing crashes. The code is correct; the data is wrong. A classic case is predicting fraud using a chargeback_filed column, when chargebacks only get filed after fraud is confirmed.
The cost is time and trust. Weeks of tuning on a leaked feature buy nothing, and business decisions get made on a number that was never true.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓