
Normalisation applies an explicit transformation to numerical values so data with very different magnitudes can enter the same calculation or comparison. Its value is not that it wipes out difference, but that an untreated numerical scale can masquerade as importance.
Suppose several homes are being compared. Commuting time may range from 15 to 60 minutes, while weekly rent may range from 400 to 800 Australian dollars. Adding the raw numbers would let rent dominate merely because its figures are larger. Mapping each feature to a range from 0 to 1 and then assigning weights better represents the real judgement: how much rent am I willing to exchange for how much time? Yet the chosen limits and treatment of outliers still affect the outcome. A common scale does not remove the value judgement.
Normalisation is also not identical to standardisation. Standardisation commonly centres values on the mean and scales them by the standard deviation. In machine-learning usage, normalisation can instead mean scaling each vector to unit length, and terminology varies across fields. Nor is it ranking, which preserves order while potentially discarding distance. The stable point is that normalisation changes representation to clarify the conditions of comparison; it does not prove that the original features are sensible or turn income, time and experience into the same kind of thing.
https://scikit-learn.org/stable/modules/preprocessing.html
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.