
A latent variable is not simply something mysterious that exists but cannot be proved. It is a variable in a model that is not observed directly but is used to account for relationships among observed data. We see patterns in what can be measured and infer an underlying structure through the model. “Language ability”, for example, cannot be measured like height, yet performance across listening, speaking, reading and writing tasks may share a relatively stable source of variation that a statistical model represents as latent.
A latent variable is not the same as an ordinary hidden fact. A key under a sofa is merely unseen for the moment; a latent variable is first of all a construct within a model, and whether it corresponds to a distinct structure in reality requires further evidence. Nor is it the same as a proxy measure: one test score may serve as a proxy for ability, whereas a latent variable is commonly estimated from relationships among several observable variables.
The distinction matters because AI, psychology and the social sciences often infer unobserved structure from visible outcomes. Inferring a latent variable, however, does not by itself establish that a natural entity has been discovered. A model may predict observations well without proving that its interpretation of the hidden structure is uniquely correct.
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.