
Coverage is not the number of times something succeeded. It begins by defining a range that ought to be considered, then asks which parts of that range have been checked, practised or included. Without a defined range, a percentage has no stable meaning: covering eight cases out of ten is not the same as covering eight out of one hundred.
Suppose a speech-recognition system correctly transcribes one hundred sentences recorded in a quiet room. That shows good performance under that condition. If real use also includes street noise, varied accents and different speaking distances, those untested conditions remain blank areas. Adding more quiet-room recordings may increase the sample size without expanding coverage.
Coverage therefore differs from accuracy. Accuracy asks how many tested cases were handled correctly; coverage asks which regions of the relevant case space were reached. It is not the same as representativeness either: a test set may name many categories while giving a rare but important one only a single example. Broader coverage can reveal omissions, but it cannot by itself show that testing was deep enough or that future performance will be reliable.
The value of the concept is that it turns “we tried it” into sharper questions: under which conditions, within what boundary, and what remains untested? In learning, it warns against mistaking repeated success on one type of problem for mastery of a whole topic. In evaluating AI or everyday plans, it keeps absent cases in view. Coverage is not a substitute for correctness; it is a map of explored regions and remaining blanks.
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.