When Train ‘Punctuality’ Improves, Do Cancelled Services Count as Success?

When Numbers Start Making Decisions · Season Two, “When the Measure Becomes the Target” · Article 5

1. Can a train that never runs be forever on time?

A passenger waits for an 8:10 train. It is cancelled, and the next service arrives twenty minutes later. For the passenger this is an obvious delay. Under an indicator that counts only how many trains that actually ran were punctual, however, the cancelled service could disappear from the denominator. More cancellations might make the remaining trains look more reliable.

That is the problem in the title, but the current New South Wales rail punctuality measure provides a positive answer. A 2017 Audit Office review confirmed that the punctuality indicator introduced in 2013 treats cancelled services and skipped stops as late. It also stopped excluding events considered beyond the operator’s control, such as severe weather and police operations, from the public result. NSW Audit Office: “Passenger Rail Punctuality”

Rail punctuality is therefore a useful counterexample for this season. Organisational adaptation does not have to end in gaming. A measure can anticipate an obvious route of evasion and include it within the definition of failure. The harder question then appears: even if cancellation cannot improve the score, does a train arriving at its terminus four minutes and fifty-nine seconds late mean every passenger arrived punctually?

2. How is “on time” constructed?

For Sydney Trains suburban services, a train must stop at all stations in the timetable and reach the final station within five minutes of its scheduled arrival to count as punctual. Intercity services receive a six-minute tolerance. The suburban system has long operated with a target of 92 per cent. The 2023 Sydney Trains Review also identified a service-availability target of delivering 99 per cent of scheduled services. Transport for NSW: “Sydney Trains Review Final Report”

This is more realistic than insisting that a train arrive within the exact timetable minute. Rail operations contain small variations, and zero tolerance would turn seconds with little practical impact into failures. Five minutes is nevertheless not a natural boundary at which the passenger’s loss changes character. It is an operational standard.

The numerator is the number of services meeting the definition during a period; the denominator is the relevant set of scheduled services. The unit being recorded is a train, not a passenger. A peak service carrying a thousand people and a late-night service carrying fifty count as one failure each if both are ten minutes late. A passenger affected across two interchanges does not automatically become three units of delay.

The calculation is not wrong. It compresses service reliability into a train-level proxy. The mistake occurs when managers and readers forget the object of the proxy and interpret train punctuality as passenger punctuality.

3. How an older measure rewarded the wrong conduct

Before 2013, the on-time-running indicator contained a notable weakness. A train could skip scheduled stations to recover time and still be counted as on time if it reached the terminus within the tolerance. For passengers unable to board or forced to continue to the wrong station, the service plainly failed. For the final timestamp, it succeeded.

The replacement definition requires all scheduled stops and counts skipped stops and cancellations as late. It did not abandon quantification. It moved the statistical object closer to the purpose of the service. This reflects an important design principle: when an organisation can improve a score by sacrificing something the score omits, the definition should bring that sacrifice back into the measure.

Closing one route leaves a further boundary. A railway can increase the time allowed in its timetable, so that more trains count as punctual. The timetable introduced after the 2003 Waterfall accident was slower, and punctuality improved for a sustained period. Part of that change represented safer and more realistic running times. Padding is not automatically manipulation; an impossibly tight timetable creates constant failure and unsafe pressure. But if planned journey time grows while punctuality rises, passengers need to know whether they received a better service or merely a more reliable version of a slower one.

Operational control can also prioritise the peak trains most important to the 92 per cent target during a disruption, displacing delay to off-peak times or other lines. The network total improves while the distribution of loss becomes less fair.

4. Why current data need a second measure

Sydney Trains reported 24-hour suburban punctuality of 87.5 per cent in 2024–25, below the 92 per cent target. The intercity result was 75.0 per cent. Monthly suburban performance ranged from 81.3 to 93.4 per cent. Sydney Trains Annual Report 2024–25

The figures reveal a substantial reliability problem but not the whole passenger loss. Two systems can each record 90 per cent punctuality. In one, the remaining 10 per cent of trains are six minutes late. In the other, a small number of major disruptions make 10 per cent of trains an hour late. The same pass rate conceals radically different time losses, missed connections and social costs.

The Audit Office therefore commended the Customer Delay Measure. It estimates the difference between each passenger’s expected and actual arrival, including early and late arrival, changed sequence, skipped stops and cancellations. Passenger volumes can be incorporated, giving more weight to delays on heavily loaded services and moving the object of measurement closer to the traveller’s final outcome.

Passenger delay is not perfect. It depends on journey data, route inference and a counterfactual timetable. If unreliable service causes people to drive, those who leave the railway stop appearing in rail data. Yet it corrects the assumption that each train has equal social importance.

5. The measure changes priorities in the control room

A rail network is highly interdependent. A late train occupies a path, affecting later services and connections. During disruption, controllers decide which train moves first and whether to skip stops, turn a service short or cancel it. Punctuality targets become part of those decisions.

A well-designed target encourages maintenance of failing equipment, better timetables, earlier recovery action and control of cascading delay. A narrow target can encourage optimisation at the compliance boundary instead of reduction in total passenger time. A system might even run fewer services to reduce variation, while passengers experience greater crowding and longer waits.

Frequency and punctuality involve a real trade-off. The Audit Office’s Transport 2019 report noted that punctuality fell after more than 1,500 weekly services were added in the 2017 timetable. More trains reduced planned waiting but left less space in the network to recover after failures. NSW Audit Office: “Transport 2019”

An operator should not be rewarded for punctuality while frequency deteriorates, nor for service volume while reliability collapses. A weighted overall score can still hide a severe failure. It is better to establish conditions that cannot compensate for one another: deliver enough services, limit extreme delay and protect safety, then improve punctuality within those constraints.

6. How do we verify that better punctuality is a real improvement?

Rail organisations should publish together:

  • the proportion of scheduled services delivered, with cancellations and skipped stops;
  • train punctuality and the distribution of minutes late, rather than only passage of the five-minute boundary;
  • passenger-weighted total delay and excess waiting time;
  • differences among lines, times, directions and stations;
  • planned journey time before and after timetable changes;
  • recovery time after major incidents and the timeliness of passenger information; and
  • safety, crowding and frequency, so they cannot be traded silently for punctuality.

Definitions and their versions also matter. If reporting moves from the peak to the entire day, the tolerance changes, or cancellations receive a new treatment, a historical line cannot be joined without explanation. A metric preserves comparable institutional knowledge only when its rules remain stable or its breaks can be traced.

Conclusion: a good measure counts the attempt to evade it

The direct answer to the title is that under the stronger New South Wales definition, a cancelled service or skipped stop is treated as late. It does not disappear from failure. That design deserves to be preserved and shows that metric governance can learn.

Train punctuality still cannot independently represent passenger punctuality. My judgement is to retain the five- and six-minute standards and the 92 per cent operational target, while publishing service delivery, passenger-weighted delay, extreme delay, frequency and safety alongside it. Any widening of the timetable should be visible in its own right rather than presented as a punctuality gain.

Once a number enters a control room, it helps determine which train moves and which loss is repaired first. A reasonable metric does not eliminate adaptation. It makes the adaptations that best improve the score align as closely as possible with the interests of passengers. Counting cancellations and skipped stops as failures is one step in that direction.

Principal sources


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.