Method
How a delay becomes a number
TfL publishes two quite different things. One is a declared status — “Good Service”, “Minor Delays” — set by people in a control room, and it is authoritative about what they have decided. The other is a live feed of arrival predictions: every train the signalling system can see, where it is, and when it expects to reach each platform ahead of it.
Delayometer ignores the first and measures the second. Everything below is what happens between that feed and the number on the front page.
Three signals, not one
A journey gets longer in three different ways, and confusing them makes a status product useless. So we measure each of them separately:
Waiting time. At each platform, in each direction, we sort the trains due by how far away they are. The gaps between them are the headways — how long you would wait, arriving at random. Trains that appear twice in the same feed are deduplicated first, because otherwise we would invent gaps nobody experiences.
Running time. Where TfL gives a train an identifier we can follow, we watch the same train reach one station and then the next, and take the difference. That is how fast the line is actually moving, as opposed to how often something arrives.
Trains that never came. The other two are measured from the trains that ran, so a train cancelled outright contributes nothing to either. We count the trains actually in the feed against the number the published timetable says are due, over the same window. A line running half its service is failing its passengers however punctual the surviving half is, so a shortfall caps the score rather than being averaged into it.
A line can fail any of the three ways, and the set is what separates “busy but moving” from “stopped” from “half the trains are missing”. They are combined into one score per section of track, then smoothed so that a single odd reading does not repaint the map.
Absence is a measurement
The hardest thing for a system like this to see is nothing at all. A platform with one train and nothing behind it produces no gap to measure, and a platform with no trains produces nothing whatsoever — so for a while the parts of a line in the most trouble were simply the parts we had least to say about, and a line could lose a whole branch while its score went up.
Now an empty platform is recorded as what it is: a statement that no train is coming within the next half hour. That is a stronger claim than any single measured gap, and it is scored as one.
Normal is a measurement, and a promise
“Four minutes between trains” means nothing on its own. At Oxford Circus on a Tuesday at 08:30 it is a problem; at Upminster on a Sunday evening it is the timetable. So every comparison is against a baseline built for that station, that direction, that kind of day, and that fifteen minutes of it — from weeks of our own measurements, in London local time so the clock changes do not smear the morning peak.
Periods where TfL had declared disruption are excluded from the baseline. Without that, a line that is chronically late slowly teaches the system that late is normal, and then stops being reported as late at all — which is the single easiest way to build a status product that is quietly worthless.
That exclusion has a cost we found by looking for it. At quiet hours nearly every observation is an absence, and absences are kept out of the baseline for the same reason disruption is — so what survives is the handful of minutes that happened to have two trains close together, and the learned “normal” for half past midnight comes out looking like the middle of the day. Compared against that, the last trains of the night read as a severe delay.
So the published timetable is the second reference, and where it promises a longer gap than we have learned, the promise wins. You cannot be late relative to a service that was never offered. Where we have measured a service better than the timetable — a line running more trains than promised — the measurement wins instead, because that is the standard the line has actually set.
The score
Each section of track gets 0–100, where 100 is running as it usually does. Those roll up to a direction and then to a line, and the number is always paired with words, because “62” is not a thing anyone can act on and “about two minutes longer than usual” is.
For a section of track the bands are fixed on the score alone: 90 and above is running normally, 75 and above is a minor delay, 35 and above is delayed, and below that is severe. A line is not coloured on its score alone, because a line carries a sentence as well and the colour must not contradict it: the chip walks exactly the same branches as the words, in the same order. Under a minute added to a journey is inside the noise, and quoting it would be a precision the measurement does not have.
A line short of trains is coloured for that too, on the same threshold that makes us say so — including when the shortfall is confined to one branch. A median cannot express a dead branch on an otherwise working line, so the share of the line that is short of trains is reported beside it, and the words become “part of the line” rather than a figure that would be true of the line as a whole and useless to the people on that branch.
What we do not know, and say so
Grey is not green. A section nobody has been measured across yet is grey and reads Measuring. This matters more than any other rule here: a status product that renders missing data as a good service is worse than no status product at all, so nothing on this site turns an absence into a reassurance.
A line in constant trouble becomes harder to measure, not easier. Because declared-disruption periods are kept out of the baselines, a line TfL has flagged for much of the week loses most of its samples for those hours — and with no picture of normal for, say, a Tuesday evening, there is nothing left to compare a Tuesday evening against. Over half the Central line’s measurements are currently excluded this way, which is why its evenings often read Measuring rather than carrying a score. We would rather keep the exclusion and lose the coverage than let a chronically late line quietly redefine what late means, but the cost lands exactly where a measurement would be most useful, and we would rather you knew that than wondered.
The DLR has no train identifiers. Its feed does not name its trains, so we can measure how long you wait but never how fast the train then goes, and never count how many distinct trains turned up. It is scored from waiting times alone, permanently, and its confidence is capped to say so. It does publish a timetable, so we can still tell the end of its service from a failure of it.
The Elizabeth line publishes no timetable to us. Every station returns nothing, so on that line we can say how long the gaps are but never how many trains were supposed to fill them. Everywhere a timetable is missing, we assume trains are due rather than assume they are not — the opposite would quietly excuse an empty platform, which is the failure mode this whole page exists to argue against.
Shared track is not yet separated. The Circle, District, Hammersmith & City and Metropolitan lines share rails for much of central London. Congestion there is currently counted against every line that runs over it, which is right often enough to be useful and wrong often enough to be worth telling you about.
Engineering works are not modelled. The nightly end of service now is — the timetable tells us when the last train has gone, so a shut railway is no longer reported as a broken one. A weekend closure or a part suspension is different: it is a change to the timetable that the published timetable does not always carry, and we do not yet recognise one by name. Declared disruption is kept out of the baselines, which limits the damage without fixing it.
The timetable is what was promised, not what was planned for today. We compare against the published schedule, so a service running to a temporary timetable will look short of trains to us even when it is running exactly as intended. That is the right way round — a passenger holding the published times has the same complaint — but it is a reason to read “fewer trains than usual” as a statement about the promise rather than about anyone’s competence.
Old data is hidden, not dimmed. If our collector stops, the measurement disappears rather than lingering. A twenty-minute-old score presented as current is a fabrication, and the fastest way to lose the only thing this product sells.
Where the data comes from
Everything is derived from TfL’s Unified API, the same open feed that powers every other app in this category. What differs is that we keep the raw record and do the arithmetic, rather than relaying the status TfL declares over the top of it.
Powered by TfL Open Data. Contains OS data © Crown copyright and database rights. Delayometer is not affiliated with Transport for London, and the measurements here are ours, not theirs.
Check it rather than believe it
Everything above is an argument. The declared-against-measured timeline is the same argument with the evidence attached: a week of TfL’s declared status for every line, drawn against what we measured over the same hours, on one clock. Find a morning where the two bands differ and decide for yourself which of us was watching. It reports the times we disagreed in both directions, including the ones where TfL saw a problem we missed.