Our track record
We log every prediction we show and score it once the flight lands, misses included. Scored on 4.1M flights from Jan–Jul 2026 that the model never saw, using weather forecasts as they were issued at the time. Trained on 2013–2025. Test period: 2026-01-01..2026-07-31, 4,120,343 flights.
When we say 30%, does it happen 30% of the time?
Dots on the dashed line mean our percentages are honest. Bigger dots stand for more predictions. The vertical axis is how often it actually happened.
Sharper as departure gets closer
How much better than “this route is usually late X% of the time”, by how far ahead we predicted whether a flight would arrive 15+ minutes late.
Our worst days
What we predicted a day and two days out vs. what actually happened — misses included, not just the good calls.
| Airport | Date | Actual | 2d out | Day ahead |
|---|---|---|---|---|
| EWR | 2026-02-23 | 100% | 64% | 63% |
| DCA | 2026-01-25 | 100% | 98% | 97% |
| BOS | 2026-02-23 | 100% | 94% | 95% |
| JFK | 2026-02-23 | 100% | 69% | 70% |
| LGA | 2026-02-23 | 100% | 72% | 75% |
| IAD | 2026-01-25 | 99% | 72% | 72% |
| BWI | 2026-01-25 | 98% | 90% | 89% |
| PHL | 2026-01-25 | 98% | 86% | 84% |
| DFW | 2026-01-24 | 95% | 99% | 100% |
| CLT | 2026-01-31 | 92% | 69% | 77% |
| LGA | 2026-01-25 | 91% | 86% | 84% |
| EWR | 2026-01-25 | 91% | 82% | 79% |
Methodology, in plain English
“How often we’re right” means: out of every 100 times we said a 30% chance, close to 30 of them actually happened — not one prediction alone, but the whole set. We score every prediction we ever showed, good or bad, against what actually happened once the flight landed or was cancelled, on flights the model never saw during training. “Sharper as departure gets closer” compares our forecast to a simple baseline (how often this route is usually delayed) and shows how much more we add as more of the picture — live weather, the FAA’s plan, the incoming aircraft — comes into view.