NYC TLC · 2023–2024 · quantile modeling
Why a 20-minute trip sometimes becomes a 55-minute trip
Which NYC taxi routes have deceptively ordinary median travel times but unusually severe upper-tail delays, and what spatial and temporal characteristics explain that unreliability?
The finding
Among 3,966 well-supported route pairs whose medians agree within 10% but whose P95s differ by at least 40%, the less reliable route was faster at P10 in 82% of comparisons — and slower at P95 in 100%.
| Quantile | Share of pairs where the unreliable route is SLOWER | Median difference (min) |
|---|---|---|
| P10 | 18.4% | -0.95 |
| P25 | 20.8% | -0.57 |
| P50 | 84.4% | +0.67 |
| P95 | 100.0% | +8.89 |
Across the whole sample the criterion identifies 47,402 qualifying pairs spanning 1,071 distinct unreliable routes.
Example
Medians 6% apart; P95s 42% apart. The 7.8 min gap is 63 standard errors wide, so it is not sampling noise.
Time of day
Each route is compared against its own overnight baseline. Aggregates are trip-weighted across supported routes.
| Bucket | Δ median | Δ P95 | Amplification |
|---|---|---|---|
| morning commute | +4.1 | +7.9 | 1.92× |
| midday | +6.0 | +10.9 | 1.84× |
| evening commute | +6.4 | +11.2 | 1.76× |
| evening | +2.6 | +4.3 | 1.65× |
The evening commute adds +6.4 min to the median but +11.2 min to P95 — an amplification of 1.76×. Unreliability peaks at 16:00 (TailSpread 12.3 min) and bottoms at 03:00 (5.6 min).
Correction · spatial normalisation
| Borough | Stable zones | Trips | Median P50 | TailSpread | P95/P50 |
|---|---|---|---|---|---|
| Manhattan | 59 | 67,232,156 | 12.4 | 9.9 | 1.83× |
| Queens | 15 | 6,680,652 | 18.0 | 11.9 | 2.09× |
| Brooklyn | 22 | 572,098 | 23.4 | 19.1 | 2.32× |
| Bronx | 1 | 13,809 | 20.0 | 18.4 | 3.23× |
Ranked scale-free, the naive reading inverts: Manhattan is the most reliable borough, not a centre of unreliability. Its trips are short and its P95 is about 1.83× its median. The genuine outliers are the airports, which combine long trips with wide tails.
Only 97 of 261 zones clear the stability bar; the other 164 are drawn hatched and excluded from the colour scale. They carry just 0.82% of trips.
Modelling
LightGBM and a linear quantile-regression baseline, both at q ∈ {0.50, 0.90, 0.95}, on a strictly chronological split: train 2023-01-01 → 2024-07-01 (5,000,477 rows), test 2024-07-01 → 2025-01-01 (1,998,222 rows), Bernoulli-sampled.
| Quantile | Target | LightGBM coverage | Linear coverage | LightGBM pinball | Linear pinball |
|---|---|---|---|---|---|
| q = 0.5 | 50% | 44.69% | 46.47% | 1.73 | 2.42 |
| q = 0.9 | 90% | 87.13% | 88.40% | 1.02 | 1.43 |
| q = 0.95 | 95% | 93.21% | 94.07% | 0.67 | 0.92 |
Coverage at q = 0.95 is 93.21% against a 95% target. It under-covers — the dangerous direction for a risk model — and the cause is not a mis-specified model but a moving target: the test period is genuinely slower, with P95 rising from 42.3 to 45.5 min. A random split would have hidden this entirely.
Fitting the three quantiles independently left 1.54% of rows violating Q50 ≤ Q90 ≤ Q95 (largest violation 17.9 min). Row-wise rearrangement removes all of them.
Distance tells you how long a trip usually takes; time tells you how badly it can go wrong. Trip distance loses 13.3 percentage points of gain share from q = 0.50 to q = 0.95 (44.9% → 31.6%), while hour-of-day features roughly double (8.4% → 15.4%).
Secondary result · near-null
Hourly ERA5 reanalysis at Central Park, joined on local pickup date and hour (100.0% match rate). The marginal comparison is non-monotonic — moderate rain appears faster at P95 than dry conditions — which is a sign of confounding, not a result.
| Band | Hours | Trips | P50 | P95 | Δ P50 | Δ P95 |
|---|---|---|---|---|---|---|
| dry | 14,944 | 63,533,610 | 13.0 | 43.0 | 0.00 | 0.00 |
| light | 2,310 | 10,371,411 | 13.1 | 44.1 | +0.07 | +1.03 |
| moderate | 239 | 1,033,647 | 12.9 | 42.2 | -0.08 | -0.80 |
| heavy | 51 | 178,317 | 12.9 | 46.2 | -0.07 | +3.17 |
Matching within (route, time bucket) across 9,364 cells removes the composition effect and recovers a consistent but tiny effect:
Method
24 monthly TLC Parquet files are read by DuckDB directly from disk and rewritten as a ZSTD dataset partitioned by year and month (1.689 GB, 24 partitions). Nothing above 100k rows is ever loaded into pandas.
Raw distributions are profiled before thresholds are chosen. Rules are applied as an ordered sequence in which each rule is credited only with rows that survived every prior rule, so removals sum exactly to raw − clean. That identity is asserted at runtime.
| Rule | Rows before | Removed | % of remaining |
|---|---|---|---|
| complete record | 79,479,946 | 0 | 0.00% |
| timestamp ordering | 79,479,946 | 29,079 | 0.04% |
| within window | 79,450,867 | 144 | 0.00% |
| duration bounds | 79,450,723 | 1,573,349 | 1.98% |
| distance bounds | 77,877,374 | 918,088 | 1.18% |
| speed bounds | 76,959,286 | 12,531 | 0.02% |
| fare bounds | 76,946,755 | 916,818 | 1.19% |
| valid zones | 76,029,937 | 912,948 | 1.20% |
Retention: 94.51% (4 further rows dropped as exact duplicates).
Route-level claims require ≥500 trips. The floor discards 46,008 thin routes at a cost of only 2.6% of trips, because NYC taxi demand is extremely concentrated. The median supported route has a TailSpread of 14.1 min and a P95/P50 ratio of 1.71×.
GROUP BY (pu_zone, do_zone) -> count, mean, P50, P95 of duration_min, median of 5 trials, each in a fresh subprocess so peak RSS is per-trial and no engine benefits from a warm cache.
| Rows | pandas s | Polars s | DuckDB s | pandas MB | Polars MB | DuckDB MB |
|---|---|---|---|---|---|---|
| 100,000 | 0.180s | 0.068s | 0.202s | 120 | 76 | 128 |
| 1,000,000 | 0.297s | 0.076s | 0.207s | 307 | 178 | 169 |
| 5,000,000 | 0.814s | 0.120s | 0.272s | 764 | 588 | 344 |
| 10,000,000 | 1.510s | 0.171s | 0.349s | 1,327 | 1,104 | 551 |
At 10M rows Polars is 8.8× and DuckDB 4.3× faster than pandas, and DuckDB uses 2.4× less memory — which is why the 75M-row aggregations run on it. At 100k rows pandas actually beats DuckDB; benchmarking a single size would have produced a misleading conclusion.
Apple M4 Pro · 12 logical cores · 25.8 GB RAM · Python 3.12.13
Limitations