We hit this while tuning an agent harness for a customer: an agent that plans the monthly rounds of about 50 field workers across mainland France. The scheduling logic was not the hard part. The hard part was a cheap, trustworthy answer to a simple question, asked thousands of times per run: how long is the drive from here to there?
That question is not specific to one customer. Any agent that plans visits, deliveries, or rounds needs it. A real routing engine is exact but too slow inside the optimizer's loop: Lee and Chae (2023) (ouvre un nouvel onglet) measured a full origin-destination matrix from GraphHopper taking 505 seconds at 1,000 nodes. And a language model is right there, so an agent might simply reach for its own sense of geography.
So we measured three options on 1,200 French address pairs, against a self-hosted routing engine: a distance-only curve, a residual neural network with 29 map features, and a language model guessing from memory. The network wins at every distance: 9.53% mean absolute percentage error (MAPE), versus 13.70% for the curve. The language model keeps up with the curve past about 20 km, and even beats it on long trips. Under 20 km, its error is an order of magnitude worse. The customer's planner now uses the residual network.
Why a national mix is harder than one city
Our customer's addresses make the problem harder than most published work. The agent covers roughly 30,000 addresses a year. Only about 3,000 need a route each month, but that monthly batch is spread across the whole country, not one city.
Most published VRP (vehicle routing problem) travel-time work studies a single metropolitan area: fewer very long legs, more stops close together, no need to think about a motorway. Planning a month of visits nationwide mixes short in-town hops and 500 km cross-country legs in the same optimization run. There, a motorway can cut a long leg's duration far more than it helps a short one.
The standard answer is a surrogate model, an instant lookup once trained. Lee and Chae's surrogate, a small neural network on raw coordinates, reached 7.68% MAPE against GraphHopper on Korean road data. We ran the same kind of experiment on this more mixed French setting.
We built a haversine-only rational curve as a cheap baseline, then a residual neural network that corrects it using 29 offline features (distance to the fast road network, commune density, distance to Paris) computed from OpenStreetMap and French address data. Ground truth is a self-hosted OSRM server on the France OpenStreetMap extract, queried 3.5 million times and downsampled to a 1.19-million-pair training table. This part is not new science: it repeats the residual-correction idea used elsewhere for travel-time estimation (Boeing and Zhou, 2026 (ouvre un nouvel onglet)) and confirms it holds for France.
The new part is the language model. We asked gpt-5.6-luna (OpenAI) to estimate driving duration, straight-line distance, and road distance between two French addresses, from memory alone, on the same 1,200 pairs the neural network is scored on. Speed was never the question: contraction hierarchies already answer an exact routing query in about 110 microseconds (Bast et al., 2016 (ouvre un nouvel onglet)). The question is whether an agent's instinct about geography can be trusted at all. GPT4GEO (Roberts et al., 2023 (ouvre un nouvel onglet)) asked the same of straight-line distance and found 12% to 51% relative error on city-pair distances.
Two address pools, not one messy national list
We used town halls and a sample of street-level addresses instead of this customer's real stops, both to protect their data and because the two are geographically similar anyway: a logistics operator's stops and a country's town halls are both spread out to cover the territory, not clustered in a few neighborhoods.
Town halls exist in almost every French commune, one per commune, which makes them a convenient, verifiable proxy: each one sits on a real, addressable, drivable street, unlike a purely random point that might land on a field or a private driveway. Our pool has 35,152 of them. In dense areas neighboring town halls can sit 1-2 km apart, and much farther in rural France, which represents medium and long trips well but leaves very short trips under-represented: two random town halls are rarely a 5-minute drive apart.
To fill that gap, we added a second pool: Base Adresse Nationale (BAN) street-level addresses, sampled from 12 departments (299,125 points after dropping addresses more than 150 m from the OSRM road graph). We did not use the full national BAN dataset directly. It holds tens of millions of individual house-number entries, many of them apartment numbers or entries that do not resolve to a distinct drivable point, which would have meant handling a long tail of special cases (duplicate entries, unreachable addresses, several house numbers sharing one entrance) just to build a sample. Restricting to 12 departments and sampling from what was left was the cheaper, still representative choice.
What raw trip data looks like, before any model
Before fitting anything, look at what a random pair of French addresses actually gives you: straight-line (haversine) distance on one axis, OSRM's uncongested driving duration on the other. The raw per-pair cloud alone is too noisy to show much by eye at this point density, so both panels also plot the median duration in each distance bin, an aggregate, not a fitted model.
Two things show in the trend line, not in the raw cloud, which is noisy at this scale. Zoomed into short trips, the trend rises fast for the first few kilometers, then bends over sharply. Zoomed out to the full range, that same bend means the trend approaches a straight line far from the origin, an asymptote in slope rather than in duration itself; duration keeps growing with distance, just at a near-constant rate past a few hundred kilometers. A straight line through the origin cannot produce both shapes at once. That is the case for a curve with more than one regime.
Why a saturating curve, and not a straight line
A trip's average speed is not constant. It rises with distance, then saturates. Gallotti et al. (2016) (ouvre un nouvel onglet) measured this directly on 780,000 Italian vehicles: mean speed grows from about 18 km/h on short trips to a plateau near 55 km/h past 2 hours, because a longer trip has more chances to reach a faster road layer, and France has only so many layers (city street, primary road, motorway). We see the same shape in our own 30,000 ground-truth pairs.
That saturating-speed shape is why the baseline we fit is a 3-parameter rational function of haversine distance h, not a straight line through the origin:
$$
Fit by least squares on the training table, the coefficients for mainland France (Corsica excluded) are:
$$
This single formula, fed nothing but the straight-line distance between two addresses, already reaches 13.70% MAPE on our 1,200-pair evaluation sample. For a VRP solver that just needs a reasonable cost to rank candidate routes, distance alone gets most of the way there.
Closing the rest of that gap needs more than a better fit of the same one input. The remaining error correlates with things haversine distance cannot see: how far a trip's endpoints sit from a motorway on-ramp, how dense the origin and destination communes are, whether the trip has to cross a city center to reach the fast road network. That points to adding features and a model that can combine them, which is what the rest of this post does.
Three estimators, one routing engine as ground truth
Data. 35,152 town halls plus addresses from the Base Adresse Nationale (BAN) across 12 departments, 334,277 points total on mainland France (Corsica and overseas territories excluded), as described above. A self-hosted OSRM server, on the Geofabrik France extract, answered 3.5 million raw /table queries in about 2.5 minutes, downsampled to a 1.19-million-pair training table, stratified so every haversine band from 50 meters to 1,300 km is represented.
Models. The rational curve above. A residual multilayer perceptron (MLP, input 29, hidden layers [256, 128, 64], GELU, AdamW) predicting a log-ratio correction on top of the curve, trained with an 80/10/10 split. We also tried a gradient-boosted tree model (HistGBDT) on the same 29 features as a quick sanity check against a different model family: it lands close to the MLP (8.74% versus 7.73% MAPE on the in-distribution test split, see results/ablation_gbdt.csv), so the result is not an artifact of one architecture. All three (rational curve, MLP, HistGBDT) were trained and saved in a separate project and imported here as finished artifacts: we did not retrain or tune any of them for this post.
LLM probe. Two arms, both on gpt-5.6-luna at reasoning.effort=medium:
- LLM bare: only the two addresses.
- LLM + distance: the two addresses plus the true haversine distance stated in the prompt. This arm exists to separate not knowing where a city is from not knowing how to turn a known distance into a duration.
We picked a moderate reasoning effort on purpose. "How many minutes between these two addresses" is a narrow, well-specified task, not a problem that needs deep reasoning.
The point is not to build a production estimator this way: the residual network stays the right tool for that. The point is to check whether the geography a general-purpose model already holds is directionally trustworthy. A customer's own planning agent could lean on that intuition without anyone measuring it first.
The system prompt tells the model to assume free-flowing traffic with no congestion, since OSRM's duration is uncongested, and to answer from memory with no tool or browsing. Every call uses OpenAI's structured-output mode with a strict JSON schema, so a malformed response is close to impossible; what does happen is a response status="incomplete", meaning the model's reasoning ran out of token budget before it produced an answer. We retry those with a doubled token budget, up to 7 attempts, at 10 requests in parallel.
We rejected the Batch API: it is 50% cheaper but takes up to 24 hours and reports no per-request latency, and latency was part of what we wanted to measure.
Sample. 1,200 pairs, 100 drawn from each of 12 haversine bands (same bands as the training table, seed 271828), scored identically for all three estimators.
How the residual network fits on top of the curve
The rational curve only sees one number, the straight-line distance. The residual network sees 29 features per pair (fast-road-network access, commune density, geometry) and learns the part the curve cannot see: instead of predicting minutes directly, it predicts log(true duration / rational curve's estimate), the multiplicative correction the curve is missing. This is residual modeling, an established move in physics-guided machine learning (Willard et al., 2022 (ouvre un nouvel onglet)) and, under the name delta-machine learning, in quantum chemistry (Ramakrishnan et al., 2015 (ouvre un nouvel onglet)), where a correction model trained on a small fraction of the data reaches the accuracy of a far more expensive method. The network only has to learn what a 3-number curve misses, not the whole function.
Figure 5 shows where that delta comes from conceptually: the rational curve is a baseline that already tracks the general shape of the data, but for any single real trip, the true duration usually sits a bit off that baseline. The network's whole job is to learn to predict that gap.
At inference time, the baseline and the predicted delta combine (prediction = baseline * exp(delta)) into the final duration.
The residual network beats the curve at every distance
On the full 1,200-pair sample:
| Estimator | MAE (min) | MAPE (%) |
|---|---|---|
| Rational curve | 14.17 | 13.70 |
| Residual MLP | 7.20 | 9.53 |
Figure 7 checks the same result a different way: does each estimator's prediction actually sit on the predicted = true diagonal, or just in the right order of magnitude?
A 4.6-minute miss can be a 195% error
A relative error means very different things depending on the trip. We picked one real pair per range (short, medium, long), each close to that range's median duration, and show what every estimator actually predicted.
| Range | Trip | True duration | Estimator | Predicted | Error | Error (%) |
|---|---|---|---|---|---|---|
| Short (0.9 km) | Baziège to Baziège | 2.4 min | Rational | 2.6 min | 0.2 min | 9.4% |
| MLP | 2.6 min | 0.2 min | 7.7% | |||
| LLM bare | 7.0 min | 4.6 min | 194.9% | |||
| LLM + distance | 4.0 min | 1.6 min | 68.5% | |||
| Medium (12.7 km) | Virming to Wuisse | 23.0 min | Rational | 20.1 min | 3.0 min | 12.9% |
| MLP | 23.3 min | 0.3 min | 1.3% | |||
| LLM bare | 19.0 min | 4.0 min | 17.4% | |||
| LLM + distance | 17.0 min | 6.0 min | 26.1% | |||
| Long (777.4 km) | Osenbach to Plouguernével | 625.9 min (10.4h) | Rational | 704.3 min | 78.4 min | 12.5% |
| MLP | 620.2 min | 5.7 min | 0.9% | |||
| LLM bare | 638.0 min | 12.1 min | 1.9% | |||
| LLM + distance | 580.0 min | 45.9 min | 7.3% |
On the long trip, the rational curve misses by 78.4 minutes, more than an hour, but that is only 12.5% relative error. On the short trip, the bare LLM misses by 4.6 minutes, which sounds negligible, and that is 194.9% relative error, almost three times the true duration. The same number of minutes or the same percentage means opposite things depending on which end of the distance range it happened on.
If you compare travel-time estimators, report errors by distance range, not one overall MAPE.
Four metrics, because MAPE alone misleads on short trips
MAE and MAPE tell part of the story, but both have known blind spots on this kind of data (see the warning above). We also report Median Symmetric Accuracy (MSA) and Symmetric Signed Percentage Bias (SSPB), from Morley, Brito and Welling (2018) (ouvre un nouvel onglet), computed from the median of log(predicted / true) instead of a mean of raw percentage errors, which keeps them well-behaved near zero and symmetric between over- and under-prediction.
| Metric | What it represents, concretely | Sensitive to |
|---|---|---|
| MAE (min) | The typical size of the miss, in minutes you can picture: "on average, off by X minutes." | Long trips dominate: a 5% miss on a 500-minute trip outweighs a 50% miss on a 2-minute one. |
| MAPE (%) | The typical miss as a share of the true duration: "on average, off by X%." Easy to compare across very different trip lengths. | Explodes on very short trips, and mechanically rewards estimators that under-predict (Tofallis, 2015). |
| MSA (%) | The same idea as MAPE, a typical relative error, but computed so a 2x overestimate and a 2x underestimate count as the same size of miss, and short trips do not dominate the number. | Stays stable and interpretable even in the 0-1 km band, where MAPE is close to meaningless. |
| SSPB (%) | Whether the estimator runs systematically early or late: positive means it tends to over-predict, negative means it tends to under-predict, independent of how big any single miss is. | Only shows a systematic direction. A model can have a good SSPB (no bias) and still miss individual trips badly. |
On the full 1,200-pair sample:
| Estimator | MAE (min) | MAPE (%) | MSA (%) | SSPB (%) |
|---|---|---|---|---|
| Rational curve | 14.17 | 13.70 | 12.07 | -5.17 |
| Residual MLP | 7.20 | 9.53 | 6.59 | -0.18 |
| LLM bare | 13.03 | 50.36 | 14.29 | +0.37 |
| LLM + distance | 12.49 | 21.06 | 12.05 | -2.63 |
Reading this table alongside Figure 6: the rational curve runs slightly early on average (SSPB -5.17%), the MLP is close to unbiased (-0.18%), and the bare LLM is essentially unbiased overall (+0.37%) even though its MAPE (50.36%) looks the worst of the four. That gap between MAPE and the other three metrics is exactly the near-zero-band artifact the warning above describes, not a sign the LLM is systematically wrong in one direction.
The language model mostly does not know where the addresses are
The bare LLM also guesses the straight-line distance itself. On the 1,200 pairs, that guess has 57.99% MAPE, just above the 12% to 51% range GPT4GEO found for city-pair distances worldwide, on finer-grained addresses rather than city centroids. Giving the model the correct distance, instead of asking it to guess one, drops its road-distance MAPE from 53.04% to 12.88%, a 4x improvement, for the same reasoning effort. Most of the bare arm's error is not an inability to convert a known distance into a plausible duration. It is not knowing where the two addresses are relative to each other in the first place.
A language model is more trustworthy on long trips than short ones
A frontier language model is far too large for a task this narrow. That mismatch is exactly why we ran this arm. Our customer's planning agent, and agentic tools generally, sit on top of models like this one. Before trusting an agent's intuition about a trip's length, it is worth knowing whether that intuition is any good at all. A purpose-built estimator would still win on cost and latency, easily.
Grouping the same 1,200 pairs into three coarse ranges, short (under 20 km), medium (20-200 km), and long (200-1,300 km), makes the split visible:
| Range | Estimator | MAE (min) | MAPE (%) | MSA (%) | SSPB (%) |
|---|---|---|---|---|---|
| Short (n=500) | Rational | 1.70 | 17.70 | 17.14 | -9.54 |
| Short (n=500) | MLP | 1.19 | 13.94 | 10.78 | -0.69 |
| Short (n=500) | LLM bare | 4.54 | 106.34 | 52.74 | +43.01 |
| Short (n=500) | LLM + distance | 2.33 | 37.29 | 26.00 | +18.85 |
| Medium (n=400) | Rational | 9.33 | 12.27 | 12.03 | -11.23 |
| Medium (n=400) | MLP | 5.57 | 7.60 | 6.46 | -1.11 |
| Medium (n=400) | LLM bare | 9.36 | 12.80 | 11.15 | -5.41 |
| Medium (n=400) | LLM + distance | 8.10 | 10.72 | 9.16 | -5.49 |
| Long (n=300) | Rational | 41.41 | 8.96 | 8.71 | +6.85 |
| Long (n=300) | MLP | 19.39 | 4.75 | 3.41 | +1.14 |
| Long (n=300) | LLM bare | 32.08 | 7.13 | 7.10 | -6.30 |
| Long (n=300) | LLM + distance | 35.29 | 7.78 | 7.97 | -7.66 |
The short-range SSPB tells its own story: the bare LLM over-predicts short trips by 43.01% on average (positive SSPB), consistent with a model reasoning about a plausible-sounding trip ("a few towns apart must be at least 10 minutes") rather than actually knowing the two addresses are 900 meters apart. On long trips the bias flips sign and shrinks (-6.30%), closer to the rational curve's own bias in that range (+6.85%).
None of this changes the practical answer. A language model call takes seconds and cannot signal when its knowledge is stale, so it is not a candidate to sit inside a VRP loop. It does mean an agent's estimate of a long inter-city leg is more trustworthy than its estimate of a short in-town hop.
If your agent reasons about trip lengths on its own, do not let it guess short in-town trips from memory. Give it a real estimator.
Limitations
- The rational curve, MLP, and LLM go stale at very different costs. A new motorway changes nothing the rational curve depends on except the target, so it refits on a small fresh sample. Nine of the MLP's 29 features come from a frozen OpenStreetMap snapshot, so a new interchange changes both the inputs near it and the target, requiring a fresh extract and a retrain. The LLM has a hard knowledge cutoff and no update path at all, and it cannot signal when a route it describes no longer exists.
- One call per pair, default temperature. We did not repeat any LLM call or fix a temperature. Per-band and per-range LLM numbers have no variance estimate; the overall numbers (n=1,200) are more stable.
- A retrospective staleness gap. Ground truth was collected in August 2026, about 6 months after the model's February 2026 knowledge cutoff. Any road opened in that window is a source of LLM error unrelated to reasoning ability.
- Concurrency may inflate measured latency. Ten LLM calls ran in parallel; measured wall time can include server-side queuing, not pure generation time.
- The MLP's geography features have their own known weaknesses. A network's detour depends more on angular position relative to a city center than on absolute bearing (Lee et al., 2017 (ouvre un nouvel onglet)), which the model's bearing features do not capture, and circuity by distance band follows a U-shape rather than a straight trend with distance to Paris (Levinson and El-Geneidy, 2009 (ouvre un nouvel onglet)). The model is used as shipped, not retrained, so this is a known limit on the baseline, not something this post fixes.
- The BAN sample is not evenly spread. The 2,994 street-level addresses that ended up with a resolved label cluster in 3 of the 12 source departments (see Figure 1), an artifact of which pairs the original sampling happened to draw, not a deliberate regional choice.
- Mainland France only, one language model, one prompt. Corsica and overseas territories are excluded. Results should not be assumed to hold for a different country's road network, a different model, or a differently worded prompt.
- None of this measures live traffic. Every duration here, ground truth included, is an uncongested estimate.
So, what should a routing agent use?
For a VRP loop that needs a fast, offline travel-time estimate, the residual network remains the right tool. It reaches 9.53% overall MAPE against the rational curve's 13.70%, a lead that holds at every distance range we tested. Its inference is in the same class of speed as contraction hierarchies, not seconds. This is what our customer's planner now uses.
A general-purpose language model, asked to guess purely from memory, is a poor fit for short in-town trips: an order of magnitude worse than either baseline under 20 km. On medium and long trips it lands close to, and sometimes past, the rational curve.
For a customer planning a mix of short local stops and cross-country legs every month, that split matters more than any single headline number. An agent's intuition about a long inter-city drive deserves more trust than its intuition about a 5-minute hop across town. Neither should be trusted with the latency or the staleness risk of an LLM call inside an actual optimization loop.
Open questions
- Would angular position relative to a city center, instead of absolute bearing, close more of the residual network's gap?
- How much do the per-range language model numbers move with repeated calls, which this run did not make?
- Do the results hold for another country's road network, another model, or another prompt?