← All posts[engineering]

Can an LLM feel the road? Fast, cheap drive-time estimates across France

While tuning an agent that plans field rounds across mainland France, we needed a fast, cheap answer to “how long is this drive?” We compared a distance-only curve, a residual neural network, and a language model guessing from memory: the network cuts error from 13.70% to 9.53% MAPE, and the language model falls apart under 20 km.

Cominty Engineering11 September 202615 min read

We hit this while tuning an agent harness for a customer: an agent that plans the monthly rounds of about 50 field workers across mainland France. The scheduling logic was not the hard part. The hard part was a cheap, trustworthy answer to a simple question, asked thousands of times per run: how long is the drive from here to there?

That question is not specific to one customer. Any agent that plans visits, deliveries, or rounds needs it. A real routing engine is exact but too slow inside the optimizer's loop: Lee and Chae (2023) (opens in a new tab) measured a full origin-destination matrix from GraphHopper taking 505 seconds at 1,000 nodes. And a language model is right there, so an agent might simply reach for its own sense of geography.

So we measured three options on 1,200 French address pairs, against a self-hosted routing engine: a distance-only curve, a residual neural network with 29 map features, and a language model guessing from memory. The network wins at every distance: 9.53% mean absolute percentage error (MAPE), versus 13.70% for the curve. The language model keeps up with the curve past about 20 km, and even beats it on long trips. Under 20 km, its error is an order of magnitude worse. The customer's planner now uses the residual network.

Why a national mix is harder than one city

Our customer's addresses make the problem harder than most published work. The agent covers roughly 30,000 addresses a year. Only about 3,000 need a route each month, but that monthly batch is spread across the whole country, not one city.

Most published VRP (vehicle routing problem) travel-time work studies a single metropolitan area: fewer very long legs, more stops close together, no need to think about a motorway. Planning a month of visits nationwide mixes short in-town hops and 500 km cross-country legs in the same optimization run. There, a motorway can cut a long leg's duration far more than it helps a short one.

The standard answer is a surrogate model, an instant lookup once trained. Lee and Chae's surrogate, a small neural network on raw coordinates, reached 7.68% MAPE against GraphHopper on Korean road data. We ran the same kind of experiment on this more mixed French setting.

We built a haversine-only rational curve as a cheap baseline, then a residual neural network that corrects it using 29 offline features (distance to the fast road network, commune density, distance to Paris) computed from OpenStreetMap and French address data. Ground truth is a self-hosted OSRM server on the France OpenStreetMap extract, queried 3.5 million times and downsampled to a 1.19-million-pair training table. This part is not new science: it repeats the residual-correction idea used elsewhere for travel-time estimation (Boeing and Zhou, 2026 (opens in a new tab)) and confirms it holds for France.

The new part is the language model. We asked gpt-5.6-luna (OpenAI) to estimate driving duration, straight-line distance, and road distance between two French addresses, from memory alone, on the same 1,200 pairs the neural network is scored on. Speed was never the question: contraction hierarchies already answer an exact routing query in about 110 microseconds (Bast et al., 2016 (opens in a new tab)). The question is whether an agent's instinct about geography can be trusted at all. GPT4GEO (Roberts et al., 2023 (opens in a new tab)) asked the same of straight-line distance and found 12% to 51% relative error on city-pair distances.

Two address pools, not one messy national list

We used town halls and a sample of street-level addresses instead of this customer's real stops, both to protect their data and because the two are geographically similar anyway: a logistics operator's stops and a country's town halls are both spread out to cover the territory, not clustered in a few neighborhoods.

Town halls exist in almost every French commune, one per commune, which makes them a convenient, verifiable proxy: each one sits on a real, addressable, drivable street, unlike a purely random point that might land on a field or a private driveway. Our pool has 35,152 of them. In dense areas neighboring town halls can sit 1-2 km apart, and much farther in rural France, which represents medium and long trips well but leaves very short trips under-represented: two random town halls are rarely a 5-minute drive apart.

To fill that gap, we added a second pool: Base Adresse Nationale (BAN) street-level addresses, sampled from 12 departments (299,125 points after dropping addresses more than 150 m from the OSRM road graph). We did not use the full national BAN dataset directly. It holds tens of millions of individual house-number entries, many of them apartment numbers or entries that do not resolve to a distinct drivable point, which would have meant handling a long tail of special cases (duplicate entries, unreachable addresses, several house numbers sharing one entrance) just to build a sample. Restricting to 12 departments and sampling from what was left was the cheaper, still representative choice.

Figure 1: The two address pools behind the ground-truth pairs. 35,152 town halls (small dots, one per commune) cover the whole country evenly. 2,994 BAN street addresses (larger dots) resolve to human-readable labels only where they ended up as an endpoint of one of the 30,000 sampled ground-truth pairs, which is why they cluster in only 3 of the 12 source departments (Loire-Atlantique, Haute-Garonne, Côtes-d'Armor) rather than spreading over all 12.

What raw trip data looks like, before any model

Before fitting anything, look at what a random pair of French addresses actually gives you: straight-line (haversine) distance on one axis, OSRM's uncongested driving duration on the other. The raw per-pair cloud alone is too noisy to show much by eye at this point density, so both panels also plot the median duration in each distance bin, an aggregate, not a fitted model.

Figure 2: OSRM driving duration against haversine distance, 6,000 random pairs from the 30,000-pair ground-truth table (light gray), with a binned median trend line (teal) computed from all 30,000 pairs. Left: full range, 0 to about 950 km. Right: zoomed to 0-20 km.

Two things show in the trend line, not in the raw cloud, which is noisy at this scale. Zoomed into short trips, the trend rises fast for the first few kilometers, then bends over sharply. Zoomed out to the full range, that same bend means the trend approaches a straight line far from the origin, an asymptote in slope rather than in duration itself; duration keeps growing with distance, just at a near-constant rate past a few hundred kilometers. A straight line through the origin cannot produce both shapes at once. That is the case for a curve with more than one regime.

Why a saturating curve, and not a straight line

A trip's average speed is not constant. It rises with distance, then saturates. Gallotti et al. (2016) (opens in a new tab) measured this directly on 780,000 Italian vehicles: mean speed grows from about 18 km/h on short trips to a plateau near 55 km/h past 2 hours, because a longer trip has more chances to reach a faster road layer, and France has only so many layers (city street, primary road, motorway). We see the same shape in our own 30,000 ground-truth pairs.

Figure 3: Median implied speed (haversine distance divided by OSRM duration) by distance band, with the interquartile range shaded, on 30,000 pairs of French addresses. Dashed and dotted lines mark Gallotti et al.'s Italian reference values. The French curve follows the same rising-then-saturating shape, plateauing slightly higher.

That saturating-speed shape is why the baseline we fit is a 3-parameter rational function of haversine distance h, not a straight line through the origin:

$$

Fit by least squares on the training table, the coefficients for mainland France (Corsica excluded) are:

$$

This single formula, fed nothing but the straight-line distance between two addresses, already reaches 13.70% MAPE on our 1,200-pair evaluation sample. For a VRP solver that just needs a reasonable cost to rank candidate routes, distance alone gets most of the way there.

Closing the rest of that gap needs more than a better fit of the same one input. The remaining error correlates with things haversine distance cannot see: how far a trip's endpoints sit from a motorway on-ramp, how dense the origin and destination communes are, whether the trip has to cross a city center to reach the fast road network. That points to adding features and a model that can combine them, which is what the rest of this post does.

Three estimators, one routing engine as ground truth

Data. 35,152 town halls plus addresses from the Base Adresse Nationale (BAN) across 12 departments, 334,277 points total on mainland France (Corsica and overseas territories excluded), as described above. A self-hosted OSRM server, on the Geofabrik France extract, answered 3.5 million raw /table queries in about 2.5 minutes, downsampled to a 1.19-million-pair training table, stratified so every haversine band from 50 meters to 1,300 km is represented.

Models. The rational curve above. A residual multilayer perceptron (MLP, input 29, hidden layers [256, 128, 64], GELU, AdamW) predicting a log-ratio correction on top of the curve, trained with an 80/10/10 split. We also tried a gradient-boosted tree model (HistGBDT) on the same 29 features as a quick sanity check against a different model family: it lands close to the MLP (8.74% versus 7.73% MAPE on the in-distribution test split, see results/ablation_gbdt.csv), so the result is not an artifact of one architecture. All three (rational curve, MLP, HistGBDT) were trained and saved in a separate project and imported here as finished artifacts: we did not retrain or tune any of them for this post.

LLM probe. Two arms, both on gpt-5.6-luna at reasoning.effort=medium:

  • LLM bare: only the two addresses.
  • LLM + distance: the two addresses plus the true haversine distance stated in the prompt. This arm exists to separate not knowing where a city is from not knowing how to turn a known distance into a duration.

We picked a moderate reasoning effort on purpose. "How many minutes between these two addresses" is a narrow, well-specified task, not a problem that needs deep reasoning.

The point is not to build a production estimator this way: the residual network stays the right tool for that. The point is to check whether the geography a general-purpose model already holds is directionally trustworthy. A customer's own planning agent could lean on that intuition without anyone measuring it first.

The system prompt tells the model to assume free-flowing traffic with no congestion, since OSRM's duration is uncongested, and to answer from memory with no tool or browsing. Every call uses OpenAI's structured-output mode with a strict JSON schema, so a malformed response is close to impossible; what does happen is a response status="incomplete", meaning the model's reasoning ran out of token budget before it produced an answer. We retry those with a doubled token budget, up to 7 attempts, at 10 requests in parallel.

We rejected the Batch API: it is 50% cheaper but takes up to 24 hours and reports no per-request latency, and latency was part of what we wanted to measure.

Sample. 1,200 pairs, 100 drawn from each of 12 haversine bands (same bands as the training table, seed 271828), scored identically for all three estimators.

How the residual network fits on top of the curve

The rational curve only sees one number, the straight-line distance. The residual network sees 29 features per pair (fast-road-network access, commune density, geometry) and learns the part the curve cannot see: instead of predicting minutes directly, it predicts log(true duration / rational curve's estimate), the multiplicative correction the curve is missing. This is residual modeling, an established move in physics-guided machine learning (Willard et al., 2022 (opens in a new tab)) and, under the name delta-machine learning, in quantum chemistry (Ramakrishnan et al., 2015 (opens in a new tab)), where a correction model trained on a small fraction of the data reaches the accuracy of a far more expensive method. The network only has to learn what a 3-number curve misses, not the whole function.

Diagram
Figure 4: The residual network's inputs and output. It takes the distance and the 28 other pair features in, and outputs a single number, the correction (`delta`) the rational curve is missing.

Figure 5 shows where that delta comes from conceptually: the rational curve is a baseline that already tracks the general shape of the data, but for any single real trip, the true duration usually sits a bit off that baseline. The network's whole job is to learn to predict that gap.

Figure 5: A single trip's true duration (red) sitting above the baseline curve. The network is trained to predict `delta`, the distance between the two, so that `baseline * exp(delta)` lands back on the true duration.

At inference time, the baseline and the predicted delta combine (prediction = baseline * exp(delta)) into the final duration.

The residual network beats the curve at every distance

On the full 1,200-pair sample:

EstimatorMAE (min)MAPE (%)
Rational curve14.1713.70
Residual MLP7.209.53
Table 1
Figure 6: MAPE by haversine band, log scale, for all four estimators on the same 1,200 pairs. The MLP leads at every distance. The rational curve's error is flattest across bands, between 7.5% and 25.4%, because it is a function of distance alone.

Figure 7 checks the same result a different way: does each estimator's prediction actually sit on the predicted = true diagonal, or just in the right order of magnitude?

Figure 7: Predicted versus true duration, rational curve and residual MLP, log-log scale, 6,000 random pairs from the 30,000-pair table, with the `predicted = true` diagonal as a dashed reference. Both estimators track the diagonal closely; the MLP's cloud sits visibly tighter around it, especially below 10 minutes.

A 4.6-minute miss can be a 195% error

A relative error means very different things depending on the trip. We picked one real pair per range (short, medium, long), each close to that range's median duration, and show what every estimator actually predicted.

RangeTripTrue durationEstimatorPredictedErrorError (%)
Short (0.9 km)Baziège to Baziège2.4 minRational2.6 min0.2 min9.4%
MLP2.6 min0.2 min7.7%
LLM bare7.0 min4.6 min194.9%
LLM + distance4.0 min1.6 min68.5%
Medium (12.7 km)Virming to Wuisse23.0 minRational20.1 min3.0 min12.9%
MLP23.3 min0.3 min1.3%
LLM bare19.0 min4.0 min17.4%
LLM + distance17.0 min6.0 min26.1%
Long (777.4 km)Osenbach to Plouguernével625.9 min (10.4h)Rational704.3 min78.4 min12.5%
MLP620.2 min5.7 min0.9%
LLM bare638.0 min12.1 min1.9%
LLM + distance580.0 min45.9 min7.3%
Table 2
Figure 8: Absolute error in minutes (bars) with relative error labeled above each bar, one panel per range, on one real example pair per range.

On the long trip, the rational curve misses by 78.4 minutes, more than an hour, but that is only 12.5% relative error. On the short trip, the bare LLM misses by 4.6 minutes, which sounds negligible, and that is 194.9% relative error, almost three times the true duration. The same number of minutes or the same percentage means opposite things depending on which end of the distance range it happened on.

If you compare travel-time estimators, report errors by distance range, not one overall MAPE.

Four metrics, because MAPE alone misleads on short trips

MAE and MAPE tell part of the story, but both have known blind spots on this kind of data (see the warning above). We also report Median Symmetric Accuracy (MSA) and Symmetric Signed Percentage Bias (SSPB), from Morley, Brito and Welling (2018) (opens in a new tab), computed from the median of log(predicted / true) instead of a mean of raw percentage errors, which keeps them well-behaved near zero and symmetric between over- and under-prediction.

MetricWhat it represents, concretelySensitive to
MAE (min)The typical size of the miss, in minutes you can picture: "on average, off by X minutes."Long trips dominate: a 5% miss on a 500-minute trip outweighs a 50% miss on a 2-minute one.
MAPE (%)The typical miss as a share of the true duration: "on average, off by X%." Easy to compare across very different trip lengths.Explodes on very short trips, and mechanically rewards estimators that under-predict (Tofallis, 2015).
MSA (%)The same idea as MAPE, a typical relative error, but computed so a 2x overestimate and a 2x underestimate count as the same size of miss, and short trips do not dominate the number.Stays stable and interpretable even in the 0-1 km band, where MAPE is close to meaningless.
SSPB (%)Whether the estimator runs systematically early or late: positive means it tends to over-predict, negative means it tends to under-predict, independent of how big any single miss is.Only shows a systematic direction. A model can have a good SSPB (no bias) and still miss individual trips badly.
Table 3

On the full 1,200-pair sample:

EstimatorMAE (min)MAPE (%)MSA (%)SSPB (%)
Rational curve14.1713.7012.07-5.17
Residual MLP7.209.536.59-0.18
LLM bare13.0350.3614.29+0.37
LLM + distance12.4921.0612.05-2.63
Table 4

Reading this table alongside Figure 6: the rational curve runs slightly early on average (SSPB -5.17%), the MLP is close to unbiased (-0.18%), and the bare LLM is essentially unbiased overall (+0.37%) even though its MAPE (50.36%) looks the worst of the four. That gap between MAPE and the other three metrics is exactly the near-zero-band artifact the warning above describes, not a sign the LLM is systematically wrong in one direction.

The language model mostly does not know where the addresses are

The bare LLM also guesses the straight-line distance itself. On the 1,200 pairs, that guess has 57.99% MAPE, just above the 12% to 51% range GPT4GEO found for city-pair distances worldwide, on finer-grained addresses rather than city centroids. Giving the model the correct distance, instead of asking it to guess one, drops its road-distance MAPE from 53.04% to 12.88%, a 4x improvement, for the same reasoning effort. Most of the bare arm's error is not an inability to convert a known distance into a plausible duration. It is not knowing where the two addresses are relative to each other in the first place.

A language model is more trustworthy on long trips than short ones

A frontier language model is far too large for a task this narrow. That mismatch is exactly why we ran this arm. Our customer's planning agent, and agentic tools generally, sit on top of models like this one. Before trusting an agent's intuition about a trip's length, it is worth knowing whether that intuition is any good at all. A purpose-built estimator would still win on cost and latency, easily.

Grouping the same 1,200 pairs into three coarse ranges, short (under 20 km), medium (20-200 km), and long (200-1,300 km), makes the split visible:

Figure 9: MAPE by coarse distance range, same four estimators, same 1,200 pairs. Short trips (under 20 km) are where the bare LLM is worst by far, 106.34% MAPE, an order of magnitude above the other three. Past 20 km, the gap closes fast: medium-range MAPE is 12.80% (LLM bare) versus 12.27% (rational) and 7.60% (MLP), and on long trips the bare LLM (7.13%) actually beats the rational curve (8.96%).
RangeEstimatorMAE (min)MAPE (%)MSA (%)SSPB (%)
Short (n=500)Rational1.7017.7017.14-9.54
Short (n=500)MLP1.1913.9410.78-0.69
Short (n=500)LLM bare4.54106.3452.74+43.01
Short (n=500)LLM + distance2.3337.2926.00+18.85
Medium (n=400)Rational9.3312.2712.03-11.23
Medium (n=400)MLP5.577.606.46-1.11
Medium (n=400)LLM bare9.3612.8011.15-5.41
Medium (n=400)LLM + distance8.1010.729.16-5.49
Long (n=300)Rational41.418.968.71+6.85
Long (n=300)MLP19.394.753.41+1.14
Long (n=300)LLM bare32.087.137.10-6.30
Long (n=300)LLM + distance35.297.787.97-7.66
Table 5

The short-range SSPB tells its own story: the bare LLM over-predicts short trips by 43.01% on average (positive SSPB), consistent with a model reasoning about a plausible-sounding trip ("a few towns apart must be at least 10 minutes") rather than actually knowing the two addresses are 900 meters apart. On long trips the bias flips sign and shrinks (-6.30%), closer to the rational curve's own bias in that range (+6.85%).

None of this changes the practical answer. A language model call takes seconds and cannot signal when its knowledge is stale, so it is not a candidate to sit inside a VRP loop. It does mean an agent's estimate of a long inter-city leg is more trustworthy than its estimate of a short in-town hop.

If your agent reasons about trip lengths on its own, do not let it guess short in-town trips from memory. Give it a real estimator.

Limitations

  • The rational curve, MLP, and LLM go stale at very different costs. A new motorway changes nothing the rational curve depends on except the target, so it refits on a small fresh sample. Nine of the MLP's 29 features come from a frozen OpenStreetMap snapshot, so a new interchange changes both the inputs near it and the target, requiring a fresh extract and a retrain. The LLM has a hard knowledge cutoff and no update path at all, and it cannot signal when a route it describes no longer exists.
  • One call per pair, default temperature. We did not repeat any LLM call or fix a temperature. Per-band and per-range LLM numbers have no variance estimate; the overall numbers (n=1,200) are more stable.
  • A retrospective staleness gap. Ground truth was collected in August 2026, about 6 months after the model's February 2026 knowledge cutoff. Any road opened in that window is a source of LLM error unrelated to reasoning ability.
  • Concurrency may inflate measured latency. Ten LLM calls ran in parallel; measured wall time can include server-side queuing, not pure generation time.
  • The MLP's geography features have their own known weaknesses. A network's detour depends more on angular position relative to a city center than on absolute bearing (Lee et al., 2017 (opens in a new tab)), which the model's bearing features do not capture, and circuity by distance band follows a U-shape rather than a straight trend with distance to Paris (Levinson and El-Geneidy, 2009 (opens in a new tab)). The model is used as shipped, not retrained, so this is a known limit on the baseline, not something this post fixes.
  • The BAN sample is not evenly spread. The 2,994 street-level addresses that ended up with a resolved label cluster in 3 of the 12 source departments (see Figure 1), an artifact of which pairs the original sampling happened to draw, not a deliberate regional choice.
  • Mainland France only, one language model, one prompt. Corsica and overseas territories are excluded. Results should not be assumed to hold for a different country's road network, a different model, or a differently worded prompt.
  • None of this measures live traffic. Every duration here, ground truth included, is an uncongested estimate.

So, what should a routing agent use?

For a VRP loop that needs a fast, offline travel-time estimate, the residual network remains the right tool. It reaches 9.53% overall MAPE against the rational curve's 13.70%, a lead that holds at every distance range we tested. Its inference is in the same class of speed as contraction hierarchies, not seconds. This is what our customer's planner now uses.

A general-purpose language model, asked to guess purely from memory, is a poor fit for short in-town trips: an order of magnitude worse than either baseline under 20 km. On medium and long trips it lands close to, and sometimes past, the rational curve.

For a customer planning a mix of short local stops and cross-country legs every month, that split matters more than any single headline number. An agent's intuition about a long inter-city drive deserves more trust than its intuition about a 5-minute hop across town. Neither should be trusted with the latency or the staleness risk of an LLM call inside an actual optimization loop.

Open questions

  • Would angular position relative to a city center, instead of absolute bearing, close more of the residual network's gap?
  • How much do the per-range language model numbers move with repeated calls, which this run did not make?
  • Do the results hold for another country's road network, another model, or another prompt?