Logistics
"95% ETA accuracy" is unfalsifiable — make it checkable
A percentage with no window and no horizon cannot be wrong. Specify hit rate within a stated window at a stated horizon, MAE in minutes, and coverage first.
- Published
- Reading
- 7 min
- Based on
- ACEA rFMS 5.0.0, project44's published ETA architecture and dwell/detention benchmarks; method rather than a delivered line — see the boundary at the end
A visibility vendor says 95% ETA accuracy. Your largest carrier says 95% ETA accuracy. The module in your TMS says 95% ETA accuracy. All three can stand at once and describe systems with nothing in common, because not one of them can be shown to be wrong.
The accuracy of a prediction is undefined until somebody says how close counts as right and how far ahead the question was asked. Omit both and the sentence survives all possible evidence — a claim with no truth conditions.
The specification is three numbers, not one
State it as: hit rate within ±W minutes, at a horizon of H hours before arrival, plus mean absolute error in minutes, per lane, measured against realised arrivals.
project44 publishes its own result roughly in that shape — +28 percentage points at a 10-hour horizon within a ±2-hour window, from an ensemble split between short-haul and long-haul at 200 km, over 150 input features drawn from ELD pings, EDI and API carrier feeds and a driver app, with an error-correction layer over a pseudo-deterministic drive-time baseline contributing 8–10 points on its own. It names its window and its horizon, which puts it ahead of most of the market. It is still a delta rather than a level: nobody tells you the hit rate before those 28 points, so you cannot tell whether the finished system lands at 60% or 90%.
Coverage is the denominator, and it is the number nobody volunteers
Accuracy is computed over the shipments that had a usable position feed. In a Bulgarian operation — over 77% of activity international, cross-trade alone above 44% — much of the fleet is subcontracted to small carriers with an aftermarket box, a phone, or nothing, and visibility platforms are weakest across exactly that long tail.
Those shipments do not go dark at random. A truck in a border queue with a driver who has stopped answering is simultaneously the least visible and the most late. Dropping it lifts the accuracy figure and degrades the product. Ask for the hit rate and the share of shipments it was computed over in the same sentence; without the second number the first is a claim about an unknown fraction of your freight.
The physical ceiling is in the ACEA specification, not in the model
ACEA’s rFMS 5.0.0, published 25 July 2025, is the manufacturer-neutral truck data API behind Volvo Connect, Scania Fleet Management, Fleetboard, MAN DigitalServices, DAF Connect and IVECO ON. Its guaranteed refresh floors are vehicle position at least once every 15 minutes and vehicle status at least once every 60.
At 85 km/h a fifteen-minute-old position is already 21 km stale before any model has made an error. No architecture undoes that. A 30-minute delivery window therefore cannot be served from OEM-native telematics alone at any claimed accuracy — the last leg needs driver-app pings or an aftermarket unit at higher frequency. That constraint sits in the specification, not in our assumptions, and anyone promising minute-level ETAs from manufacturer portals has not read it.
How the feed is consumed matters for the same reason — pagination by receivedDateTime + 1 second behind the moreDataAvailable flag, X-Rate-Limit headers respected, HTTP 429 backed off. A dropped page and a stationary truck look identical downstream, and an ingestion bug is indistinguishable from a model error unless somebody instrumented the gap.
Accuracy decays non-linearly with horizon, so one figure is a point on a curve you were not shown
At a 30-minute horizon an ETA mostly reports GPS. At ten hours it forecasts a border queue, where the driver places the 45-minute break under Reg. 561/2006, and dwell at every intermediate stop. Error does not grow with distance; it grows with the number of discrete events between now and arrival, each carrying its own variance.
So the useful request is not a number but a shape: hit rate at 1, 4, 10 and 24 hours, on the same shipments, in the same window. Flat and then a cliff means a static drive-time table plus a GPS position. Graceful decay means something is genuinely modelling rest and dwell. The 200 km split exists because the mechanisms differ — short-haul dominated by urban traffic and stop duration, long-haul by rest scheduling and border variance.
Dwell variance dominates multi-stop error, and it compounds
Target live-load dwell is 60–90 minutes; congested facilities routinely exceed three hours, largely because 60% or more of a site’s daily volume can land in the first three hours of the shift.
That distribution is right-skewed, with a commercial cliff in it: the standard convention is a two-hour free window, then €50–100 per started hour — Standgeld, seisuraha — priced in national transport conditions and CMR-based contracts, not by a regulator. Predict the mean of a right-skewed distribution and you get a system accurate on average and wrong on precisely the days that cost money.
On a multi-stop route the error at stop four is the accumulated dwell error at stops one to three plus drive time — which is why project44 attributes 35-plus points to a dedicated dwell-prediction module on multi-stop work, against 8–10 for the error-correction layer over drive time. The honest output is an interval: arrival between 14:10 and 15:40, 80% coverage, calibrated per lane. An interval can be scored. A point estimate can only be argued about.
Which timestamp counts as arrival moves the answer by an hour
Geofence entry, gate-in, dock-in, unload complete, POD signature. One truck, five defensible truths, separated by exactly the dwell the model was trying to predict. A vendor scoring against geofence crossing and a customer scoring against POD are not disagreeing about a model.
Two further definitions decide the result before any mathematics. Which prediction is scored — an ETA refreshing every five minutes produces hundreds per shipment, and scoring the last guarantees a beautiful number. And whether early counts as a hit: at a booked dock slot, four hours early is a refused entry and a driver burning duty time in the yard. Report the signed error distribution; the two tails have different prices.
Eight questions that turn the claim into something checkable
- Hit rate within what window, at what horizon, over which date range and lanes?
- What share of shipments had a usable feed, and how are the excluded ones counted?
- Which timestamp is treated as arrival?
- Which of the many predictions per shipment is scored?
- What is the signed error distribution — how much of the miss is early?
- MAE in minutes per lane, including the worst lane, not the average.
- What does week one look like on a new lane or a new subcontractor, with no history?
- Will you run in shadow mode against our realised arrivals for a month first?
A vendor who can answer these has instrumented their own product. The informative part is which question gets answered with a case study instead of a number.
Where this stops applying
A better ETA is worth nothing if no decision downstream changes because of it. If the receiving site cannot re-book a slot, if dispatch cannot re-tender, if customer service is never told, a tighter interval is a nicer number on a screen. Specify the decision first and the accuracy target follows — a two-hour window is enough for slot rebooking and useless for a cross-dock transfer.
And the target stops at the physical floor of the data. Spending on model architecture while the binding constraint is a 15-minute refresh and a third of loads on phone-only subcontractors is spending on the wrong half. Coverage, then window, then model.
One boundary about us. Palamed has not shipped a freight visibility platform. Our four deliverable engagements are a European car marketplace with 300,000+ listings, a platform for an AI automation agency, the Ministry of Education and Science dictionary at beron.mon.bg, and email-marketing automation for a beauty brand, where the roughly 60% reduction in outreach time is the client’s own number by the client’s own method. Everything above is the published pattern — the ACEA refresh floors, project44’s disclosed architecture, the dwell and detention conventions — plus what we would measure first. Which would not be a model. It would be reconstructing your realised arrivals from data you already hold, then plotting the hit-rate curve of the ETAs you run today. Often enough the incumbent beats the replacement being sold, and nobody has checked.
