01 Sep 2026
Your OTD says 94%. Your supplier says 99%. Neither of you is wrong.
On-time delivery has no single agreed definition. Two respected standards give opposite answers on which date counts. Here is what that does to a scorecard, and the one number almost nobody calculates.

By Matthew Spencer, Co-founder and CEO of FlockScore
Slide four of the quarterly review. You have the supplier at 94% on time. They arrive with a deck showing 99%.
Twenty minutes go on whose extract is correct. Somebody promises to align the data offline. The three deliveries that actually stopped the line never get discussed, and the meeting moves to the cost agenda.
Both numbers are usually right. That is what makes the argument unwinnable.
The standards contradict each other
Most people assume there is a definition somewhere, and that if both sides checked it the argument would end.
There is no single standard definition of on-time delivery, and two of the most widely used standards measure it against different dates.
TL 9000, the quality management standard published by the TIA QuEST Forum and used across telecoms, measures against the customer requested date. It states plainly that "the company's quoted lead time or ATP has no bearing on the TL OTD measurement", and that the supplier may not initiate a change to that date. SCOR, maintained by ASCM, has a level two reliability metric called Delivery Performance to Customer Commit Date. The date the supplier gave back.
A supplier that quotes sixteen weeks and delivers in sixteen weeks scores 100% under one standard and can score close to zero under the other.
CIPS does not attempt to resolve it either. Its guidance says simply that "there are no standard measurements that can be universally applied across several industries". The ambiguity goes all the way down to the benchmark data: APQC's figure, drawn from 4,648 organisations, puts the median for supplier on-time delivery at 90.0%, and defines on time as delivery "by the agreed delivery time".
Six choices, none of them about the supplier
Before the supplier does anything, the number has already been shaped by decisions inside your own business.
1. Which date you compare against. What you asked for, what they confirmed back, or a third date your ERP maintains separately from both.
2. Which physical event counts as the delivery. Goods leaving the supplier, or goods arriving at your dock. On long-haul lanes this alone is worth several points.
3. Which timestamp you use for that event. The lorry arriving, or somebody in receiving posting the goods receipt. If receiving posts in a Monday batch, every Thursday arrival is late by system definition.
4. How much lateness you allow, and whether early counts as a miss. voestalpine Automotive Components publishes its supplier rating: on time or one day early scores full marks, two days early scores 80, more than four days early scores 1. Early is a miss under TL 9000 too, unless the customer authorised it. Buyers who have not read their own rules tend to assume early is free.
5. What counts as one unit of measurement. An order line, a whole order, a quantity, or a value. A supplier who misses one line on a fifty line order is at 98% or at 0% depending on which you chose.
6. What happens to partial deliveries. TL 9000 requires the entire quantity of a line on the due date, or the whole line is late. Others score the proportion delivered. The automotive standard counts per part number per period, so a shortfall and its catch-up can land in different periods.
Six choices, each with several defensible answers. Your supplier picked a different combination, and had no reason not to.
Walmart proved the point in public, twice
Walmart's on-time in-full programme is the clearest demonstration available, because the definition kept changing while the suppliers did not.
In 2016 suppliers had a four day shipping window. In 2017 it was two days. By 2018 the window was gone, replaced by a must-arrive-by date where turning up early is a miss in its own right. In 2020 the threshold went to 98% across all categories, with a penalty of 3% of the cost of goods sold on non-compliant orders.
Then in February 2024 Walmart reversed, splitting the single 98% figure into two separate measures: 90% on time and 95% in full.
A supplier operating in exactly the same way throughout would have watched its score fall for eight years and then partly recover, without changing anything. Which tells you what a delivery definition actually is. Not a technical setting, but a commercial lever that moves in both directions and takes real money with it.
Can you even measure it the way the standard says?
LK03, the Odette and AIAG standard for automotive supply chain KPIs, ties the measurement point to the Incoterm. Under FCA or EXW, delivery accuracy is measured on dispatch at the supplier. Under DAP or DDP, on arrival at the customer. Same supplier, same shipment, different measurement point, decided by a commercial term someone negotiated years ago.
The standard is candid about where that data comes from. For FCA and EXW it notes that "as basis for the data in general the ASN is available". For DAP and DDP, "in general, the physical goods receipt is available as base data".
Which is the whole problem. Your goods receipt is yours and you always have it. The dispatch date belongs to the supplier, reaches you only if they send an advance shipping notice, and then only if it lands somewhere you can report on. In SAP that notice creates an inbound delivery, a separate document in logistics execution, while vendor evaluation reads the goods receipt. Getting the dispatch date into the score is a development, not a setting.
So most teams do one of two things. They normalise everything to goods receipt, which quietly adds their own inbound transit and receiving time to the score of every FCA and EXW supplier. Or they drop those Incoterms out of the measurement altogether.
Both are defensible. Neither tends to be written down anywhere the supplier can see it. And a supplier scored on a metric that silently includes your inbound freight will lose an argument it does not know it is having.
Your ERP already made these choices for you
Here is the part that tends to end the argument in the room.
In SAP's classic vendor evaluation, the delivery date variance score does not use the actual delivery date on the purchase order. It uses the statistics-relevant delivery date, a separate field that lets you move the expected date for planning while the scoring date stays where it was. Whether that field tracks the confirmed date or holds the original is itself a configuration decision, and one almost nobody in procurement knows has been taken.
The variance is then converted using a standardising parameter defining, in SAP's own words, "how many days correspond to a percentage variance of 100%". Somebody set that at go-live. It counts in working days if a calendar is configured, scores nothing unless a minimum delivery quantity check passes, and is smoothed against the previous score with a weighting factor.
So the number on the slide is a smoothed, scaled, calendar-adjusted comparison between a date maintained separately from the commitment and the moment somebody pressed post.
And in S/4HANA there are now two of them. Classic evaluation still runs the logic above. Alongside it sits an analytical KPI, Supplier Evaluation by Time, on a flat scale where each day of variance costs one percent, early and late count alike, and the window is year to date. Two supplier delivery scores, both shipped by SAP, calculated differently. Which one reaches your slide depends on which app somebody opened.
Those specifics are SAP. The pattern is not. Every system makes these choices somewhere, usually in configuration nobody has opened in years. It is worth an hour with whoever owns it, and in my experience that hour changes the conversation more than any amount of supplier escalation.
Before you blame the supplier, look at what you asked for
All of that is about how the number is computed. Underneath it sits a harder question: whether the miss was the supplier's at all.
The requested date is not usually a statement of what you needed. It is what the MRP run produced from a planned lead time somebody typed into the material master years ago. If that parameter is wrong, request-date OTD is measuring your own master data and calling it supplier performance.
Schedule stability is the same story. Move the call-off three weeks out, then score the supplier against the original date, and the miss is yours. TL 9000 encodes that asymmetry explicitly: only the customer may change the reference date. That rule is not neutral. It protects the buyer, and every supplier reading their scorecard knows it.
The number almost nobody calculates
Once you accept there is no single correct definition, the useful move is not to pick one. It is to run two, and look at the distance between them.
Against your requested date: did I get what I needed, when I needed it? That measures your supply chain.
Against their confirmed date: did they do what they said they would? That measures the supplier.
The gap between them is the diagnostic, and says more than either number alone.
| Confirmed date | Requested date | What it usually means |
|---|---|---|
| High | High | Genuinely reliable. Rarer than you would hope. |
| High | Low | They quote long and hit it. Either a lead time negotiation, or your planning parameter is wrong. Routinely misdiagnosed as a performance problem. |
| Low | High | They promise anything and it works out, because you hold stock or somebody expedites. Looks healthy. Fails without warning. |
| Low | Low | A real problem, and at least an unambiguous one. |
The third row is the one worth hunting. A supplier can sit there for years on a comfortable headline score, while the buffer quietly absorbing them gets cut for working capital reasons.
This is a reporting change, not a project. If you cannot get it built, pull a year of purchase order lines carrying both dates and do it once in a spreadsheet. The answer is usually uncomfortable enough that the reporting change funds itself.
Why the definition is worth more than the number
Write yours down. Which date, which event, which timestamp, which tolerance, which unit, how partials count, which Incoterms are in or out. Put it on the scorecard itself rather than in a procedure nobody opens.
That ends the argument on slide four. But it is not the real reason to bother.
A number without its definition cannot travel. It cannot be properly explained to the supplier being judged by it, set against how that supplier performs for anyone else, or fed into anything that needs to reason about reliability. It is a private figure in a private dialect, and every company in the chain is speaking a different one.
Automotive worked this out a while ago. That is what LK03 is for, and why it ships with standard messages for reporting delivery accuracy back to the supplier. Not because the metric is clever, but because a shared definition is the thing that lets performance data move between companies at all.
The rest of us mostly have the number and not the definition. Which is the wrong way round, because the definition is the part that makes the number worth anything to anyone else.
Sources: TL 9000 on-time delivery Q&A, TIA QuEST Forum · SCOR RL.1.1 Perfect Order Fulfilment, ASCM · Odette and AIAG, LK03 Key Performance Indicators for Automotive Supply Chain Management, v2.01 · CIPS, managing supplier performance · voestalpine Automotive Components supplier rating · Walmart OTIF requirements and Walmart's 2024 reduction · SAP, setting up vendor evaluation using the Logistics Information System and SAP, KPIs in supplier evaluation · APQC, percentage of supplier on-time delivery