VWTVWT
← All insights

Measuring On-Time Delivery Honestly, and Moving It

Published September 12, 2026 · 11 min read

Almost every operation has an on-time delivery number, and almost none of them can be compared with anyone else's. The number is a definition before it is a fact, and the definition is usually the part nobody wrote down.

That matters in two directions. Reporting a number that flatters the operation means the real problem never gets a project. And accepting a carrier's number without asking how it was built means paying for a performance that was never measured.

This is about one measure: getting it honest, and then moving it.

Six choices that set the number before any truck moves

Every on-time figure rests on six decisions. Change any one and the percentage changes without a single delivery changing.

  • Which promise is being kept. The date the customer asked for, the date you confirmed back, or the date on the dispatch plan. These are three different numbers and the last one is the easiest to hit.
  • What gets counted. A whole order, a line on an order, or one delivery. An order with ten lines is one failure or one out of ten depending on the choice.
  • Where the clock stops. Arrival at the gate, release from the dock, or the signature on the receipt.
  • How much tolerance. On the day, within a window, within an hour. And whether arriving early counts as on time or as a failure.
  • Whose fault is excluded. Most operations quietly remove the deliveries the customer caused, then compare the result with a number that did not remove them.
  • What sits in the denominator. This one is the largest and it is dealt with below.

Write these six down before arguing about the percentage. Two operations moving the same goods with the same trucks can honestly report figures five points apart, and the whole difference is on that list.

The Thai formula, and the thing it multiplies

Thailand has an official answer to what this measure looks like. The Ministry of Industry runs an industrial logistics performance index covering nine logistics activities in three dimensions, cost, time and reliability, for twenty-seven indicators in total. Factories assessed on it are benchmarked against a national database of more than 1,200 others.

The delivery reliability indicator is DIFOT, delivered in full and on time. The formula is not what most people assume. It is not the share of orders that were both complete and punctual. It is two separate fractions, multiplied:

DIFOT = (orders delivered complete ÷ orders delivered) × (orders delivered on time ÷ orders delivered) × 100

The department's own worked example uses a year of 6,340 delivered orders, of which 6,280 were complete and 6,300 were on time. That gives 6,280 ÷ 6,340 = 99.05% and 6,300 ÷ 6,340 = 99.37%, and the two multiplied give 98.43%.

Two things follow from the shape of that formula.

Shortages and lateness are the same score. Two lines each at 99% produce 98.01%, not 99%. An operation that fixes its punctuality and leaves its short shipments alone will watch the composite barely move and conclude the project failed.

In full means every line. The official note is explicit: if any one item on a purchase order cannot be delivered, the whole order is not complete. A ten-line order missing one line is a full failure, not a 10% one.

The denominator only holds orders that arrived

Here is the part worth knowing before quoting the number to anybody.

The department's own explanation states that only orders already delivered to the customer are counted. A company may have 10,000 orders on its books, and the measure ignores every one that has not yet reached the customer.

So the order that is three weeks late and still sitting in your warehouse is not a failure. It is not in the arithmetic at all. It joins the denominator on the day it finally goes out, and if it then arrives on the revised date, it can even count as on time.

Take a month of 1,000 confirmed orders. 960 are delivered inside the month and 40 are still undelivered at month end. Of the 960 delivered, 900 arrived on the date agreed and 930 arrived complete.

  • Scored the official way: (930 ÷ 960) × (900 ÷ 960) × 100 = 90.82%
  • Scored against every order confirmed: (930 ÷ 1,000) × (900 ÷ 1,000) × 100 = 83.70%

The gap is 7.12 points, and every point of it is an order nobody received. That is not an argument for abandoning the official measure. Report it, because it is what a Thai industrial benchmark will compare you against. But calculate the second number beside it every month, because the difference between the two is a pure count of the work that never left the building.

The same month, scored at the transport department

The same index carries a second version of the measure, one level down. The transportation indicator counts delivery instructions the transport department received, how many it delivered complete against the instruction, and how many it delivered by the date agreed.

Note what changed. The denominator is no longer the customer's order. It is the instruction handed to transport, and the clock is the date on that instruction.

Carry on with the same month. Transport received instructions covering all 960 delivered orders. It hit the required date on 930 of them and delivered 950 of them complete against the paperwork it was given. The 30 orders that arrived late at the customer but on time against the instruction are the orders that reached transport after the customer's date was already unreachable, so the instruction carried a later date from the start.

  • Transport department score: (950 ÷ 960) × (930 ÷ 960) × 100 = 95.87%

Three defensible numbers, one month, one set of trucks:

Scored as Result
Transport department, against its instructions 95.87%
Customer orders delivered, official denominator 90.82%
Every order confirmed, including those never shipped 83.70%

The gaps are the whole point.

  • 83.70 to 90.82, or 7.12 points: orders that never shipped.
  • 90.82 to 95.87, or 5.05 points: orders that reached the truck too late to save.
  • 95.87 to 100, or 4.13 points: what the vehicles and the road actually did.

Most operations only ever compute one of the three, argue about the last 4.13 points, and never see the 12.17 above it. Anyone pressing a carrier on a number while the top two rows go unmeasured is pushing on the smallest piece.

Where the official clock starts, and why it flatters

The same index also measures time, and the boundaries are worth copying or deliberately rejecting.

The order cycle indicator counts days from the company confirming receipt of the customer's order until the customer receives the goods. The clock starts at your confirmation. Everything that happens to an order between the customer sending it and your desk accepting it is outside the measure, so a slow order desk improves the score. If orders sit unacknowledged, measure from receipt as well, and the difference is a number nobody currently owns.

The delivery indicator counts days from the goods being loaded onto the vehicle, drawn from the warehouse, until the customer receives them. Two things fall out. It excludes the wait between the order being picked and the vehicle being loaded, which in a busy dispatch yard is not small. And it is counted in days, which is too coarse for any promise made in hours. A same-day or window-based promise needs its own clock, in minutes.

The handbook that defines these adds a consistency check worth running once a year: the order cycle time should not exceed the sum of the processing, procurement, handling, warehousing and delivery times, because they are the same process measured in pieces. If your total is longer than its own parts, there is time being spent that no department has claimed.

Where you stop the clock deserves the same care. A truck that reaches the gate at 13:40 and is released at 16:20 is punctual or nearly three hours late depending on which timestamp you write down, and both are defensible. Pick one, and know that the delivery window a customer sets is a reservation on a block of the vehicle's day rather than an arrival time, so a window that is never met is a fleet size problem before it is a punctuality problem.

The promise itself has to be built on how long the trip really takes. A date set from a map answer and not from the real door to door time on the lane produces a number that was never achievable, and no improvement project can rescue it.

Getting the arrival time out of the argument

None of this works if the arrival time is whatever the driver says it was.

In Thailand it does not have to be. Goods vehicles under the journey data recorder rules carry a device recording position, speed and time, transmitting to the Department of Land Transport at least every five minutes, with six months of data retained. That means the arrival and departure time of every trip for the past six months already exists, independently of the delivery note, and the same record is what settles an argument about waiting time at a gate.

Use it as the measurement source, not the driver's word and not the customer's complaint. A measure people can argue with will be argued with instead of improved.

What actually moves it

Two Thai factories, both assessed on this index, show the pattern.

A large milk production factory in the north serving over 500 customers on a two-day delivery cycle had its routes drawn by hand from planner experience. Transport ran at 9.5% of sales and DIFOT sat at 80%. The routes were rebuilt as a vehicle routing problem with capacity and time window constraints, solved in a spreadsheet. Cost fell 23.7% and DIFOT rose to 95%.

A medium-sized medical device trader in Chiang Mai, roughly 1,000 product codes, 450 customers and 1,300 orders a month, delivered with its own vehicles and assigned work on paper from experience. DIFOT sat at 85%. A transport management system took over order handling, job assignment, tracking and reporting. Management time fell 80%, cost fell 25%, traceability reached 100%, DIFOT rose to 97%, and the whole thing paid for itself in five months.

Neither bought a truck. Neither leaned on its drivers. Both changed a decision that was made before the vehicle moved, and in both cases the cost went down while the reliability went up. That is the opposite of the assumption sitting behind most of these discussions, which is that punctuality has to be bought. Below a certain standard it is not bought at all. It is the return on removing guesswork from the dispatch decision.

Where to stop

Reliability gets expensive near the top. The cost of providing a given service level rises steeply as it approaches 100%, so buying two points between 95 and 97 costs far more than buying two points between 70 and 72, and the customer may not notice the difference at all.

The composite makes this worse. Take five service lines running at 95% on time, 98% complete, 99% damage-free, 97% picked accurately and 94% invoiced accurately. Multiplied together, the share of orders with nothing wrong at all is 84%. Chasing the composite upwards means fighting on five fronts at once, and the costly front is rarely the one the customer is complaining about.

So pick the line the customer actually notices, set the target there, and leave the others measured but alone.

The procedure

  1. Write the definition down, all six choices, on one page. Date the page.
  2. Score last month three ways: transport against its instructions, delivered orders the official way, and every order confirmed. Compute the two gaps.
  3. Put every failure in one of three buckets: never shipped, reached transport too late to save, transport's own. The three buckets and the three gaps are the same split.
  4. Work the largest bucket first, which for most operations is not the third one.
  5. Measure how late, not only how often. Twenty minutes late and two days late are the same tick on a percentage, and they are not the same thing to the person receiving the goods. Count the hours.
  6. Freeze the definition for twelve months. Changing it resets the series, and an improvement that came from a redefinition is the one thing worse than no measurement at all.

The honest number that comes out of this is almost always worse than the one being reported today. That is the useful part. The three gaps are not a scorecard. They are a work list, in order, with the size of each job already attached.