Blogs/Manufacturing insights

Manufacturing insights8 min read · Aug 31, 2026

Why high machine utilization can increase manufacturing lead time

High machine utilization increases lead time by inflating queue time. Above roughly 85% utilization, waiting stops rising gently and starts rising steeply.

Factory floor with machines running at high utilisation while jobs wait in queue

High machine utilization increases manufacturing lead time because a machine kept close to full load has no slack to absorb variability, so jobs wait in queue instead of being worked on. The machine looks more productive on the report. The order takes longer to reach the customer.

This is not an opinion. It falls out of queueing theory, which is the branch of applied mathematics that describes what happens when arrivals are irregular and service times vary. Trooba Flow, our Factory Flow Intelligence software, models a plant as an open queueing network for exactly this reason: to show what a utilization target will do to lead time before anyone changes a schedule.

The two things utilization actually buys you

First, definitions, because these words get used loosely on the shop floor.

Utilization

The fraction of available time a resource spends doing productive work. 6.8 hrs of an 8-hr shift is 85%.

Queue time

Time a job spends waiting at a resource before work begins. Nothing is happening to the part.

Work in process

The total number of jobs released into the plant but not yet finished.

Lead time

Elapsed calendar time from job release to completion — or MCT, door to door.

Higher utilization buys you two things. It spreads fixed cost over more output, and it raises throughput up to the point where the resource saturates. Both are real. Neither has anything to do with speed of delivery for an individual order.

What high utilization costs you is buffer. A machine at 98% has 2% of its time uncommitted. When a job arrives early, or a setup runs long, or an operator is pulled to another cell, there is nowhere for that disruption to go except into the queue behind it.

The mechanism: variability, utilization, time

The relationship between utilization and waiting is described by Kingman's formula, also called the VUT equation because it has three factors: variability, utilization, and time. In plain language:

Queue time ≈ variability factor × utilization factor × average process time

The utilization factor is the part that matters here. It is ρ / (1 − ρ), where ρ is utilization expressed as a decimal. Look at what that fraction does. At 50% utilization it equals 1. At 90% it equals 9. At 98% it equals 49. The denominator is shrinking toward zero, so the whole term runs away. This is why waiting time does not scale with utilization in any intuitive way. It bends.

The variability factor is the average of the squared coefficients of variation of arrivals and of service times. If arrivals are perfectly regular and every job takes exactly the same time, this factor is zero and there is no queue at all, at any utilization. That factory does not exist. In high-mix, low-volume work it is the opposite: mixed routings, uneven order sizes, variable setups, and rush jobs push the variability factor well above 1.

Variability and utilization multiply. Reducing one makes the other less punishing, and a plant with high variability is punished much earlier on the utilization curve than a plant with stable, repetitive flow.

Why the curve is the whole story

Most capacity discussions assume a straight line. Push utilization from 75% to 85% and you assume you have given up a little responsiveness for a lot of efficiency. Push it to 95% and you assume you have given up a little more.

You have not. The last five points of utilization cost more than the first fifty combined.

A worked example

The following numbers are illustrative, generated from Kingman's approximation, not measured in a specific plant.

Take one workcell. Average process time is 30 minutes per job, with a variability factor of 1.

Total time at the cell, by utilization · process time is 30 min in every column

75% 85% 90% 95% 98% 2.0h 3.3h 5.0h 10.0h 25.0h
UtilizationAvg. queue timeTotal time at the cellvs. 75% baseline
75%1.5 hr2.0 hr1.0×
85%2.8 hr3.3 hr1.7×
90%4.5 hr5.0 hr2.5×
95%9.5 hr10.0 hr5.0×
98%24.5 hr25.0 hr12.5×

Processing time never changed. It is 30 minutes in every row. Everything else is waiting. Going from 90% to 95% doubles the time a job spends at that cell. You did not add work. You did not slow the machine. You removed slack.

Now put that cell in a routing. If a typical job passes through eight operations and every one of them is held at 95%, the accumulated time is about 80 hours. Run the same eight operations at 75% and it is about 16 hours. Same machines, same process times, same people. The difference is entirely queue.

Raise the variability factor to 4, realistic for a job shop with irregular order arrivals and long, uneven setups, and every queue figure above multiplies by four. At 90% utilization that cell now holds a job for roughly 18.5 hours instead of 5.

Little's Law tells you where to look

WIP = Throughput × Lead Time

Rearranged, it becomes a diagnostic: Lead Time = WIP ÷ Throughput. It is one of the few results in operations that holds regardless of distribution, sequencing rule, or process detail.

If a plant completes 40 jobs a day and there are 400 jobs on the floor, average lead time is 10 days. To cut lead time to 5 days at the same throughput, WIP has to come down to 200 jobs. There is no third option.

This is why high utilization and high WIP always travel together. Keeping every machine loaded requires a queue in front of every machine, because an empty queue means an idle machine. So when a plant manager says "we need more WIP to keep the machines busy," that statement is correct — and it is also a statement that lead time will be long.

The measurement problem

Most plants are instrumented to see utilization and blind to queue time. Machine monitoring reports spindle hours. ERP reports operation completions. Neither one reports how long a job sat on a rack. That asymmetry produces a predictable failure: teams optimise the thing they can see.

Utilization-firstFlow-first (QRM)
Primary metricResource utilization, machine hoursMCT, queue time, WIP
Treats idle time asWaste to be eliminatedCapacity to absorb variability
Response to long lead timeAdd shifts, add machines, push harderReduce batch size, variability, WIP
Typical batch policyLarge batches to amortise setupSmall batches, smaller transfer lots
Effect on deliveryLead time grows quietlyLead time falls, often without capex

Neither column is morally superior. Utilization-first thinking is correct for a dedicated high-volume line running a stable product at predictable demand. It is the wrong model for high-mix, low-volume and make-to-order work, where the customer is buying responsiveness.

QRM's practical guidance is to plan critical resources at roughly 70% to 80% of capacity rather than at 90% or above, precisely so the utilization factor stays in the flat part of the curve.

A measured result

MEASURED RESULT — TARINIKA

Tarinika, a jewellery manufacturer, reduced manufacturing lead time using the same factory and the same machines. No machine was made faster; no capacity was added. The changes acted on batch and transfer-lot policy and on how work was released.

28 → 7days · >75% reduction

What to do instead

If lead time is the problem, work these in order.

1. Measure queue time

Walk a representative part and compare total elapsed time to total touch time.

2. Find what's above 85%

Resource by resource, including inspection, setup crews, and shared skills.

3. Attack variability first

Level order release and standardise routings before touching capacity.

4. Cut batch and transfer lot size

Let a downstream cell start before the upstream batch finishes.

5. Set a target, defend it

Decide the utilization critical resources will run at, and stop treating idle as failure.

6. Test before you change it

A routing or shift change moves load somewhere else — model it first.

Where Trooba Flow fits

Trooba Flow models your plant as an open queueing network: products, routings, resources, demand patterns, and variability. From that model it predicts where bottlenecks will form, how long queues will be, how much WIP will sit on the floor, what utilization each resource will run at, and what manufacturing lead time results.

That lets you ask the questions this article raises without running the experiment on live orders — a second shift on the press, a halved transfer lot, a shift in customer mix — as a projection, labelled as such every time, built from your data and established queueing relationships.

See what your own utilization targets are doing to your lead time.

Request a Flow Analysis

In short

Utilization and lead time are not aligned goals. They pull against each other, and the tension gets sharply worse above about 85% because the utilization term in Kingman's formula grows non-linearly as it approaches full load. Variability multiplies that effect, which is why high-mix plants feel it first and hardest.

Little's Law then makes the trade visible: the WIP required to keep machines fully loaded is the same WIP that sets your lead time. The practical move is to stop treating idle capacity as waste and start treating it as the thing that buys you speed.

If you want to see what your own utilization targets are doing to your lead time, you can request a Flow Analysis at trooba.com/flow-analysis.

FAQs

What is the ideal machine utilization rate for a job shop?

There is no single correct number, but Quick Response Manufacturing generally recommends planning critical resources at roughly 70% to 80% of capacity rather than 90% or higher. Plants with irregular order patterns and long, uneven setups need more reserved capacity than plants with steady, repetitive flow.

Does running machines at 100% utilization actually reduce cost per part?

It reduces the fixed-cost share allocated to each part at that operation, which is why it looks good on a costing report. It also raises WIP, queue time, and lead time across the whole routing. Those costs appear as expediting, obsolescence, and lost orders, usually recorded somewhere other than the operation that caused them.

Why does adding a machine sometimes not reduce lead time?

Because the queue may not be at that machine. Adding capacity where utilization was already moderate changes very little, while the real constraint stays where it was. Queues also form at shared people, inspection, and approval steps. Model the whole routing before buying equipment.

What is the difference between queue time and cycle time?

Queue time is the waiting portion only, the time a job sits before work starts. Cycle time is usually the total time at a step, including queue time, setup, run time, move time, and inspection. In high-mix plants, queue time typically makes up the large majority of cycle time.

How can lead time drop without adding capacity or speeding up machines?

By changing what happens between operations rather than during them. Smaller batches and transfer lots, controlled order release, lower WIP, and reduced variability all shrink queues without touching process speed. The Tarinika result — 28 days to 7 days — came from lot sizing and queue dynamics on the same factory and the same machines.

Is high utilization always a problem?

No. On a dedicated high-volume line producing a stable product against predictable demand, variability is low, so queues stay small even at high load. The problem is specific to high-mix, low-volume and make-to-order environments, where mix itself generates the variability that turns a high utilization target into long, unstable lead times.

See what your utilization targets are doing to your lead time.

Request a Flow Analysis