Motionwell Automation builds custom machines in Singapore and modernises the control systems on existing ones, so the question of how to reduce downtime in production usually reaches us as a specific machine. A multi-station rotary indexing assembly machine, where one jammed escapement holds every station on the dial. A palletizing cell where the robot runs its motion on its own controller beside the cell PLC, so the first diagnostic question after a stop is which of the two stopped first. Or an older machine whose controller is out of support, where the outage lasts as long as the hunt for a replacement card. Machines are designed, assembled and tested at our Woodlands Link facility, and the company has delivered more than 150 special purpose machines since 2014 under ISO 9001:2015 and bizSAFE Level 3.
The short answer. Unplanned downtime in manufacturing is not one problem. It is at least five, they have different owners, and only some of them are fixable by working on the machine. Jams, feeder faults, nuisance safety stops, slow recovery and long changeovers are machine problems. Material starvation, operator availability and scheduling are not, even though they land in the same column of the same report. The sequence that works is to log every stop with its trigger for a week before touching anything, split what you find into those two piles, and then spend money only on the machine pile. Everything below is the detail behind that.
This page is written from the builder’s side of the problem: which stops a machine can be designed not to have, which ones it can be designed to recover from quickly, and which ones are somebody else’s to fix. The arithmetic that turns stops into a percentage is in our explainer on what OEE actually measures.
What Is Downtime in Production, and Which Kind Do You Have?
It is five different things, usually, filed under one heading because they all look the same on a report: the machine was scheduled to run and it was not running. They separate cleanly once you ask who was standing at the machine and what they were waiting for.
Micro-stops are the ones nobody records, because recording them takes longer than fixing them: a jam at an escapement, a gripper that missed, a sensor that read a reflection, cleared by hand and running again inside a minute. Breakdowns are the opposite, and they are the stops everybody remembers, because something failed, the machine cannot run until it is repaired or replaced, and somebody had to be called. Changeover sits between the two: part of that time is planned and part of it is the overrun, and the two get reported together far more often than they get separated.
The last two belong to the line and never to the machine. Under starvation and blocking the machine is fine and there is nothing to feed it, or nowhere for its output to go. Under waiting the machine is fine and the material is there, and the person, the tool, the approval or the paperwork is not: waiting for a fitter, waiting for a first-article check, waiting for a cleaning sign-off.
Idle time is the everyday name for those last two, and it maps onto them instead of forming a sixth category: the machine is capable and available, and it is not producing because material, a person, a tool or an approval is missing. Reducing it is supply, buffering and shift structure work, and a machine builder reaches it only by taking the feeding of the machine on as a separate scope.
| Category | What it looks like at the machine | Who resolves it | Typical duration |
|---|---|---|---|
| Micro-stop | A jam, a missed pick, a false sensor read; operator clears it and presses start | Operator, on the spot | Seconds to a minute |
| Breakdown | A failed drive, a broken change part, a controller fault that will not reset | Maintenance, sometimes with a purchase | Hours to days |
| Changeover | Format parts swapped, recipe selected, first article approved | Setter, operator, quality | Minutes to hours, per product pair |
| Starvation or blocking | Machine ready, no material or no downstream space | Planning, logistics, the neighbouring machine | Whatever the upstream problem lasts |
| Waiting | Machine ready and material present, person or approval absent | Supervision and shift structure | Unpredictable, and rarely logged |
Two things follow from the table and they set up everything after it. The categories have different owners, so a single downtime number tells you nothing about who to talk to. And the durations differ by three orders of magnitude, so a list sorted by total time and a list sorted by frequency will not agree, and you need both.
Which Losses Respond to Machine Design, and Which Do Not?
This is the split worth making before any money is committed, because it decides whether the work is ours or yours.
A machine builder has levers on the mechanism, the controls, the recovery behaviour and the changeover design. A machine builder has no lever at all on whether material arrives, whether a second operator exists on night shift, or whether the schedule puts three product changes into one shift. Both kinds of loss land in the same availability column, which is exactly why the column is a poor place to start.
| Loss | Does machine work move it? | What actually moves it |
|---|---|---|
| Jams at a feeder or escapement | Yes | Escapement sized for the part, correct track geometry, orientation verification before the pick |
| Missed picks and dropped parts | Yes | Gripper geometry for the part orientations that occur, and sensing that confirms the pick |
| Nuisance safety stops | Yes | Layout and zone design, so the stop is not triggered by normal work |
| Slow recovery after any stop | Yes | Recovery sequences and alarm text designed in, covered below |
| Long or fiddly changeover | Yes | Tool-free format parts, keyed so no re-teach is needed, recipe-driven adjustment |
| Reduced speed running | Sometimes | Find the one station that could not hold tolerance at rated speed and fix that station |
| Breakdowns from obsolete controls | Yes, as a project | Modernisation, because the outage lasts as long as the hunt for a replacement card |
| Material starvation | No | Supply, buffering, or automating the feeding of the machine as a separate scope |
| Operator or fitter availability | No | Shift structure and staffing |
| Scheduling and product mix | No | Planning, though changeover design changes what the schedule can afford |
Two rows deserve a note.
Reduced-speed running is in the middle because it is usually a symptom. Machines get slowed for a reason, and the reason is often one station that could not hold tolerance at the rated cycle. Slowing the whole machine is the cheapest available workaround, and it stays in place for years because it works. Fixing the station is machine work; discovering that the station is the reason is measurement work.
Material starvation is at the bottom because it is not a machine fault, and it still lands on the machine’s numbers. It can be attacked, but as a different scope: the four-year QA laboratory automation programme we run uses AMRs, a cobot and server-based task scheduling largely so that instruments are kept fed without a technician standing there. That is a materials-handling project that happens to improve an availability figure.
Why Do the Stops Nobody Logs Cost More Than the Ones Everybody Remembers?
Because they are individually trivial, collectively large, and structurally invisible.
A jam cleared in fifteen seconds, forty times a shift, is ten minutes of stopped machine. Nobody writes it down, because writing it down takes longer than clearing it. Worse, most collection schemes only log a stop once it exceeds a threshold, and that threshold is often set at sixty seconds. Every one of those forty stops falls below it. The time does not disappear from the output figure, but it disappears from the downtime list and reappears as a performance shortfall that nobody can locate, because performance shortfalls have no reason code attached.
The fix is a measurement setting, not an engineering change: set the threshold at one second, record the trigger sensor or alarm bit with each stop, run it for a week and sort by frequency. What comes back is usually short and dull. One feeder escapement. One gripper losing a particular part orientation. One sensor mounted where it catches a reflection off a passing part. Three items, each of which is an afternoon of work, and together they can be most of what the machine loses in an hour.
Feeding deserves separate attention because it is where a lot of these live. A vibratory bowl with an escapement sized for the part, correct track geometry and orientation verification before the pick is a solved problem; the same bowl asked to handle a part variant it was never tooled for is a jam generator. A feeder that presents a wrong-way part every fiftieth cycle costs more than it looks like it should, because each occurrence is a stop, a clear, a restart and sometimes a scrapped part. How feeders are matched to parts is on our part feeding and presentation page.
What Makes a Stop Recoverable by the Person Who Found It?
Recovery is a design property, not an operating procedure, and it is the part of downtime work that gets left out of specifications entirely.
Consider what has to be true for a night-shift operator to clear a jam and restart without calling anybody. They have to know which station stopped and why. They have to be able to reach the jam. They have to be able to remove the part without a tool they do not have. The machine has to be able to resume from wherever it stopped, with no full home cycle that scraps everything in process. And the guard they had to open has to close and reset without a sequence nobody remembers.
Each of those is a decision taken during design or during a controls rebuild, and each has a cheap version that looks fine at acceptance and costs output for the rest of the machine’s life.
| Recovery requirement | Designed in | The cheap version, and what it costs later |
|---|---|---|
| The operator knows what stopped | Alarm names the station, the device and the condition, mapped to a downtime reason at the controller | A numeric fault code and a laminated sheet, so every stop starts with a search |
| The jam can be reached | Access designed for the parts that actually jam, with the guard opening where the hand has to go | Access through a panel held by fasteners, so clearing needs a tool trolley |
| The part comes out without damage | Manual release or jog on the axis that is holding it | Power off and lever it out, which is how change parts get bent |
| The machine resumes from where it stopped | Sequence written so every station can be recovered from any state | Home-all-axes only, so each micro-stop scraps the parts in process |
| The guard closes and the cell resets | One reset device, sited outside the hazard zone with a full view of it | A reset sequence across two panels that only the day shift knows |
The row that matters most is the fourth. Writing a sequence so the machine can be recovered from any state is engineering effort at the program stage, and it is the difference between a fifteen-second stop and a five-minute one. On a modernisation this is the moment to buy it, because the program is being written anyway; the reasoning about why carrying old logic forward defeats this is on our PLC migration and upgrade page.
The failure mode we plan against is a cell that stops safely but can only be cleared by somebody who is not on the floor at the time. That arrangement does not survive contact with a three-shift operation. What happens instead is improvisation: a door propped, an interlock defeated, a machine run in manual for a whole shift because manual is the mode that lets people reach the problem. Every one of those is a worse outcome than the stop it was working around, and none of them appear in any report until something goes wrong. Designing the sanctioned recovery to be faster than the improvised one is the only control that holds.
Why Does a Safe Stop Turn Into a Defeated Interlock?
Because a safety function that fires during normal work is indistinguishable, from the operator’s side, from a machine that keeps breaking.
The pattern is specific and it repeats. Light curtain trips because operators have to reach across a conveyor to do their job. An interlocked door opened several times a shift for a routine reach-in that the layout never accounted for. A laser scanner zone set wider than the actual hazard, so people walking past a cell stop it. In each case the safety device is working exactly as designed, and the design was made without watching how the work is really done.
The fix is layout and zone design, and it is never an override. Reposition the guard so the reach is outside the detection field. Add a properly rated muting arrangement where material has to pass through a curtain. Convert a full stop into a safe reduced-speed state where the task needs presence. We build safety doors, interlocks and laser scanner systems with LVD and CE testing on every machine, and a well-zoned cell shows up directly in the availability column without anybody calling it a downtime project.
Two facts constrain how far that can be pushed. Safety distances follow from measured stopping performance, and stopping performance is measured on the built machine, which is why a control system change re-opens the guard positions in both directions. And the calculations behind each safety function are made against a current standard edition: ISO 13849-1:2023 is what a present-day design is argued against, and a design still documented to the 2015 edition will need its performance levels restated when the machine is reassessed. Neither of those is optional to make a stop less annoying.
Which Downtime Is Obligatory, and What Makes Spares an Availability Question?
Some downtime is not a failure. Cleaning, calibration, planned maintenance and the planned portion of a changeover are time the machine is not producing because you decided it should not be. The builder’s lever on obligatory downtime is design: access for cleaning, format parts that come off without tools and go back on keyed so no re-teach is required, recipe-driven adjustment, and a first-article check built into the sequence so approval is not a separate errand.
Spares are the other half of this, and they are where an availability problem hides as a procurement problem.
An obsolete controller does not cause downtime until the day a card fails. Then the outage is not as long as the repair; it is as long as the hunt for a replacement part. That is a lead-time and stock argument you can quantify, and it is a better business case for modernisation than obsolescence anxiety on its own. The same logic applies further down: sensors and safety devices whose replacements come in a different body need a new bracket, which is fabrication work discovered during a shutdown.
| Spares exposure | What makes it a downtime risk | What removes the risk |
|---|---|---|
| Controller, I/O modules, communications cards | Lifecycle status is published per catalogue number, and they run out at different times | Ask the vendor in writing for each part separately, and keep the answer with the handover pack |
| Program source and passwords | A password nobody recorded, a dead memory backup battery, a programming cable that has gone missing, or a tool that only runs on an operating system no laptop still has | Confirm the file on the shelf matches the controller, and archive it somewhere that survives staff changes |
| Drives and motors | A modern drive on old mechanics does not reproduce old behaviour; acceleration, following error and run-down all change | Tune on the machine with the real load, and record the parameters as delivered |
| Field devices | Sourcing and sinking inputs are not interchangeable, and an analogue loop wired for a card nobody sells becomes a conversion | Keep the I/O list with the actual device named at the end of each point |
| Mechanical wear parts | The ones that wear are rarely the ones on the spares list | Identify them from the machine, not from the bill of materials |
The second row carries a warning worth stating plainly, because it converts a repairable situation into an unrecoverable one. Where a controller holds the only copy of the logic and cannot be uploaded, do not power it down. A machine that has run for fifteen years can be switched off permanently by accident, and the recovery work after that is tracing field wiring and reconstructing the sequence from observed behaviour. What each of these states costs to recover is set out on our machine retrofit and modernisation page.
There is a documentation item that belongs here and not in a paperwork section. The handover pack that ships with a machine or a modernisation includes the alarm list with each alarm mapped to a stop reason, the I/O list, a program archive, spare parts lists and a component list with firmware versions. Those are not administrative extras. They are the reason a small fault stays a small fault, because on a machine that nothing describes accurately, diagnosis is the long part of the outage.
How Do You Measure Stops Before Changing Anything?
By deciding what a signal means before it is wired, and by accepting a smaller dataset you can trust over a rich one you cannot.
What you can learn about stops depends on how much of the control system you are allowed to touch. Reason codes are the thing that separates a useful downtime record from a total, and they only exist where the controller can tell you why it stopped.
| Where the signal comes from | What you learn about a stop | What it will not tell you |
|---|---|---|
| The controller’s own tags and alarm words | That it stopped, when, and which alarm or interlock caused it | Anything the program never computed |
| A protocol gateway on an older controller | Most of the above, without rewriting the program | Values the controller holds but does not publish on that interface |
| Bolt-on sensing: current, vibration, a stack light tap | That it stopped, and roughly how often | Why it stopped, and which product was running |
| A hardwired dry contact off an existing circuit | Running, stopped, fault | Reasons, counts, product identity |
Three disciplines decide whether the numbers survive contact with an argument.
Say what the contact means before it is wired. A contact off a main contactor says the machine is energised. A contact off the cycle-start latch says something much closer to producing. Those two produce very different availability figures from the same machine, and the difference is invisible once it is in a database. Write the sentence, check it against the circuit, and get maintenance to agree with it.
Know how each inferred signal lies. A current transformer reads a machine that idles hydraulics, vacuum or heaters as running while it produces nothing. An accelerometer picks up the machine next door through the floor. A stack light tap reads a local convention, and a flashing amber can mean three different things on three machines in one bay. Bolt-on sensing is a legitimate rung, and the baselining day where somebody stands at the machine and compares reality against the inference is what makes it trustworthy. The full ladder, and which machine belongs on which rung, is on our machine data acquisition page.
Split the setup bucket before you collect it. A single category covering changeover absorbs the format change itself, waiting for a fitter, waiting for material, first-article approval and cleaning. Those five have different owners and different fixes, and nothing downstream can separate them later, because only the machine knows which condition it was in at the moment it was in it. Splitting them is a collection requirement.
One more thing decides whether two machines can be compared at all. Real-time clocks in controllers, HMIs and industrial PCs drift independently, and two devices on the same machine can be minutes apart after a year in service. If you want to know whether one machine waited for another, sync everything that produces a timestamp to one source and stamp events at the source. Otherwise the gap between two clocks is the floor on any conclusion you draw.
When Is Downtime Reduction the Wrong Thing to Buy?
The machine you want to fix is not the constraint. Lifting a station that is not limiting output produces inventory. Which machine to work on is a question about the line, and it is worth settling before any hardware is quoted.
A week with a clipboard would answer it. Where the question is narrow and current, a paper log and a stopwatch are faster and cheaper than a collection scheme, and they sometimes show that the answer was already known on the floor.
The mechanism is the actual complaint. New controls fitted to worn ways, tired bearings or a stretched drive train buy accuracy the mechanism cannot hold. A control system cannot compensate for backlash it cannot see, and a machine that fails the same way with better diagnostics is not an improvement.
Nothing is forcing a controls change. Spares available, no compliance driver, no data requirement, no product change the machine cannot follow. If a drive is faulting, an HMI is dead or a sensor is unreliable, replace what failed. The age of the controller is not by itself evidence that it is the fault.
Nobody owns the output. A downtime log with no named owner and no standing meeting produces a database. No amount of hardware quality prevents that, and it is a common reason these schemes go quiet in their second year.
The stops belong to the part, not the machine. A station jamming on components that arrived out of tolerance is doing its job, and the loss belongs upstream or to the supplier. Tightening the incoming specification, or moving the inspection one step earlier so the bad part never reaches the station, removes those stops; work on the station that found them does not.
What Should You Bring to Get an Answer on One Machine?
Six items, and they are ordered so that the first three decide what kind of project this is and the last three decide what it costs.
- Which machine, and why that one. In one sentence: what makes you believe this machine is where the output is being lost.
- Your current stop record, in whatever form it exists, including a whiteboard or a paper log. If none exists, say so; that changes the first step.
- The control platform: controller make, model and vintage, drive types, HMI, and whether program source is available and confirmed to match the machine.
- What a stop looks like on the floor. Who clears it, what they open, what tool they need, and whether the machine can resume or has to home.
- The changeover picture: how many product changes per week, how long each is allowed, and whether format parts need tools.
- The shutdown window, as hours and frequency, and whether the machine splits into stations that can be worked on separately.
With those we can say which of your losses are machine losses and which sit upstream of the machine, whether the next step is a week of logging or a controls project, and what the recovery design on this particular machine would have to change.
Talk to a Motionwell engineer with those six items and we can scope it from there.