The Two-Cent Part That Shut Down a Factory for a Month

It wasn't one of the failure modes that anyone had thought to instrument.

I trained as an industrial engineer, and my first job out of university was at a large manufacturer, on a team that collected and analyzed data from sensors on industrial machines. The effort was straightforward and valuable. These machines ran expensive, high-volume production, and every day a machine sat idle, the customer lost serious money. Our job was to see a failure coming before it happened, so we could send a technician to fix the machine during planned downtime instead of losing days to a breakdown.

The system worked well. It watched the parts that tended to wear, learned the patterns that came before the common failures, and got good at predicting them. For the failures we had designed it to see, it did exactly what it was supposed to.

Then one day, a machine went down because a small bolt broke. The bolt cost about two cents. There was no sensor on it, no data stream watching it, nothing in the model that accounted for it. Why would there be? It rarely broke. It wasn't one of the failure modes anyone had thought to instrument, because instrumenting for a two-cent part that fails once in a blue moon makes no sense when you're focused on the expensive, predictable stuff.

So the system never saw it coming. It couldn't. That failure lived completely outside what the system was built to watch.

And here's where a two-cent problem turned into a catastrophe. The bolt was a specialized part, not something sitting in a drawer. It had to be ordered from a specific supplier, and while everyone waited for it to arrive, the machine sat dead. By the time the part came and a technician installed it, the customer's machine had been down for the better part of a month.

A two-cent component took a production line offline for weeks and cost a fortune in lost output. The cheapest, rarest failure in the whole system turned out to be the most expensive one that year.

I've spent the years since moving from industrial engineering into data and AI, and I've watched this same pattern play out in system after system, across industries that have nothing to do with printing or bolts. It's worth naming plainly, because it's costing manufacturers real money and it's almost invisible until it bites.

Every predictive system is built around the failures you expect. You instrument the parts you know wear out, you train the model on the patterns you've seen before, and you optimize for the common case. That's rational, and it works, right up until the failure that nobody modeled.

The problem is structural, not a mistake anyone made. A system tuned to catch the average, expected failure is blind by design to the rare one, because the rare one was never in the data it learned from. It doesn't show up as a warning. It shows up as a machine that's suddenly dead, with no alert that ever fired.

The deeper trap is that we tend to prioritize what to monitor based on how often something fails. That instinct is exactly backwards for the failures that hurt most. The cost of a breakdown has almost nothing to do with how often it happens. A two-cent bolt that fails once can cost far more than a wear part you replace on schedule every quarter because the damage isn't in the part; it's in how long the line stays down and how ready you are to respond.

Frequency tells you what to expect. It tells you nothing about what will actually hurt you.

So the question worth asking—before your next predictive maintenance investment or the next round of tuning your monitoring—is not, "Are we catching the failures we know about?" You probably are. The better question is "What could take a line down that we have no sensor for, and no plan to respond to?"

Walk the machine and look for the parts that aren't instrumented, not because they're unimportant, but because they seemed too small or too rare to bother with. Ask which of them, if they failed at the worst moment, you couldn't fix quickly, because the part is specialized or the response isn't ready.

That intersection, the failure you can't see and can't quickly recover from, is where your next month-long outage is hiding right now.

None of this means monitoring the common failures is wrong. It's necessary, and the systems that do it well earn their keep. But a system optimized only for the average, expected failure will always leave you exposed to the one you didn't think to watch—and on a factory floor, that exposure is measured in weeks of downtime and a number on a profit-and-loss statement that nobody saw coming.

The machine that's most likely to blindside you isn't the one failing in ways you understand. It's the one about to break in a way your system was never built to see.

About the Author

Michael Podgortsev

Michael Podgortsev

Director of Data and AI Strategy

Michael Podgortsev is a director of data and AI strategy and former CTO. He holds a degree in industrial engineering and began his career building predictive-maintenance systems from sensor data on industrial machines.

Sign up for our eNewsletters
Get the latest news and updates

Voice Your Opinion!

To join the conversation, and become an exclusive member of IndustryWeek, create an account today!