The Little Black Book of Reliability Management

On several occasions in the past, I have been cruising along doing all the things I thought were necessary to steadily improve reliability when out of the blue I was surprised by a problem I did not know existed. In those situations, I came to understand that having all the conventional reliability processes (RCM, RCA, Lifecycle Analysis, etc.) in place does not necessarily prevent surprises associated with day-to-day problems that cost the company a lot of money. Sometimes individuals working in plants, or in tightlyknit organizations, think it would be impossible for big-ticket problems to get past their awareness. That theory may be true, but there are plenty of examples to the contrary.
I recall one situation in which short, but chronic, outages were being caused by a single equipment item. This state of affairs was viewed as a nuisance and it was thought to result from a variety of different sources. When a total incident-tracking system was installed, it was found:
The source of all failures had a common cause.
Although the unit itself was down for only a short time, the outage drove product off-spec. It took a long time after re-start to get the product back onspec. It required even longer to re-run the off-spec material. All the time the off-spec material was being re-run, the fresh feed capacity was severely limited.
As it turned out, this "nuisance" problem was the most expensive problem confronting the plant. Until all the incidents...