Frame the equipment failure
Identify the asset, location, operating mode and observed failure. For example, “Conveyor #3 stopped following a motor overload trip” describes an event; “poor lubrication caused it” assumes a cause. Record when it happened, how long it lasted if known, and what was observed before restarting.
Preserve the actual condition
Capture alarm history, operating trends, inspection records, component condition and any relevant photos before evidence is lost. Contain immediate risk and record what was changed during recovery so that the investigation does not confuse a temporary fix with a validated explanation.
Test competing explanations
Use Fishbone to group potential contributors such as operating load, inspection process, equipment condition, material and measurements. Then use Why-Why on the most plausible branches. An overload alarm, for instance, may reflect a real mechanical load or an instrumentation problem; only the evidence can distinguish them.
Close the loop on reliability
Assign actions with owners and due dates. Define an effectiveness check that matches the failure mode: a follow-up measurement, inspection or observation under comparable operating conditions. Record the result rather than declaring a fix successful when the work order closes.