Why Root Cause Analysis Matters

Root cause analysis (RCA) is the difference between fixing a problem once and fixing it over and over. I have seen the same broken drill, the same oversized hole, and the same surface finish problem recur because someone fixed the symptom instead of the cause.

A symptom is what you see. A broken drill is a symptom. The root cause is why the drill broke. Fixing the symptom means putting in a new drill. Finding and fixing the root cause means the new drill does not break again.

I have applied RCA to deep hole drilling problems for years. The approach works across every type of problem: tool breakage, hole deviation, surface defects, and process variation. The process is the same. The specific causes change.

The cost of not doing RCA is high. Repeated tool breakage costs time and materials. Repeated quality problems cost customer confidence. A single thorough RCA pays for itself many times over.

The RCA Process

Step 1: Define the Problem Clearly

A vague problem statement leads to a vague solution. I write down exactly what happened, when it happened, and on which job. A good problem statement includes measurable details.

  • Bad statement: The drill broke.
  • Good statement: A 10mm gun drill broke at 200mm depth on job 4172, 30 minutes into the shift on June 15. It was the third consecutive breakage at the same depth on the same job.

The good statement gives me a starting point. I know the depth, the tool, the job, and the frequency. That data focuses the investigation.

Step 2: Gather Data

I gather all available data about the event before I start guessing causes. The data includes the machine parameters at the time of the event, the tool condition, the material batch, and the coolant condition.

Data SourceWhat to Collect
Machine controlSpindle load, feed rate, coolant pressure at time of event
Tool recordsTool age, regrind count, previous issues
Material certHardness, grade, heat number
Coolant logPressure, temperature, filter condition
Previous eventsSimilar problems on other jobs or machines

I keep data collection sheets for every job. The sheets make it easy to gather the relevant information without having to remember what to check.

Step 3: Identify Possible Causes

With the data in hand, I list all possible causes that could produce the observed symptoms. I do not filter at this stage — every plausible cause goes on the list.

For a drill breaking at a consistent depth, possible causes include coolant pressure drop at that depth, a hard spot in the material at that location, chip packing at that depth, thermal expansion closing the chip clearance, and tool fatigue failure.

Step 4: Test the Most Likely Cause

I rank the possible causes by likelihood based on the data I collected. Then I test the most likely cause first. The test should be a simple check that confirms or eliminates the cause.

In the case of the three consecutive drill breakages at 200mm depth, the data showed coolant pressure dropping from 1000 psi to 600 psi at that depth. The most likely cause was a clogged coolant filter that caused pressure loss. Checking the filter pressure gauge confirmed a 40 psi pressure drop across the filter — double the normal value.

Step 5: Implement the Fix

Once the root cause is confirmed, I implement the fix. In the filter example, replacing the filter restored coolant pressure to 1000 psi. But the fix should also prevent recurrence.

I added a weekly filter pressure drop check to the preventive maintenance schedule. The check takes 2 minutes and catches filter clogging long before it causes drill breakage.

Fix ElementExample
Immediate corrective actionReplace coolant filter
Preventive actionWeekly filter pressure drop check
Trigger for actionReplace filter when pressure drop exceeds 25 psi

Step 6: Verify the Fix Worked

The last step is verification. I run the job under the same conditions and confirm the problem does not recur. One successful cycle is not enough — I monitor for at least 10 cycles or one full shift.

In the filter example, the job ran for a full week without a single breakage after the filter replacement and weekly check were implemented. The problem was solved.

Real-World Examples

Example 1: Tool Breakage at Fixed Depth

A customer called about consistent gun drill breakage at 150mm depth on a hydraulic cylinder job. The operator was replacing drills and running again without investigating.

I gathered the data. Coolant pressure dropped from 900 psi at entry to 500 psi at 150mm depth. The filter was clean. The coolant line had a kink at the point where the drill reached 150mm.

The root cause was a twisted coolant hose that pinched when the drill slid reached that position. The fix was rerouting the hose. The breakages stopped immediately.

Example 2: Inconsistent Surface Finish

A production job showed good surface finish for the first 20 parts each shift and then degraded. The operator adjusted feeds and speeds, which helped temporarily but the problem returned the next day.

The root cause was thermal drift. The machine was cold at shift start and the first 20 parts drilled well. As the machine warmed up, the spindle alignment shifted and the finish degraded.

The fix was a consistent warm-up procedure. After implementing 30-minute warm-up, the surface finish was consistent from the first part to the last.

Key Takeaways

  • Define the problem with measurable details before investigating.
  • Gather data from all sources before listing possible causes.
  • Rank causes by likelihood and test the most likely first.
  • Fix the root cause, not the symptom, to prevent recurrence.
  • Verify the fix over multiple cycles, not just one.
  • Add preventive measures to catch the problem before it happens again.