A deep hole drilling machine that is down costs more per hour than almost any other machine in the shop. The hourly burden rate is high, the parts going through it are usually high-value, and the whole production schedule backs up when one machine stops. I have spent years learning what keeps these machines running, and most of what works is not in the maintenance manual.

Here is my approach to reliability — what I actually do, not what the textbook says.

Measuring Reliability: MTBF, MTTR, and OEE for Deep Hole Drilling

I track three metrics religiously: Mean Time Between Failures (MTBF), Mean Time To Repair (MTTR), and Overall Equipment Effectiveness (OEE). Each tells me something different about the machine’s reliability.

MTBF measures how long the machine runs between unplanned stops. For a deep hole drilling machine, I calculate it as total operating hours divided by the number of unplanned stops. A well-maintained gun drilling machine should run 300 to 500 hours between failures. A BTA machine running heavy cuts might be lower, around 200 to 350 hours, because the loads are higher and the coolant system works harder. When I see MTBF dropping below 200 hours, I know something systemic is wrong.

MTTR measures how long it takes to get the machine running again after a stop. This is where preparation matters. A rotating union seal replacement should take 45 minutes if you have the part and the tools ready. Without the part on the shelf, that same repair takes two to five days. My target MTTR is under two hours for common failures and under eight hours for anything else. I track MTTR separately for mechanical, electrical, and coolant-system failures because each category requires different response times and skill sets.

OEE combines availability, performance, and quality into a single number. For deep hole drilling, I find availability is the most impactful component. A machine that runs 95% of the time but produces scrap 5% of the time has a quality loss, but a machine that is down 20% of the time because of reliability problems is losing production that can never be recovered. My OEE target for a deep hole drilling machine is 85%, with availability being the primary focus.

I record every stop in a logbook that sits on the machine control panel. The operator writes down the time, the reason, and how long it took to fix. I review this log weekly. Patterns emerge that the daily firefight obscures — three coolant hose failures in two weeks tells me I need to check the hose routing, not just replace the hose.

Critical Failure Modes: What Fails Most Often and Why

After years of tracking failures across multiple machines, I have a clear picture of what breaks and how often. The coolant system dominates the failure list because it runs at high pressure continuously and handles contaminated fluid. Here is the breakdown by subsystem:

SubsystemFailure ModeFrequencyDowntime ImpactRoot Cause
Coolant systemRotating union seal failureHigh (every 3-6 months)1-4 hoursNormal wear, coolant contamination
Coolant systemHigh-pressure hose burstMedium (every 6-12 months)1-3 hoursPressure cycling fatigue, abrasion
Coolant systemPump seal leakMedium (every 8-14 months)2-6 hoursWear from fine chip particles
Coolant systemFilter cloggingHigh (weekly)30-60 minHigh chip load, oversized chip load
Spindle assemblyDrive belt failureMedium (every 12-18 months)1-2 hoursFatigue, oil contamination
Spindle assemblyBearing wearLow (every 3-5 years)8-24 hoursNormal wear, coolant ingress
Guide bushingOD/ID wearHigh (every 1-3 months)30 minNormal cutting wear
Way systemWiper seal failureMedium (every 6-12 months)2-4 hoursChip abrasion, coolant exposure
Way systemGib strip looseningLow (every 12-18 months)1-3 hoursVibration, normal wear
ElectricalDoor interlock failureMedium (every 6-12 months)30-60 minCoolant contamination of switch
ElectricalCable track fatigueLow (every 2-3 years)2-6 hoursContinuous flexing, chip damage
Hydraulic systemSeal leakageMedium (every 8-14 months)1-4 hoursNormal wear, contamination
Control systemSensor failureLow (every 1-2 years)1-4 hoursCoolant ingress, mechanical damage

The table tells a clear story: the coolant system is the machine’s weakest link. More than half of my unplanned stops come from coolant-related failures. The rotating union seal alone accounts for about 20% of all downtime events. I spend more time on coolant system maintenance than on everything else combined, and that allocation is correct.

Spindle bearing failures are rare but catastrophic — eight to 24 hours of downtime and thousands of dollars in parts. I detect bearing degradation early by monitoring spindle load and vibration during the warm-up cycle. A bearing that is starting to fail shows a 5-10% increase in no-load spindle power consumption. When I see that trend, I schedule the bearing replacement during a planned shutdown rather than waiting for the bearing to seize mid-production.

Preventive Maintenance Schedule Optimization

A factory-recommended PM schedule is a starting point, not the final answer. Every machine runs different materials, different depths, and different coolant conditions. The PM schedule needs to adapt to actual wear patterns.

I group PM tasks into three intervals:

Daily (operator-performed, 10-15 minutes): Coolant level and concentration check, visual inspection of hoses and fittings for leaks, check rotating union for drip, listen for unusual spindle or pump noise, verify way cover condition, check chip conveyor for jams. These checks catch about 40% of developing failures before they cause a stop.

Weekly (maintenance-performed, 1-2 hours): Coolant filter element inspection and replacement if needed, check and lubricate guide bushings, inspect drive belt tension and condition, check way wiper seals for damage, verify coolant pressure at the spindle, check rotating union seal for wear patterns, inspect electrical cabinet for coolant ingress.

Quarterly (maintenance-performed, 4-8 hours): Replace rotating union seal (proactive, not reactive), inspect spindle bearings via vibration analysis and temperature check, check and adjust gib strips, inspect ball screw nut preload, clean coolant tank and inspect pump strainer, check all limit switches and sensors, verify emergency stop function, perform coolant analysis and replace if degraded.

The most important optimization I made was switching the rotating union seal from a reactive replacement to a proactive quarterly replacement. The seal costs $50 to $200 depending on the machine. The downtime to replace it proactively during a scheduled PM is about an hour. Replacing it reactively when it fails mid-cycle costs three to four hours of emergency downtime plus the potential for secondary damage if the leaking coolant reaches the spindle bearings or way covers. The math is simple.

I also schedule major PM activities around production gaps rather than calendar dates. If a machine has a three-day gap between jobs, I use that gap for the quarterly work rather than waiting for the calendar date. This approach has eliminated the conflict between production and maintenance schedules.

Spare Parts Management: What to Stock, What to Order

Spare parts management for deep hole drilling machines is different from standard CNC machines. The specialized parts have longer lead times, and the cost of downtime is higher. Here is my spare parts inventory strategy:

ItemLead TimeCriticalityRecommended Stock LevelEstimated Cost
Rotating union seal kit2-5 daysCritical2 per machine$300
High-pressure coolant hose (common sizes)1-3 daysCritical1 set per machine$400
Guide bushings (frequently used sizes)1-5 daysCritical3 each per size$600
Coolant filter elements1-3 daysHigh6 pack$500
Spindle drive belt2-5 daysHigh1 per machine$200
Door interlock switch2-5 daysHigh2 per machine$200
Coolant pump mechanical seal3-10 daysHigh1 per machine$250
Way wiper seals (set)5-15 daysMedium1 set per machine$150
O-ring and seal kit (assorted)1-2 daysMedium1 kit$100
Spindle bearings6-12 weeksLow (order ahead)1 set on shelf$1,200
Ball screw assembly4-8 weeksLow (order ahead)N/A (order on need)$2,500+
Control board/module2-6 weeksLow (order ahead)1 for critical machines$1,500+
Hydraulic cylinder seal kit2-4 weeksMedium1 per cylinder type$150

The dividing line between “stock” and “order” is downtime cost. If a part fails and the downtime cost exceeds the part cost plus the cost of holding inventory, stock it. I walk through this calculation for every part.

Two exceptions to the rule: spindle bearings and ball screws. Both have long lead times (6-12 weeks and 4-8 weeks respectively), and both are expensive. I keep spindle bearings on the shelf for any machine that runs more than one shift. The cost is about $1,200 and the shelf life is essentially unlimited if stored properly. A failed spindle bearing on a two-shift machine without a replacement on hand means 8-12 weeks of downtime waiting for the bearing to arrive. That calculation is easy.

For ball screws, I take a different approach. I order a replacement when I see the first signs of wear — increasing position error, visible flat spots on the ball track, or backlash that exceeds the machine specification. This gives me enough lead time to have the part before the existing screw fails completely. The same logic applies to control boards. I stock one board for each machine model in the shop, but only if the board is not shared across multiple machines.

For a deeper breakdown of what I stock and why, see my article on Essential Spare Parts for Deep Hole Drilling Machines.

Operator Contribution to Reliability

Operators are the most underused resource in machine reliability. They spend eight to twelve hours a day with the machine. They hear the sounds, feel the vibrations, and see the coolant flow. No maintenance person or engineer has that depth of familiarity. I have built my reliability program around what the operator sees and hears.

The daily operator check covers five areas and takes ten minutes at the start of each shift: coolant level and appearance, rotating union drip condition, hose and fitting visual inspection, spindle warm-up and listen cycle, and chip conveyor operation. I designed the check to be fast enough that operators actually do it, not so long they skip steps.

Operator reporting is the backbone of my early warning system. I trained every operator to recognize three specific sounds — a bearing rumble that means the spindle is starting to wear, a pump cavitation sound that means the coolant level is low or the filter is clogged, and a high-frequency squeal from the rotating union that means the seal is about to fail. When an operator reports any of these sounds, I investigate immediately. Over the past year, operator reports have caught five failures before they caused a stop.

I also ask operators to log any unusual cutting conditions — an unexpected change in spindle load, a surface finish that looks different, or chips that are not breaking properly. These subtle changes often point to a machine problem before any alarm triggers. A gradual increase in spindle load over a week might mean the guide bushing is wearing tight, the coolant pressure is dropping, or the belt is slipping. Without operator input, I would not see that trend until the machine faulted.

Operators need to know that their reports lead to action. When an operator reports a problem and nothing happens, they stop reporting. I make sure every report gets a response, even if the response is “I checked it and it is within spec, keep monitoring.” The feedback loop keeps the information flowing.

For a complete walk-through of the daily inspection routine I use, see my Daily Machine Inspection Checklist.

Reliability Improvement Case Studies

Here are two cases where I applied the principles above and saw measurable improvement.

Case Study 1: Coolant System Reliability on a Three-Shift Gun Drilling Machine

The machine was a dedicated gun drilling cell running 24 hours a day, five days a week. MTBF was 180 hours, and the coolant system caused 70% of the failures. The rotating union seal was failing every 6 to 8 weeks, and hose bursts happened every 3 to 4 months.

I made three changes: switched the rotating union seal replacement from reactive to a proactive 8-week interval, replaced all high-pressure hoses with braided stainless steel lines that handle pressure cycling better, and installed a secondary coolant filter to reduce particle load on the main filter. The result was an MTBF improvement from 180 hours to 420 hours over six months. The proactive seal replacement added two hours of PM time every eight weeks and eliminated 12 hours of emergency downtime per quarter. The hose upgrade cost $800 and eliminated hose bursts entirely.

Case Study 2: Spindle Bearing Failure Prevention on a BTA Machine

The machine was a large BTA machine drilling 2-inch diameter holes in steel bar stock. The machine had been in service for eight years and the spindle bearings had never been replaced. I noticed a gradual increase in no-load spindle power — about 3% per quarter over two years. The vibration trend showed a similar increase. I scheduled a bearing replacement during a planned production gap. The old bearings showed visible spalling on the raceways and one cage had a crack. The replacement took 16 hours during a three-day production gap. Without the early detection and planned replacement, the bearing would likely have failed mid-production, causing at least 24 hours of emergency downtime plus the risk of spindle damage from bearing debris.

Key Takeaways

  • Track MTBF, MTTR, and OEE separately. MTBF tells you about failure frequency, MTTR tells you about your response capability, and OEE tells you the overall production impact.
  • The coolant system causes more than half of all unplanned stops on deep hole drilling machines. Focus your reliability efforts there first.
  • Switch rotating union seal replacement from reactive to proactive on a fixed schedule. The cost-benefit ratio is overwhelmingly in favor of proactive replacement.
  • Stock spare parts based on downtime cost, not part cost. Keep spindle bearings and critical seals on the shelf for any machine running more than one shift.
  • Operators are your best early warning system. Train them to recognize specific failure sounds and build a reporting process that gives them feedback within the same shift.
  • Planned bearing replacement based on trend data (spindle load and vibration) costs half the downtime of emergency bearing replacement.
  • SMAT RT in the PM schedule around production gaps rather than calendar dates to eliminate the maintenance-versus-production conflict.
  • For safety considerations that affect reliability decisions, see Deep Hole Drilling Machine Safety Standards.