RELIABILITYMETHOD

Reliability Engineering

Condition Monitoring & Predictive Maintenance

Condition Monitoring and Predictive Maintenance (PdM) is the systematic process of evaluating the health of physical assets using condition-based technologies, inspections, and data analysis to detect developing failures before functional failure occurs.

Status: PublishedDifficulty: BeginnerUpdated: 2026-08-03

Source and Scope Boundary

This page is jointly derived from approved internal sources `RM-MKS-7009` and `RM-MKS-6005`. `RM-MKS-7009` owns condition-monitoring technology, measurement, data quality, alarm criteria, and diagnostic method. `RM-MKS-6005` owns PdM strategy, finding-to-work workflow, governance, verification, and financial evaluation. A condition-monitoring alert is not automatically a diagnosis, work priority, remaining-useful-life estimate, or work order.

Definition

Condition Monitoring and Predictive Maintenance (PdM) is the systematic process of evaluating the health of physical assets using condition-based technologies, inspections, and data analysis to detect developing failures before functional failure occurs.

The objective is to perform maintenance only when asset condition indicates it is necessary while minimizing risk, downtime, and lifecycle cost.


Executive Summary

Traditional preventive maintenance assumes equipment requires service based on time or usage.

Predictive Maintenance uses actual asset condition to determine when intervention is required.

Without an effective condition monitoring program, organizations often experience:

  • Unexpected equipment failures
  • Unnecessary preventive maintenance
  • Increased maintenance costs
  • Reduced equipment availability
  • Poor maintenance planning
  • Missed early warning signs
  • Shortened asset life

An effective PdM program is designed to help organizations:

  • Detect defects early
  • Reduce exposure to some detectable failure modes when findings are acted on in time
  • Optimize maintenance intervals
  • Improve maintenance planning
  • Support data-driven decisions

Realized reliability, availability, emergency-work, and lifecycle-cost benefits depend on program maturity, technology-to-failure-mode fit, and disciplined work conversion (see "Program Performance Measurement" below); they are not automatic, and this page does not promise a specific downtime, cost, or reliability outcome.


Why Condition Monitoring Matters

Most equipment failures develop over time.

Condition Monitoring identifies degradation before it becomes functional failure, allowing maintenance to be planned instead of forced by breakdowns.


What Condition Monitoring & Predictive Maintenance Is

A complete program includes:

  • Asset selection
  • Technology selection
  • Inspection routes
  • Data collection
  • Data analysis
  • Alarm management
  • Work identification
  • Corrective action
  • Continuous improvement

What Condition Monitoring & Predictive Maintenance Is Not

Condition Monitoring is not:

  • Replacing all preventive maintenance
  • Purchasing expensive technology without a strategy
  • Collecting data without analysis
  • Predicting every failure
  • A one-time implementation project

PdM is a reliability strategy that complements other maintenance approaches.


Objectives

An effective program should aim to:

  • Detect defects early
  • Reduce exposure to some detectable failure modes when findings are acted on in time
  • Improve asset reliability
  • Reduce maintenance costs
  • Improve planning and scheduling
  • Extend asset life
  • Increase equipment availability
  • Reduce operational risk
  • Support continuous improvement

Condition Monitoring Philosophy

The best time to repair equipment is after a defect has been detected but before functional failure occurs.


Core Components

A complete PdM program includes:

  • Critical asset selection
  • Failure mode identification
  • Monitoring technology selection
  • Inspection procedures
  • Data analysis
  • Alarm management
  • Work management integration
  • Performance measurement
  • Continuous improvement

Relationship to the Maintenance Process

Condition Monitoring supports:

  • Reliability Engineering
  • Asset Criticality Analysis
  • Reliability-Centered Maintenance
  • Defect Elimination
  • Preventive Maintenance
  • Maintenance Planning
  • Maintenance Scheduling

Inputs

Successful PdM depends on:

  • Asset hierarchy
  • Criticality rankings
  • Failure history
  • FMEA results
  • Inspection data
  • Sensor data
  • CMMS history
  • Engineering expertise

Outputs

An effective program produces:

  • Early defect detection
  • Predictive work orders
  • Improved maintenance strategies
  • Reliability, availability, emergency-work-reduction, asset-life, and lifecycle-cost benefits that depend on program maturity, technology-to-failure-mode fit, and disciplined work conversion (see "Program Performance Measurement" for controlled measurement)

Selecting Assets for Condition Monitoring

Condition Monitoring should focus on assets where early defect detection provides meaningful business value.

Selection criteria include:

  • Asset criticality
  • Failure consequences
  • Failure frequency
  • Repair cost
  • Downtime cost
  • Safety risk
  • Environmental impact
  • Predictability of failure modes

Not every asset requires predictive maintenance.


Selecting Predictive Technologies

Choose technologies based on the asset and expected failure modes.

Common technologies include:

  • Vibration Analysis
  • Infrared Thermography
  • Ultrasound Inspection
  • Oil Analysis
  • Motor Circuit Analysis
  • Electrical Signature Analysis
  • Process Monitoring
  • Visual Inspection

The selected technology should be capable of detecting the failure before functional failure occurs.


Vibration Analysis

Vibration analysis is one of the most effective technologies for rotating equipment.

Typical applications include:

  • Bearings
  • Pumps
  • Motors
  • Fans
  • Gearboxes
  • Compressors
  • Turbines

Common defects detected include:

  • Bearing damage
  • Imbalance
  • Misalignment
  • Mechanical looseness
  • Gear defects

Infrared Thermography

Infrared inspections identify abnormal temperature conditions.

Applications include:

  • Electrical panels
  • Motors
  • Bearings
  • Steam systems
  • Insulation
  • Refractory systems

Thermography can identify abnormal thermal patterns associated with electrical, mechanical, insulation, or process conditions. Interpretation requires emissivity, reflected temperature, load, access, environment, instrument capability, and safety context.


Ultrasound Inspection

Ultrasound detects high-frequency sound generated by developing defects.

Applications include:

  • Bearing lubrication
  • Compressed air leaks
  • Steam traps
  • Electrical discharge
  • Vacuum leaks
  • Mechanical friction

Ultrasound can provide useful early evidence for selected friction, lubrication, leak, and electrical-discharge conditions. Relative lead time depends on the failure mode, method, interval, operating state, and detection criterion; it should not be assumed to precede vibration evidence universally.


Oil Analysis

Oil analysis evaluates lubricant condition and machine health.

Typical testing includes:

  • Wear metals
  • Particle counts
  • Water contamination
  • Viscosity
  • Oxidation
  • Additive depletion

Oil analysis can identify lubricant degradation, contamination, and wear-particle evidence when sampling and laboratory methods are controlled. It does not provide a universal or guaranteed lead time to failure.


Motor Circuit Analysis

Motor testing evaluates electrical health.

Applications include:

  • Insulation degradation
  • Rotor defects
  • Stator winding problems
  • Connection issues
  • Voltage imbalance

Electrical testing can reveal selected supply, connection, insulation, rotor, or stator conditions; which evidence appears first depends on the failure mode and method.


Inspection Routes

Inspection routes should define:

  • Assets inspected
  • Inspection frequency
  • Technology used
  • Alarm limits
  • Data collection procedures
  • Required documentation

Routes should be optimized for efficiency and consistency.


Alarm Management

Every technology should use documented alarm limits.

Alarm levels commonly include:

  • Normal
  • Alert
  • Alarm
  • Critical

Alarm criteria should be method- and failure-mode-specific and may use absolute values, change from baseline, statistical behavior, operating context, or diagnostic logic. Their basis, revision history, responder, and required action must be documented and reviewed as evidence develops.


Data Analysis

Collected data should be evaluated for:

  • Trends
  • Rate of change
  • Severity
  • Remaining useful life only where a validated prognostic model, uncertainty statement, and applicable operating assumptions support it
  • Repeat patterns
  • Correlation with operating conditions

Data becomes valuable only after it is interpreted correctly.


PdM Governance

Condition Monitoring and Predictive Maintenance should operate under documented governance with defined responsibilities, standardized inspection practices, and measurable objectives.

Governance should establish:

  • Program ownership
  • Technology standards
  • Inspection responsibilities
  • Alarm management procedures
  • Data quality requirements
  • Review frequency
  • Continuous improvement expectations

Technical governance under `RM-MKS-7009` controls methods, measurements, data quality, analyst competence, and diagnostic disposition. Program governance under `RM-MKS-6005` controls strategy, handoff to work management, priority assignment, verification, contractors, and financial evaluation. Neither authority silently absorbs the other.


Cross-Functional Collaboration

An effective PdM program requires participation from:

  • Reliability Engineering
  • Maintenance
  • Operations
  • Maintenance Planning
  • Maintenance Scheduling
  • Engineering
  • Production
  • Contractors and technology specialists, when appropriate

Successful programs depend on both technical expertise and operational support.


Predictive Work Management

Condition Monitoring findings should be integrated into the work management process.

Each identified defect should include:

  • Asset identification
  • Inspection findings
  • Technology used
  • Severity classification
  • Recommended corrective action
  • Target completion date
  • Responsible owner

Predictive work should be planned before functional failure occurs.


Work Prioritization

Not every predictive finding requires immediate action.

Prioritize work using:

  • Asset criticality
  • Defect severity
  • Failure progression rate
  • Safety consequences
  • Production impact
  • Planned outage opportunities

The objective is to perform work at the optimal time.


Program Performance Measurement

Use controlled metric contracts that define population, period, source, owner, states, exclusions, and missing-data treatment. Technical measures include collection compliance, detection-plan coverage, and actionable-alert precision. Workflow measures include finding-to-work conversion and verification completion. Financial estimates require a finance-controlled counterfactual and should be reported as ranges; do not label estimates as avoided failures without evidence. See the table below for controlled calculations; additional technical measures (e.g., missed-detection review rate) and workflow measures (e.g., decision lead time, repeat-finding rate) are defined in the governing sources `RM-MKS-7009` and `RM-MKS-6005`.


Case Study

The following is an illustrative composite drawn from common patterns across maintenance organizations, not a specific documented case.

A manufacturing facility implemented vibration analysis, infrared thermography, and oil analysis on critical rotating equipment.

The team would define the covered failure modes, monitoring points, measurement controls, alert dispositions, handoff rules, and repair-verification method before evaluating results. It would compare controlled technical and workflow measures over a suitable observation window and disclose operating, asset, staffing, and data-quality changes. This composite illustrates an evaluation method; it reports no measured outcome or guaranteed benefit.


Continuous Improvement

Improve the PdM program through:

  • Technology evaluations
  • Alarm refinement
  • Inspection route optimization
  • Reliability reviews
  • Failure investigations
  • Lessons learned
  • AI-assisted diagnostics

The program should continually evolve as equipment performance and organizational knowledge improve.


Knowledge Graph Updates

Future Knowledge Library topics introduced:

  • Vibration Analysis
  • Infrared Thermography
  • Ultrasound Inspection
  • Oil Analysis
  • Motor Circuit Analysis
  • Inspection Routes
  • Alarm Management
  • Predictive Analytics
  • Vibration Monitoring
  • Inspection Route Design
  • Predictive Data Analysis
  • Predictive Work Management
  • Inspection Route Optimization
  • Alarm Governance
  • Predictive Maintenance KPIs
  • Condition Monitoring Program Governance
  • AI-Assisted Diagnostics
  • Predictive Maintenance ROI
  • Continuous PdM Improvement

Industry Applications

Food Manufacturing

Condition Monitoring should prioritize:

  • Refrigeration systems
  • Packaging equipment
  • High-speed rotating assets
  • Utility systems
  • Food safety critical equipment
  • CIP equipment

Distribution and Warehousing

Priority applications include:

  • Conveyor systems
  • AS/RS equipment
  • Sortation systems
  • Motors and gearboxes
  • Dock equipment

Municipal Utilities

Focus on:

  • Pumps
  • Motors
  • Blowers
  • Lift stations
  • Electrical distribution
  • Backup generators

Commercial Facilities

Typical applications include:

  • HVAC equipment
  • Chillers
  • Cooling towers
  • Boilers
  • Emergency generators
  • Building automation systems

Small Manufacturing

Prioritize:

  • Production bottlenecks
  • Critical rotating equipment
  • Air compressors
  • Utility equipment
  • High-cost assets

Condition Monitoring for Small Business Owners

Small organizations should:

  • Monitor critical equipment first.
  • Begin with simple inspection routes.
  • Use affordable technologies where practical.
  • Trend inspection results.
  • Plan repairs before failures occur.

A simple predictive maintenance program is often more effective than an overly complex one that cannot be sustained.


Condition Monitoring & Predictive Maintenance Maturity Model

Level 1 — None / Reactive

  • Failures drive maintenance
  • Little or no condition monitoring
  • Limited equipment history

Level 2 — Ad Hoc / Developing

  • Basic inspections
  • Initial predictive technologies
  • Trending begins

Level 3 — Programmatic

  • Formal PdM program
  • Integrated work management
  • Technology standards
  • KPI reporting
  • Continuous improvement

Level 4 — Integrated

  • Enterprise condition monitoring
  • AI-assisted diagnostics
  • Predictive analytics
  • Digital asset health dashboards
  • Reliability-centered culture

Level 5 — Optimized / Predictive

  • Multi-technology evidence is integrated where justified
  • Validated prognostics are used only where technically defensible
  • Strategy changes follow engineering and program-governance review

Predictive Maintenance KPIs

MetricControlled calculationPrimary limitation
Collection complianceValid measurements collected within the approved window ÷ measurements due in the frozen point/period population × 100A completed but invalid sample is not compliant.
Detection-plan coverageIn-scope failure modes with a validated method, point, interval, criterion, and response owner ÷ in-scope failure modes selected for CM × 100Coverage does not establish detection probability.
Actionable-alert precisionAlerts confirmed as actionable conditions ÷ alerts fully reviewed × 100This is not a full false-positive statistic without known ground truth.
Finding-to-work conversionActionable findings requiring maintenance with a linked approved work record ÷ actionable findings requiring maintenance × 100Exclude findings dispositioned through a documented alternate path.
Verification completionApplicable post-work verifications completed ÷ completed PdM-originated work requiring verification × 100Completion alone does not prove defect removal.
Estimated economic effectApproved scenario-based benefits minus incremental program and corrective-action costsCounterfactual and downtime assumptions are uncertain; report ranges.

Common Mistakes

Organizations frequently:

  • Monitor noncritical assets first.
  • Collect data without analysis.
  • Ignore alarm conditions.
  • Delay corrective work.
  • Purchase technology without strategy.
  • Fail to trend results.
  • Separate PdM from planning and scheduling.
  • Never measure program performance.

Best Practices

  • Focus on critical assets.
  • Match technology to failure mode.
  • Establish clear alarm limits.
  • Convert findings into planned work.
  • Verify repair effectiveness.
  • Continuously optimize inspection routes.
  • Measure business value.
  • Integrate PdM into the reliability program.

Reliability Method Source Boundary

This page is governed by `RM-MKS-7009` for technical condition monitoring and `RM-MKS-6005` for PdM program strategy and work conversion. It does not assign future Reliability Method document identifiers or imply that proposed standards exist.


Product Opportunities

The items below are potential future product ideas for roadmap and planning purposes. They are not existing Reliability Method products, features, or services.

Templates

  • Inspection Route Worksheet
  • Asset Health Assessment
  • Predictive Finding Report
  • Alarm Limit Register
  • PdM Program Audit Checklist
  • Technology Selection Matrix

Calculators

  • Predictive Maintenance ROI Calculator
  • Remaining Useful Life Estimator
  • Inspection Interval Calculator
  • Downtime Avoidance Calculator

AI Tools

  • Predictive Maintenance Advisor
  • Asset Health Analyzer
  • Vibration Interpretation Assistant
  • Thermal Inspection Advisor
  • Reliability Trend Analyst

Facility Manager Features

  • Asset Health Dashboard
  • Inspection Route Manager
  • Predictive Findings Queue
  • Alarm Management Dashboard
  • Reliability Trend Analytics
  • Reliability Method Standards Library

  • Reliability Engineering
  • Reliability-Centered Maintenance
  • Asset Criticality Analysis
  • Defect Elimination
  • Preventive Maintenance
  • Maintenance Planning

References

  • `RM-MKS-7009`, *Condition Monitoring*, version 1.5, Approved Internal.
  • `RM-MKS-6005`, *Predictive Maintenance*, version 1.4, Approved Internal.
  • ISO 17359:2018, *Condition monitoring and diagnostics of machines — General guidelines*.
  • ISO 18436-2:2014, *Condition monitoring and diagnostics of machines — Requirements for qualification and assessment of personnel — Part 2: Vibration condition monitoring and diagnostics*.
  • ISO 13374 series, *Condition monitoring and diagnostics of machines — Data processing, communication and presentation*; confirm the applicable part and edition.
  • OSHA 29 CFR 1910.147, *The control of hazardous energy*; applicability depends on the task and jurisdiction.

Revision History

Version 1.0 Initial Condition Monitoring & Predictive Maintenance foundation created.

Version 1.1 Expanded predictive technologies, governance, and work management.

Version 1.2 Completed industry applications, maturity model, KPIs, Reliability Method Standards roadmap, product opportunities, references, and revision history.

Version 1.3 Reconciled to approved sources RM-MKS-7009 and RM-MKS-6005; clarified the authority split, bounded technical and outcome claims, aligned the five-level maturity model, added controlled metrics, removed unassigned roadmap identifiers, and replaced generic references.

Version 1.4 Independent review (`docs/quality/reviews/2026-08-03-claude-KN-5003-KN-8007-independent-publication-review.md`) found findings NEW-KN8007-01 (Objectives/Outputs sections retained unqualified benefit language inconsistent with the hedged Executive Summary/Case Study) and NEW-KN8007-02 (KPI-section prose named three measures the table did not define). This revision resolves both: hedged the "Prevent failures" objective and the reliability/asset-life/lifecycle-cost outputs, and narrowed the KPI prose to match the measures actually tabulated, pointing to the governing sources for the additional named measures. Claude authored this correction; a reviewer other than Claude must revalidate before any future publication authorization.

Version 1.5 Codex's independent revalidation (`docs/quality/reviews/2026-08-03-codex-KN-5003-KN-8007-nonblocking-findings-revalidation.md`) found NEW-KN8007-01 only partially resolved: the Executive Summary still listed unqualified "Improve reliability," "Increase asset availability," "Reduce emergency work," and "Lower lifecycle costs," and the Outputs section retained a bare "Reduced emergency work" bullet, inconsistent with `RM-MKS-6005` Section 5's boundary that realized value is not automatic and no specific outcome is promised. This revision resolves both: the Executive Summary now separates concrete process objectives from a single hedged benefits sentence, and the Outputs section folds the emergency-work claim into the same dependency-qualified statement used for reliability/availability/asset-life/lifecycle-cost benefits. Claude authored this correction; a reviewer other than Claude must revalidate before any future publication authorization.