Reliability Method

Reliability Engineering

Maintenance Strategy Development

Maintenance strategy development selects technically effective, risk-informed actions for asset failure modes and converts them into controlled maintenance, monitoring, redesign, or run-to-failure decisions.

Status: ApprovedDifficulty: AdvancedUpdated: 2026-08-08

Overview

Maintenance Strategy Development is the process used to determine how an asset, system, or failure mode should be managed so that required function is preserved at an acceptable level of risk and lifecycle cost.

A maintenance strategy is not simply a PM list. It is the logic that connects required function, functional failure, failure mode, consequence, criticality, detectability, technical feasibility, task selection, interval or trigger, corrective response, and ownership.

The objective is not to maximize maintenance activity. The objective is to select the least intrusive, technically effective, and economically justified action for the failure mode and operating context.

What It Is

A maintenance strategy defines:

  • what function must be protected;
  • how that function can fail;
  • what happens if it fails;
  • how significant the consequence is;
  • whether degradation can be detected;
  • whether a scheduled task is technically effective;
  • when the task should occur;
  • when run-to-failure is acceptable;
  • when redesign or procedural change is required.

Possible strategy outputs include:

  • time-based preventive maintenance;
  • usage-based preventive maintenance;
  • condition-based or predictive maintenance;
  • failure-finding;
  • scheduled restoration;
  • scheduled discard;
  • operator care;
  • corrective maintenance;
  • run-to-failure;
  • redesign;
  • operating controls;
  • spare-parts and contingency planning.

Why It Matters

A good maintenance strategy:

  • reduces unnecessary maintenance;
  • focuses effort on credible failure modes;
  • improves PM and PdM quality;
  • supports asset availability;
  • improves planning and scheduling;
  • strengthens spare-parts decisions;
  • reduces maintenance-induced failures;
  • makes run-to-failure decisions explicit;
  • identifies when maintenance cannot solve the problem and redesign is required.

Without a clear strategy, organizations often inherit OEM checklists, duplicate PMs, over-maintain low-risk assets, under-maintain critical failure modes, and continue repeating ineffective tasks.

When to Use It

Use Maintenance Strategy Development when:

  • commissioning new assets;
  • creating or rebuilding a PM program;
  • reviewing chronic failures;
  • implementing RCM or FMEA outputs;
  • deploying PdM;
  • reviewing high-criticality assets;
  • standardizing maintenance across similar equipment;
  • evaluating run-to-failure;
  • changing operating context;
  • introducing new technology or controls.

Do not use it as a substitute for work prioritization, job planning, scheduling, RCA, or capital-project selection.

Core Principles

  1. Required function comes before maintenance task.
  2. Failure modes drive strategy selection.
  3. Consequence determines the rigor of the decision.
  4. Asset criticality informs where analysis effort should be concentrated.
  5. The least intrusive effective strategy is preferred.
  6. More maintenance is not automatically better.
  7. Run-to-failure can be a valid strategy.
  8. Hidden failures require failure-finding.
  9. Redesign is appropriate when maintenance cannot adequately control risk.
  10. OEM recommendations are inputs, not automatic final decisions.
  11. Operating context matters.
  12. Strategy effectiveness must be reviewed using failures and condition findings.

Process or Lifecycle

Process or Lifecycle diagram
  1. Define Required Function leads to Identify Functional Failures.
  2. Identify Functional Failures leads to Identify Failure Modes.
  3. Identify Failure Modes leads to Evaluate Consequences.
  4. Evaluate Consequences leads to Apply Criticality and Operating Context.
  5. Apply Criticality and Operating Context leads to Select Strategy.
  6. Select Strategy leads to Build Tasks and Job Plans.
  7. Build Tasks and Job Plans leads to Configure CMMS.
  8. Configure CMMS leads to Execute and Capture Results.
  9. Execute and Capture Results leads to Review Failures, Findings, and Cost.
  10. Review Failures, Findings, and Cost leads to Select Strategy.

A strategy that is never reviewed becomes an assumption rather than a controlled maintenance decision.

Roles and Responsibilities

RoleResponsibility
Reliability EngineerLeads strategy development and optimization
Maintenance PlannerConverts approved strategy into executable work
SchedulerIntegrates strategy-generated work into schedules
TechnicianExecutes tasks and provides condition feedback
Maintenance SupervisorVerifies execution and task quality
OperationsDefines operating context and supports access
EngineeringSupports failure analysis and redesign
EHSReviews safety and environmental consequence
Quality / Food SafetyReviews product and compliance consequence
CMMS AdministratorMaintains plans, task lists, counters, and master data
Maintenance ManagerOwns performance and strategy governance
Site LeaderResolves business-risk and resource conflicts

Required Inputs

Typical inputs include:

  • asset hierarchy;
  • required function;
  • performance standard;
  • operating context;
  • asset criticality;
  • failure history;
  • failure modes;
  • OEM information;
  • drawings;
  • process data;
  • condition-monitoring data;
  • regulatory requirements;
  • quality and food-safety requirements;
  • safety and environmental consequence;
  • spare-parts lead times;
  • labor capability;
  • shutdown opportunities;
  • redundancy;
  • contingency options;
  • lifecycle cost information.

Required Outputs

A completed strategy should define:

  • asset or asset class;
  • required function;
  • performance standard;
  • functional failure;
  • failure mode;
  • failure effect;
  • consequence;
  • criticality;
  • selected strategy;
  • task type;
  • interval or trigger;
  • technical basis;
  • acceptance criteria;
  • labor and material requirements;
  • owner;
  • review date.

Step-by-Step Implementation

1. Define the required function

Describe what the asset must do and the required performance standard.

Example:

Function: Deliver 450 gallons per minute of process water at 60 psi. Functional failure: Unable to deliver at least 450 gallons per minute at 60 psi when demanded.

2. Identify functional failures

Functional failures may be complete, partial, intermittent, hidden, degraded, unsafe, inefficient, quality-related, or unavailable on demand.

3. Identify credible failure modes

Failure modes should be technically specific and actionable.

Examples include bearing lubrication contamination, seal wear, impeller erosion, winding insulation breakdown, sensor drift, valve sticking, filter blockage, and belt fatigue.

4. Evaluate consequences

Consider safety, environmental, food safety, quality, regulatory, production, customer, financial, asset damage, recovery time, and redundancy.

5. Select the strategy

StrategyAppropriate when
Time-based PMFailure is meaningfully age-related
Usage-based PMFailure relates to hours, cycles, mileage, or throughput
Condition-based maintenanceDegradation can be detected with sufficient warning
Failure-findingFailure is hidden until demand
Scheduled restorationRestoration renews resistance to failure
Scheduled discardReplacement before a life limit is effective
Operator careSimple safe checks can be performed by operators
Corrective maintenanceDefect is identified before functional failure
Run-to-failureConsequence is acceptable and response is controlled
RedesignMaintenance cannot adequately control risk
Procedural controlOperating method materially influences failure
Spare/contingency strategyFast recovery is more effective than prevention

6. Set interval or trigger

Consider age, usage, P-F interval, failure consequence, environment, regulatory requirement, historical findings, lead time for corrective work, and shutdown constraints.

7. Build executable work

Translate strategy into job plans, task lists, material requirements, acceptance criteria, condition limits, work triggers, counters, and inspection routes.

8. Review effectiveness

Use work history, failures, findings, costs, and condition data to decide whether the strategy remains effective.

Decision Rules

Decision Rules diagram
  1. Failure Mode Identified leads to Consequence Acceptable?.
  2. Consequence Acceptable?, when Yes, leads to Effective Preventive or Predictive Task?.
  3. Consequence Acceptable?, when No, leads to Can Risk Be Controlled by Maintenance?.
  4. Effective Preventive or Predictive Task?, when No, leads to Run-to-Failure with Contingency.
  5. Effective Preventive or Predictive Task?, when Yes, leads to Select Least Intrusive Effective Task.
  6. Can Risk Be Controlled by Maintenance?, when Yes, leads to Select Least Intrusive Effective Task.
  7. Can Risk Be Controlled by Maintenance?, when No, leads to Redesign or Operating Change.
  8. Select Least Intrusive Effective Task leads to Set Interval or Trigger.
  9. Set Interval or Trigger leads to Build Job Plan and CMMS Controls.

Run-to-failure is appropriate only when

  • consequence is low;
  • failure is evident;
  • secondary damage is limited;
  • repair is straightforward;
  • spares are available;
  • downtime is acceptable;
  • no better task exists;
  • contingency is defined.

CMMS Considerations

A CMMS may need to support asset hierarchy, criticality, failure-mode coding, maintenance plans, task lists, counters, condition triggers, job plans, BOMs, strategy classification, regulatory flags, version control, work-order linkage, measurements, condition findings, and strategy review dates.

The CMMS should hold the executable controls. It should not become the place where strategy logic is improvised without technical basis.

SAP PM Considerations

The MKS identifies SAP PM concepts that may support strategy execution, including equipment, functional locations, maintenance plans, maintenance items, task lists, strategy plans, counters, measurement points, notifications, maintenance orders, catalog profiles, work centers, planner groups, BOMs, permits, revisions, and status management.

Exact configuration varies by SAP release and site design.

Metrics and KPIs

Strategy Coverage

Definition: Portion of in-scope assets or failure modes with an approved maintenance strategy.

Formula:

Strategy Coverage = Assets or Failure Modes with Approved Strategy ÷ Total In-Scope Assets or Failure Modes × 100

Units: Percent

Interpretation: Indicates how much of the intended population has documented strategy logic.

Limitations: High coverage does not prove strategy quality.

Failure-Mode Coverage

Definition: Portion of significant identified failure modes with an assigned strategy.

Units: Percent

Interpretation: Indicates technical completeness.

Limitations: Depends on the quality of failure-mode identification.

PM Effectiveness

Definition: Measures whether PM tasks are detecting or preventing the failure conditions they were designed to address.

Units: Organization-defined

Interpretation: Supports PM optimization.

Limitations: Requires reliable work-order and finding data.

PdM Conversion Rate

Definition: Portion of valid condition findings converted into controlled corrective work.

Units: Percent

Interpretation: Tests whether condition monitoring produces action.

Limitations: A high rate is not necessarily desirable if alarm quality is poor.

Repeat Failure Rate

Definition: Portion of failures that repeat within the defined analysis boundary.

Units: Percent

Interpretation: Indicates whether strategy and corrective actions are eliminating recurring problems.

Limitations: Requires consistent failure coding.

Common Failure Modes

  • Starting with PM tasks instead of functions and failure modes.
  • Copying OEM checklists without validating operating context.
  • Applying identical PM programs to all identical equipment.
  • Using criticality to select task type instead of failure behavior.
  • Treating run-to-failure as neglect rather than a deliberate decision.
  • Never considering redesign.
  • Using calendar PM on random failure modes.
  • Setting intervals without technical rationale.
  • Creating PdM routes with no corrective-response process.
  • Never reviewing the strategy after failures.
  • Updating strategy decisions without updating CMMS controls.

Best Practices

  • Start with required function.
  • Analyze significant failure modes.
  • Use criticality to decide analysis depth.
  • Select the least intrusive technically effective action.
  • Document run-to-failure decisions.
  • Escalate failure modes that maintenance cannot control to redesign.
  • Document interval logic.
  • Translate strategy into executable work.
  • Review failures against the current strategy.
  • Standardize similar asset classes only when operating context is comparable.
  • Use technician and operator feedback.
  • Keep strategy connected to PM, PdM, RCA, FMEA, and RCM.

Maturity Levels

LevelCharacteristics
1 — ReactiveMaintenance response begins after failure
2 — Calendar-BasedPMs exist but often lack failure-mode basis
3 — ControlledStrategy templates, criticality, and change control exist
4 — Reliability-BasedPM, PdM, RCM, FMEA, and RCA are integrated
5 — OptimizedEnterprise strategy libraries and continuous feedback are used

Real-World Example

A process-water pump is required to deliver 450 gpm at 60 psi.

A reliability review identifies several credible failure modes: bearing lubrication contamination, mechanical seal wear, impeller erosion, and winding insulation breakdown.

The team does not assign one generic monthly PM to all four failure modes. Instead, bearing condition is monitored using condition-based methods; seal leakage is inspected and trended; impeller performance is monitored through process performance; motor condition is monitored with applicable electrical and condition techniques; low-consequence consumable failures are evaluated for run-to-failure; and recurring installation-related failures trigger redesign or precision-maintenance action.

The result is a strategy package tied to failure behavior rather than a calendar checklist.

Audit Questions

  • Are required asset functions defined?
  • Are significant functional failures identified?
  • Are failure modes technically credible?
  • Are consequences evaluated?
  • Is asset criticality used appropriately?
  • Are time-based tasks limited to age-related failure behavior?
  • Are condition-based tasks supported by sufficient warning time?
  • Are hidden functions covered by failure-finding?
  • Are run-to-failure decisions explicit?
  • Are redesign triggers defined?
  • Are intervals supported by evidence?
  • Are strategy outputs translated into job plans and CMMS controls?
  • Are failures reviewed against the strategy?
  • Are strategy changes traceable?
  • Are similar assets standardized only when operating context is comparable?

Confirmed related topics:

  • Asset Criticality Analysis
  • Preventive Maintenance
  • Condition Monitoring & Predictive Maintenance
  • Maintenance Planning
  • Maintenance Scheduling
  • Failure Analysis & Failure Mechanisms
  • Failure Modes and Effects Analysis
  • Reliability-Centered Maintenance
  • Root Cause Analysis

Potential future assets:

  • Maintenance Strategy Development Worksheet
  • Strategy Selection Decision Tree
  • Run-to-Failure Decision Form
  • PM/PdM Selection Matrix
  • Maintenance Strategy Audit
  • Strategy Library Template
  • Strategy Review Checklist

These assets are not created by this document.

References

References are limited to the authoritative source and confirmed related Reliability Method sources:

  • RM-MKS-6009 — Maintenance Strategy Development
  • RM-MKS-6008 — Asset Criticality Analysis
  • RM-MKS-6002 — Maintenance Planning
  • RM-MKS-6003 — Maintenance Scheduling

Public derivative source: Draft. Publication status: Not Authorized.