A hardware replacement SLA should describe the process, not only a time target. Define the starting event, required evidence, replacement criteria and who restores the service after the physical work.

RAID, a second port and a spare server reduce particular risks, but none creates recovery automatically. Every failure needs a tested path.

01

Define priority and the clock start

The timer may start from monitoring, ticket registration or engineer confirmation. Those moments can be far apart. For critical service, define the P1 channel and the evidence supplied in the first message.

  • server code and symptom start time;
  • IPMI/KVM reachability and power state;
  • OS, SMART, RAID or interface errors;
  • results from an external network;
  • changes immediately before the failure.
NextA drive failure includes data and rebuild
02

A drive failure includes data and rebuild

The engineer confirms the failed drive and replacement. Define the minimum class and capacity, whether the old device is returned or destroyed and who starts the array rebuild.

Replacing a drive does not restore data when RAID is absent or several drives fail. A successful rebuild also does not replace an external backup and application-level integrity check.

  • diagnosis and physical replacement target;
  • minimum model and endurance of the replacement;
  • owner of rebuild and monitoring;
  • handling of the removed media;
  • external backup used for data recovery.
NextA server failure needs a node-replacement plan
03

A server failure needs a node-replacement plan

A board, power or controller fault may take longer to repair than a move to a spare machine. State whether an equivalent node is available and how drives, addresses, VLANs and IPMI access move.

Define equivalence in advance. More cores may not replace the required CPU generation, memory type, PCIe lanes or controller.

NextA port failure requires end-to-end diagnosis
04

A port failure requires end-to-end diagnosis

The fault may be in the NIC, cable, optic, switch port, VLAN or upstream route. Replacing a cable without checking counters and link state does not complete diagnosis.

A second port helps only when redundancy is configured. Bonding or LACP needs switch support; two unrelated interfaces do not fail over by themselves.

  • link state and hardware error counters;
  • cable or optical module;
  • server NIC and driver;
  • switch port, VLAN and configuration;
  • external route and upstream reachability.
NextIPMI and backups cover different risks
05

IPMI and backups cover different risks

IPMI/KVM shows boot state and supports remote recovery, but it does not store application data. A backup restores data, but it does not replace console access or spare hardware.

Critical workloads need both mechanisms and a tested procedure: who starts recovery, where the copy is stored, how a replacement server is delivered and how the original design returns.

NextWhat the SLA should record
06

What the SLA should record

“Replacement in two hours” is incomplete without a clock start, NOC hours and exclusions. Separate acknowledgement, update, diagnosis and physical-replacement times. Include maintenance, remote hands, spares and escalation.

A service credit does not restore data or reduce downtime, but it makes the commitment measurable. Define the claim window, calculation and maximum credit.

  • NOC availability and operating hours;
  • P1/P2/P3 acknowledgement and update targets;
  • diagnosis and hardware replacement target;
  • exclusions, maintenance and data responsibility;
  • credit, claim window and escalation process.
NextGo to the topic questions