Problem Overview
An LSI MegaRAID controller (current generation: Broadcom MegaRAID 8 Tri-Mode and MegaRAID Storage Manager) reports Consistency Check corrected medium error in the controller event log, and the affected physical drive accumulates thousands of media errors while the other disks in the same RAID 6 virtual drive remain at zero. This symptom pattern indicates an isolated physical drive fault, not an enclosure, backplane, or controller fault.
Event ID 0x0121 (Consistency Check corrected medium error) is logged when the consistency check (CC) patrol operation reads a logical block that fails its initial read, the controller reconstructs the missing block from the RAID 6 parity, and the reconstructed block is then successfully written back to the suspect physical drive. The firmware parameter string is documented as:
Consistency check corrected medium error (VD 0x%x at 0x%llx, PD 0x%02x(e0x%02x/s%d) Lun 0x%llx at 0x...
Decoding the fields:
- VD 0x%x — Target ID of the affected Virtual Drive.
- 0x%llx — Logical Block Address (LBA) within the VD where the corrected block lives.
- PD 0x%02x — Device ID of the Physical Drive that contained the bad block.
- e0x%02x/s%d — Enclosure ID and Slot number of that PD.
- Lun 0x%llx — Logical Unit Number and LBA on the PD.
Event ID 0x003A (Consistency Check done) is the closing record for that CC pass and is informational; it is logged whenever a consistency check completes, regardless of whether corrections were made. Its presence in the log only confirms the scan ran, not that errors occurred.
Key MegaRAID Event Codes
| Event ID | Severity | Description | Action Indicated |
|---|---|---|---|
| 0x003A | Information | Consistency Check done on %s | Scan completed, verify corrections |
| 0x0121 | Information | Consistency Check corrected medium error (VD/PD/LUN detail) | Identify PD, schedule replacement |
| 0x0040 | Warning | Consistency Check found inconsistency | Initiate Make Data Consistent |
| 0x004A | Information | Patrol Read corrected medium error | Track recurring PD |
| 0x0051 | Warning | PD predictive failure (SMART) | Replace immediately |
| 0x0071 | Critical | PD failed | Rebuild from spare |
The canonical references for these codes are Broadcom LSA Storage Authority Event Messages (2.7) and Broadcom MegaRAID 8 Tri-Mode Event Messages (1.0).
Root Cause Analysis
A single physical drive accumulating thousands of media errors (the source case shows 2,963) while every other member disk reads zero means the fault is localized to that drive. Three mechanisms can drive this count upward:
- Uncorrectable read errors on the magnetic surface — bad sectors that the drive firmware cannot recover internally and that are corrected by the RAID 6 parity engine on the host side. Each correction increments the Media Error counter in the PD properties.
- Latent sector errors (LSE) — sectors that were written correctly but now fail on read because the magnetic signal degraded over time. These are exactly what Patrol Read and Consistency Check are designed to catch.
- Mechanical or head-stack degradation — when the head cannot stabilize over a track. These typically present with a sharp rise in media errors alongside SMART 5 (Reallocated Sector Count), SMART 197 (Current Pending Sector), and SMART 198 (Offline Uncorrectable).
RAID 6 tolerates this failure mode silently for a long time. Each bad block is reconstructed from two independent parities (P and Q), the reconstructed block is rewritten to the suspect PD, and the operating system never observes a read error. The trade-off is read latency: every corrected read takes the latency hit of two extra reads plus a write, so a high media-error count measurably degrades array I/O performance even when no application sees an I/O failure.
Confirm the Fault is the PD and Not the Backplane
Before swapping the drive, eliminate enclosure and cabling causes:
- In MegaRAID Storage Manager (MSM) or LSA Storage Authority, open Physical View and select the suspect PD.
- Read Device Descriptor and confirm the WD Red Pro (or equivalent) model and firmware revision. If the firmware is more than 18 months old, update to the manufacturer's latest revision before replacing; some media-error storms are firmware bugs.
- Capture SMART Data: compare SMART 5, 187, 197, 198 against the data sheet. Non-zero raw values on a drive reporting only media errors confirm a media-level fault rather than a transport fault.
- Check Enclosure View for amber/blue SES LEDs on that slot. If the LED is solid, the backplane is reporting the fault; if the LED is off but MSM shows errors, the drive is reporting internally.
- Compare SAS link error counters across all PDs. Errors concentrated on one device point to the drive; errors spread across an enclosure row point to the expander.
Media Error Threshold Decision Matrix
| Media Error Count | SMART 197/198 | Predicted Failure | Action |
|---|---|---|---|
| 0 – 10 | 0 | No | Continue monitoring, log weekly |
| 10 – 100 | 0 | No | Increase Patrol Read frequency to daily |
| 100 – 1,000 | Non-zero | No | Order replacement drive, schedule swap |
| 1,000 – 5,000 | Non-zero | Likely | Replace within 72 hours, run Make Data Consistent first |
| > 5,000 | Non-zero | Yes | Replace immediately, treat PD as failed |
| Any | Yes | Yes (SMART trip) | Replace immediately regardless of count |
A count of 2,963 (as in the source case) lands in the third row: SMART data is non-zero on a degraded drive, predicted failure is likely but not yet declared by SMART, and the drive should be replaced within 72 hours under a planned maintenance window.
Replacement Procedure
- Verify backup integrity. RAID 6 is not a backup. Confirm that the most recent backup of each VD on the affected controller has been successfully tested by a trial restore before touching the array.
- Force a Consistency Check and Make Data Consistent. In MSM, right-click the VD and select Consistency Check. Wait for completion; this rewrites corrected blocks to all PDs and ensures parity matches data before rebuild.
- Identify the replacement target. Use the event 0x0121 PD ID and slot number from the log to confirm the physical bay. Mark the drive with a service tag if the chassis does not have illuminated fault LEDs.
- Mark the PD as offline. In MSM: right-click the suspect PD → Mark Physical Drive as Missing. This triggers the controller to use the dedicated hot spare (if present) or to leave the VD in a degraded state pending the new disk. Do not select Set Failed unless SMART has already tripped; forcing a failed state can cause unnecessary rebuilds.
- Physically remove and insert the replacement. Hot-swap is supported on all current MegaRAID HBAs. Wait 30 seconds after insertion for the controller to detect the new PD, complete its device discovery, and rebuild the foreign state table.
- Assign the new PD as a hot spare or as a replacement member. If a dedicated hot spare is configured, the rebuild starts automatically. Otherwise, right-click the VD → Rebuild and select the new PD.
- Monitor rebuild progress. Rebuild rate is configurable in the controller properties; do not exceed the rated workload of the disk. For a 4 TB NAS-class drive in RAID 6, expect roughly 6 – 14 hours depending on the rebuild rate setting.
Compatible Replacement Drive Selection
The source mentions difficulty sourcing WD Red Pro. Compatible replacements for a RAID 6 array that mixes vendors are acceptable as long as capacity matches or exceeds the original and sector size is identical (512e or 4Kn). Acceptable families:
- WD Red Pro (WD8005FFBX, WD102KFBX, WD142KFBX family)
- Seagate IronWolf Pro (ST16000NE000, ST18000NE000 family)
- WD DC Ultrastar (WUH721414ALE6L4, WUH722020ALE6L4 family) — datacenter grade, 5-year warranty, recommended when budget allows
- Toshiba N300 (HDWG480, HDWG460 family)
Verification Steps
- Confirm rebuild completed: VD Properties → State = Optimal, PD Properties → State = Online on every member.
- Verify no new 0x0121 events on any PD after rebuild. If the new drive logs 0x0121 within 24 hours, the backplane slot is suspect; reseat the drive and re-test.
- Run a forced Consistency Check on the VD and confirm completion with zero corrections (only 0x003A in the log, no 0x0121).
- Issue a scrub at the OS level:
scrub -a /mnt/array(ZFS) ore2fsck -f /dev/mapper/mpathX(ext4) — depending on the filesystem used. Read every block and let the controller surface any latent error. - Capture a fresh SMART snapshot from all PDs and archive it for trend analysis.
- Re-enable Patrol Read with a daily schedule if it was disabled.
Long-Term Monitoring Configuration
To prevent recurrence of silent media-error storms:
- Enable Patrol Read on the controller and set it to Automatic or Continuous with a daily start window outside peak load.
- Configure Consistency Check weekly, also outside peak hours.
- Export the event log via
storcli /c0 show eventsorperccli /c0 show eventsto a syslog target, then alert on any 0x0121 in the log. - Maintain either a global hot spare or a cold spare matched to the largest drive in any VD. A cold spare on the shelf is cheaper insurance than a multi-day replacement lead time when the original manufacturer is out of stock.
- Subscribe to drive vendor firmware notifications. Several recent media-error storms on NAS-class drives were fixed by firmware updates before any PD failed.
Data Corruption Risk Assessment
RAID 6 reconstructs any single-block read error using P and Q parity, which means the application layer sees the correct data even when one PD has uncorrectable sectors. The risk of silent data corruption arises in two narrow cases:
- Two physical drives fail overlapping regions before the first is replaced — unlikely in a 72-hour swap window but possible if a second drive has a hidden weakness.
- An error is written silently to parity and propagated across rebuilds. This is the failure mode that Consistency Check and Patrol Read are designed to detect; running both weekly closes this gap.
For applications where absolute data integrity is required (databases, archival storage), layer a checksum-aware filesystem (ZFS, Btrfs) on top of the RAID 6 VD so that any block that escapes parity correction is detected at the read path.
FAQ
What does event 0x0121 actually mean on an LSI MegaRAID controller?
Event 0x0121 (Consistency Check corrected medium error) is logged when the CC patrol read a block, the block failed internal recovery on the source PD, and the controller reconstructed the block from RAID 6 parity and rewrote it. The event string identifies the VD, LBA, PD ID, enclosure, and slot, allowing you to pinpoint the failing disk. See the Broadcom MegaRAID 8 Tri-Mode Event Messages reference.
Is 2,963 media errors on a single PD reason to replace the drive?
Yes. A drive that holds the entire media-error count of a VD while every other member is at zero has localized media degradation. Replace within 72 hours under a planned window, run Make Data Consistent first, and verify with a post-rebuild Consistency Check that logs only 0x003A with no 0x0121.
Can I keep using the drive until it fails outright?
Technically yes, RAID 6 will continue to correct bad reads with parity, but every corrected read costs the latency of two extra reads plus a write, so I/O throughput drops measurably. More importantly, a second simultaneous failure of another disk during this window collapses the array into a single-parity state and risks data loss.
Does RAID 6 already protect me against the data corruption from this drive?
Against single-disk bad-sector failures, yes. RAID 6 reconstructs each affected block from P and Q parity. It does not protect against two overlapping failures, and it does not detect a silently wrong parity write — that is why Patrol Read and Consistency Check must run weekly.
Can I substitute a different drive brand for the failed WD Red Pro?
Yes, as long as capacity is equal to or greater than the original member, sector size matches (512e with 512e, 4Kn with 4Kn), and the drive is on the MegaRAID compatibility list (or an equivalent enterprise NAS class such as Seagate IronWolf Pro, Toshiba N300, or WD Ultrastar DC). Mixed 512e/4Kn in the same VD will fail foreign import.