Case record · NAS & RAID · SDR-2025-0642
Two Dead Drives in One RAID 5 Array.
A business storage shelf with a familiar injury: two drives failed in the RAID 5
. The IT contractor had a couple of spares
to hand and had already slid one in to see whether the array would just rebuild itself
. It stalled a few per cent along. The trap was half sprung already.
Seeing the same thing yourself?
0800 6890668
The translation.
RAID 5 can carry one missing member and no more. With two gone the controller has nothing left to calculate from, and a rebuild begun in that state lays down new parity over the very stripes that reconstruction has to read in the state the failure left them. Business shelves add a hazard of their own: vendor sector sizes and proprietary metadata layouts that ordinary desktop tools read as gibberish. The spares had a proper job waiting at the end of this case. It was not the one they had been handed at the start.
Kit used on this job.
What happens in a case →| Platform | Its role in this case | Why we use it |
|---|---|---|
| PC-3000 SAS/SCSI | Spoke to the SAS members directly, which no ordinary desktop controller can manage | For SAS and SCSI disks out of business shelves, which normal kit cannot reach |
| Atola TaskForce 2 | Imaged the whole set side by side, turning a week of queueing into days | Images several drives in parallel — it turns a week on an array into a few days |
| UFS Explorer RAID Recovery | Worked out the layout and built the volume back up from the copies | Understands how NAS volume managers really work, instead of a flat array |
In the lab.
Image the lot, healthy members included
Every member went onto an imager, the two dead ones beside the four that still worked. SAS disks need hardware that talks their interface properly, and the vendor sector formatting used by the shelf had to be handled right at that stage, or everything built on top of it would have been nonsense.
Bring the failed pair back far enough to read
Both dead members were assessed and coaxed into a readable state, and the usual thing proved true — neither was uniformly dead, and most of each surface still read. Holding images of both let the reconstruction pick the better copy of any block instead of leaning on parity for everything.
Derive the layout, assemble it in software
Stripe size, disk order, parity rotation and delay were all read out of the on-disk metadata rather than guessed from a default. The volume was then put together virtually across the images, and the filesystem read out of that reconstruction.
The result.
The array came back and was verified, then written onto the company's own spare disks as plain storage — the only sensible part they were ever going to play. The stripes the aborted rebuild had overwritten cost part of one project folder.
Similar cases on the index.
Also from NAS & RAID.
Does this sound like your own drive?
The rule holds as in every case above: switch it off, and let a free diagnosis come before any decision.