Data Recovery Case File · Enterprise RAID · Beyond Tolerance

Two failures past one drive of tolerance: a twelve-disk RAID 5 reassembled from its own metadata

The enquiry read like a well-kept asset register: a 36-bay enterprise chassis, a Linux system, a hardware RAID controller, and — the patient — its twelve-disk backplane array of 4TB drives, configured long ago as RAID 5, now reporting two failed members and dropping the server into emergency boot. In the same chassis, pointedly, sat a twenty-four-disk RAID 6 reporting perfectly healthy. The customer's team had already removed drives for inspection, and their controller had opinions about that too. Twelve disks, one drive of tolerance, two drives gone: arithmetic this archive exists for.

SystemEnterprise 36-bay chassis, Linux; hardware RAID controller; affected array: 12 × 4TB SATA, RAID 5; adjacent 24-disk RAID 6 healthy
Reported conditionController reports two failed members; system in emergency boot; drives removed for inspection post-failure; array offline
Fault classDouble member failure on single-parity RAID — recovery dependent on drop order and member imaging
Equipment usedAtola TaskForce 2 (parallel imaging of all twelve members) · ACE Lab Data Extractor (controller metadata analysis, virtual reconstruction)

The two questions a double failure turns on

RAID 5 survives one absence; at two, the array is over its tolerance and the controller rightly refuses to guess. Recovery past that point rests on two questions asked of the evidence. First: how dead is each "failed" drive? Controllers drop members for felonies and misdemeanours alike — a genuinely dying disk and a momentarily unresponsive one get the same verdict — so all twelve members were imaged on the TaskForce 2, the ten survivors cleanly and the two outcasts with error-tolerant patience; both yielded substantially complete images, one near-perfect, one ragged at the edges. Second: who dropped first? The moment a first member leaves, it freezes in the past while the array writes on without it — so it can never rejoin a reconstruction as an equal. The controller's own metadata, written onto every member, answered from the images: configuration, sequence and event history read and cross-checked, the two departures dated and ordered beyond doubt. (The team's post-failure drive removals, done without bay labels, cost nothing because that metadata identifies every member's position — but it's a grace not to lean on, and the intake paperwork said so.)

The reconstruction

In Data Extractor, the array was rebuilt virtually the only defensible way: the eleven members current at the final failure — ten survivors plus the second, freshest dropout's image — assembled in the controller's recorded geometry, with the ragged patches of that eleventh member recomputed from the others' parity, single-parity mathematics stretched exactly as far as it lawfully goes; the first, stale dropout excluded from the data path and consulted only as a last-resort witness for regions with no other testimony. The Linux volume layered above came up coherent from the virtual whole, filesystems mounting current to the array's final write; the estate was extracted, integrity-checked, and delivered onto new storage sized for its scale, with a written reconstruction report for the team's post-mortem.

Outcome

Effectively complete recovery, boundaries itemised — and an argument the customer's own chassis had been making silently for years: the twenty-four-disk array entrusted with double parity sat healthy through this entire story, while the twelve-disk single-parity array wrote this page. RAID 5 across a dozen ageing spindles is one silent dropout from living on luck; the rebuild-time exposure alone should retire the configuration at this scale. Their new array carries the second parity stripe its neighbour always had — and a monitoring alert on member drops, because the first departure in this story went unnoticed until the second made it unmissable.

For admins holding a multi-failed array

Stop: no force-online, no rebuild attempts, no re-seating drives to "see if they come back" — every write reshapes evidence the reconstruction needs. Label positions if the drives are still seated; if they're already out, don't guess an order. Preserve the controller's state and logs, image before analysis, and treat the drop sequence as the crown jewels of the case — with stale members in play, when matters more than whether.

Array past its tolerance?
It's usually still reconstructable — call Bristol Data Recovery on 0117 332 1137 before anything rebuilds or rejoins.
Request a quote online →

Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.

Call us — 0117 332 1137
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →