Data Recovery Case File · Enterprise · RAID 6 / iSCSI / ReFS

Four layers deep: recovering two rack QNAP arrays that stalled "onlining" mid-migration

Most of these case files fit on a bench. This one arrived on a trolley: two rack-mount QNAP arrays — a head unit and two expansion shelves each, around fifty disks in total, RAID 6 throughout — which had been mid-way through a planned migration to new storage when the configuration work went wrong. Both arrays now powered up into a permanent "onlining" state and never became accessible, holding hostage the iSCSI LUNs — 74TB, 64TB and a smaller sibling — that a Windows server environment depended on.

Devices2 × QNAP rack storage arrays (each: head unit + 2 expansion shelves); ~50 disks of mixed sizes; RAID 6
Presentation layer3 iSCSI LUNs per array (74TB / 64TB / 0.5TB), consumed by Windows Server 2016 via Storage Spaces, formatted ReFS, hosting virtualised server workloads
Reported conditionFailure during migration configuration; arrays stall indefinitely in "start-up / onlining" and never present storage; disks individually healthy
Fault classArray/volume metadata corruption above intact RAID members
Equipment usedAtola TaskForce 2 (staged parallel imaging at scale) · ACE Lab Data Extractor (RAID 6 and multi-layer reconstruction)

The enquiry

“There was a planned migration from our old QNAP storage arrays to new mirrored storage, but something went wrong while our IT team were configuring the copy. The QNAP is accessed through Windows Storage Spaces on a Server 2016 machine, formatted ReFS. When the array starts up to present itself, it remains stuck in an onlining state and never becomes accessible. RAID 6, roughly fifty disks across a head and two expansion units per array; three LUNs per array of 74TB, 64TB and 0.5TB.”

The shape of the problem

Enterprise storage fails in layers, and this stack had five: physical disks; the RAID 6 sets striping data with double parity across them; the arrays' internal volume management; the iSCSI LUNs — which to the appliance are simply enormous files or block ranges within those volumes — presented to Windows as disks; and finally Storage Spaces and ReFS on top, where the organisation's actual data lived. The permanent "onlining" stall was a fault at the third layer: array metadata left inconsistent by the interrupted migration work, so the appliances could no longer assemble a description of their own storage — while the disks beneath, as the customer's own checks suggested, remained individually healthy. The appliances could not be trusted, or coaxed; the recovery had to bypass them entirely and rebuild what they'd forgotten, from evidence.

Recovery at scale

Fifty disks make method non-negotiable. Every drive was labelled to its exact chassis, shelf and bay before removal, then imaged on the TaskForce 2 in staged parallel batches — its many simultaneous ports are what make a fifty-disk intake a schedule rather than a season — with each image hash-verified, and the original disks retired to shelves for the duration. From there the stack was rebuilt upward in Data Extractor, one layer at a time, each proven before the next was attempted. The RAID 6 sets first: member order, stripe geometry and the double-parity rotation established from on-disk structure and mathematically verified — RAID 6's second parity stream, the feature bought for resilience, doubling here as a truth-check across every stripe. Then the arrays' internal volumes reconstructed from the assembled sets, and within them the LUNs located and extracted — the 74TB and 64TB giants lifted out as intact virtual disks. Only then the top of the cake: the extracted LUNs presented together to a reconstruction of the Storage Spaces layer, the ReFS volumes walked, and the server environment's file systems — the virtual machines, shares and databases the whole edifice existed to hold — opened, extracted and verified against their own metadata.

Outcome

The organisation's data recovered in full from both arrays and delivered to its new storage — the migration completed, in the end, just not by the route anyone had planned. An engagement like this is measured in weeks, not days, and priced as the project it is; we say so plainly because enterprise clients deserve schedules they can plan incident response around. The deeper lesson belongs in every IT department's migration runbook: the most dangerous hours in a storage system's life are the ones spent changing it, and the moment a migration is scheduled is the moment the source arrays' backup coverage should be at its strictest — not, as convention tempts, winding down because "it's all moving anyway."

For IT teams mid-incident on shared storage

When an array stalls onlining: stop restarting it — repeated initialisation attempts against inconsistent metadata can convert a stall into damage. Don't accept firmware "repairs" or re-initialisation offers, don't reseat disks without labelling, and preserve the failed state exactly. RAID members that stay untouched keep every recovery option open; an appliance allowed to keep trying keeps none of its promises.

Business storage down and the vendor out of answers?
Call Bristol Data Recovery on 0117 332 1137 — enterprise RAID, iSCSI and virtualised environments recovered from evidence, not luck.
Request a quote online →

Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.

Call us — 0117 332 1137
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →