Forum Replies Created
-
[David Patterson] “confused between RAID 5 and RAID 3”
It is confusing, and some people still use it, yet RAID3 basically a dinosaur from a decade ago. Among parity RAIDs, it’s RAID5 for under six drives, RAID6 for six or more. RAID1 or 10 can also be used in certain scenarios, rarely in video though due to low space efficiency.
You can get a decent ATTO or Areca RAID controller ($500-900) but it’s an overkill for four-five internal drives. The box I linked above is a good balance between cost, capacity and performance.
If you need performance first, just use OS striping to make a RAID0 array on your internal drives. You’ll get around 500MB/s (on a four-drive RAID0) – something not easily achievable with external boxes unless you spend at least $1.5K more. Do backups of course.
-
[Alex Gerulaitis] “This is why I think formulas will help play with failure and URE rates, and perhaps also match them to available stats.”
Found a spreadsheet that seems to be doing exactly that, on zetta.net/_wp/?m=200906.
It puts RAID10 reliability at about 25% that of RAID6, with certain assumptions about drive counts, URE and failure rates. Not sure how trustworthy it is until someone more knowledgeable than me analyzes it.
(The page I linked to above has errors and may not be easily navigable depending on your browser. I put that spreadsheet and up on my GDrive if you’d like to take a look at it.)
-
External, correct?
What interface do you need it to have? USB 3.0, eSATA, FW800 are the usual suspects.
Speed requirements?
Probability of failure / data loss: RAID5 sets do fail, especially with high capacity desktop drive due to the infamous “URE on rebuild” problem but they’re in general more reliable than RAID0 sets. What I am trying to say is that it’s reasonable to expect nearly any RAID5 set to lose at least some data over its lifetime, when using high capacity desktop drives. So use it with caution and back up your data.
That said, ProAvio EB400CR has a good reputation and support.
Alex Gerulaitis
Systems Engineer
DV411 – Los Angeles, CA -
[Alex Gerulaitis] “Wonder how Quadro K5000 and the upcoming GTX-780 compare…”
May have answered my own question: GTX-480 and GTX-580 are the only GeForce cards with a 384-bit memory interface. Memory bandwidth (192GB/s) is the same as in GTX-680. Core count – much lower (512 on GTX-580 vs 1536 on GTX-680). So somehow that memory interface seems to be playing a critical role.
(Edit: scratch that: GTX Titan and GTX-780 also have a 384-bit memory interface. Not sure why Titan is slower than 580 in ray tracing.)
-
[John Cuevas] ” the 580(using Fermi architecture) seems to outperform the gtx 680″
John,
Thanks so much for compiling the benchmarks – eye opening. 🙂
Wonder how Quadro K5000 and the upcoming GTX-780 compare…
Also, are there CPU-based AE benchmarks pitching 3770K and 3930K CPUs against dual E5 Xeons?
Thanks again.
-
Thanks for the input Vadim.
[Vadim Carter] “I do have all the nitty-gritty details and formulas laying around somewhere which explain mathematically which RAID level is more reliable. If my memory is correct, RAID 10 is deemed slightly more reliable than RAID 6.”
Would be awesome if you could find those formulas and see if they account for all the reliability variables and factors:
– chances of enough simultaneous drives failures within a rebuild cycle to bring down the whole array
– UREs and how they affect reliability, especially during rebuilds
– protection against other causes of data lossFrom my research so far, reliability advantages of either RAID level depend on a balance of drives’ FR and URE rates. E.g. if there is a guaranteed URE during a mirror rebuild, then there is a guaranteed data loss incident in RAID10. 6 OTOH will be resistant to it, as it still has another place where to look for healthy data during rebuilds.
Based on that alone, resistance of RAID6 to data loss due to UREs is much, much higher vs. RAID10 on a single drive failure – thousands to millions times so. Does it make RAID6 much more reliable overall? I don’t know, and would love help with that.
At the same time, resistance to whole drive failure is really down to FR within drive replacement / rebuild window, not necessarily drive’s lifetime – which may make Wikipedia and many other formulas not too trustworthy.
If we assume that there’s an increased chance of a drive failure during rebuild (in the same mirror pair, or in RAID6 array), then RAID10 may have less of an advantage, depending on those chances.
This is why I think formulas will help play with failure and URE rates, and perhaps also match them to available stats.
[Vadim Carter] “To anyone reading this, make no mistake about it, RAID does not protect against data corruption.”
Perhaps you’re talking about “silent rot”, not all incidents of data corruption? A URE is data corruption, and redundant RAIDs do protect against them, to a degree?
[Vadim Carter] “This is why RAIDZ is an excellent choice where data integrity is paramount, each block of data is checksummed and the checksum is then written to a separate area of the disk with a pointer to the original data block.”
Are checksums a function of RAIDZ or ZFS?
[Vadim Carter] “One last thing in regard to performance, all other things being equal, RAID 10 will always outperform RAID 6. The reason is simple – parity calculation is an “expensive” operation. Calculating it twice is even more “expensive”.”
Agreed, parity calculations are expensive – yet computing power grows much, much faster than disk speeds. Perhaps at some point parity calculations on beefy systems with software RAID will get cheap enough to have a negligible effect on performance if they haven’t already? I’d be more concerned with I/O than parity overhead.
-
[Jostein Svalheim] “We have experienced server admins here, so a simple mirroring server should be a cakewalk, right?”
Haha… 🙂
The absolute first thing to do is to ask your existing admins for suggestions. They’re the ones to maintain it, right? If so, they should be consulted with.
The 2nd absolute thing to do is to get the specs on the existing array. The mirror you’re putting should be the same or faster speed than the 1st array, or else you will be slowing the 1st array down.
The 3rd thing – get an idea of an existing workload – average and peak I/Os and block sizes. Your server admins might help with that. Whatever you are adding to your infrastructure gotta accommodate that workload in the short and long terms.
Last, mirroring adds to the I/O and bandwidth overhead (controller, bus, OS levels). Usually not an issue, but gotta watch out for that, too.
After we get all that done, we can start thinking about specific chassis, drives and controllers.
Alex Gerulaitis
Systems Engineer
DV411 – Los Angeles, CA -
[David Gagne] “If you have RAID5, and a drive fails, and you replace it same day, you will not lose data. “
You might if you get a drive failure or an URE during rebuild – which is technically after replacement. Chances of an URE (non-catastrophic data loss) during rebuild are higher than 50% on a 10TB array with desktop drives – if we believe published URE numbers – which some people think are too optimistic.
Single bit URE on video files – may not be a big deal. In Pentagon – it probably is – they wouldn’t want that drone send a Hellfire to the wrong address…
So RAID5 is fine in many circumstances but it can no longer be considered truly resilient with today’s high capacity drives. On top of it, we don’t really know what happens on that near-certain URE-on-rebuild: will the RAID brain offline the drive and fail the whole array? That would mean a near-certain catastrophic data loss on a whole array with a single drive failure. Or will it just leave that URE in place meaning a single bit error in your file or file system? Either way – not pretty.
RAID6 is quite a bit more resilient, both with UREs and drive failures as chances of two or three happening within the same replacement / rebuild cycles are much, much smaller. RAID10 is also resilient to multiple drive failures and to a lesser degree – with UREs.
I just wanted to see if anyone had a math model or stats pitching 6 against 10 but it may indeed be an exercise in vain.
-
[Walter Soyka] “More RAM = more better.”
You gotta patent the hell out of that phrase, Walter.
-
[David Gagne] “A multi mirror raidz would be most resiliant to drive failure. Add in HA, UPS, and it’s never going down…”
Thanks David. Not familiar with Z although I heard good things about it. Are there any models that quantify that resiliency, and measure it against, say, 10, 6, and a 3-way RAID1, in a similar way Adam Leventhal does it with 5 and 6?
(I realize copying the same data in at least three places is the most resilient method – and RAID6 might be the most efficient implementation of it in terms of space utilization – even if not the most reliable.)
Wikipedia has an AFR formula supposedly measuring resiliency but I am not too happy with it: doesn’t account for risks of UREs during rebuilds; factors annual FR rather than the risk of simultaneous failures within drive replacement and rebuild time frames.