Felix Buechner
Forum Replies Created
-
I am having a lot of problems very similar to not all but many of the problems described in this discussion (like the Finder stalling completely after some time; depending on what size of files i read with what speed to which target; for example i was having no problem writing to the RAID5, and no problem reading off it – but when i read or wrote large amounts of data via LAN, i had immediate kernel panics. Very, very weird stuff).
I was suspecting the raid controller (which was not Highpoint) all the time, and finally, i replaced it against a Highpoint 3522.
After that, i had very similar problems. Not exactly the same, but nearly so.Let me tell you what i think and what seems to solve the problem (right now, i am conducting very extensive testing; in a couple of days, i wll be able to tell for sure if i solved the problem):
Some time ago, i built servers with external hot-swap enclosures, directly connected to SATA raid controllers having internal connectors. Without the infiniband cables, i had to use SATA cables.
In doing so, i saved tons of money while having perfecty reliable, fast servers. Only the mechanics was looking rather strange (SATA cables running quite “raw” out of the Mac into some box…)
Of course to have long cables would help, so i did try SATA cables with a lengh of 2 meters.
I knew the SATA spec said “Maximum 1 meter”, but heck! There WERE cables with 2 m, so lets try!And then, i had very similar problems as the ones described above.
Ruefully, i returned to 1 m cables and everything was ok again.Today, building with the R8ML, i used 2 m cables – just for ease of servicing the system. It allowed me to take the MacPro out of the rack while leaving it fully cabled to everything.
Only there were the above problems. Then, as i said above, i took another controller completely, just to find my problems stayed much the same.
Now this in all probability means that i was following a dead end altogether.
Since the Mac, its RAM, its PCI bus, in short everything else had been tested over and over again, i began suspecting the R8ML; or a cable inside it or the cables leading to it.
I dont know of you guys, but i should add that one very peculiar property of the problem was that nearly never i got any meaningful log entry. I mean no entry at no log, neither any log of OSX nor of any of the controllers.
This means that something happens that is out of the norm AND the system has no way to respond to it or it freezes so completely that logging is rendered impossible.
This is quite compatible with a cable that is installed, but fails. This kind of problem usually is not logged, because it is so rare a condition. BTW back in those days, i had no entrys either.Anyway, i opened the box and found:
– the two infiniband plugs soldered to a small PCB carrying 4 SATA plugs each
– 8 SATA cables leading from said PCB to the backplane PCBThe SATA cable length i did not measure, but they are not routed tightly, so they are above 50 cm, much more likely around 70 cm.
All in all, a very simple, and clean, thing.
Only: whithout any active element between infiniband and SATA cables, the lengths add up!
I know the Inifinis are better shielded than SATA. They are allowed to be much longer.But are they still allowed to be longer when they carry SATA SIGNALS, connecting to SATA drives?
If you add 0.7 meter SATA cable, then a connector, then a short lead on a PCB, then another connector, and then a 2 m cable (even a nicely shielded one), i do not expect the signal to get better underway, but the opposite.All these plugs and cables cannot but degrade the signal, each one a tiny bit.
And that in the Infiniband world the cables are allowed to be long – fine with me!
SCSI nowadays is allowed some very great leghts, too.But again: We are sending SATA signals over those cables, not SCSI oder Inifini.
And so, the SATA cable constraints do apply.Anyway, this is what i did:
– i split my 8 drives into two groups of 4 and tested them separately and consecutively to be able to identify one ore more components (since no logs are available, i had to test to exhaustion or failure).One group of four would fail after max 8 hours, the other on did not fail at all.
I swapped cables but everything remained same, so the cables in themselves were both ok or at least identical.The failing group, i split into 4 single drives and tested them separately.
Two came out ok, two not.
I made sure all these results were repeatable.Then i exchanged the cables against shorter ones, having 1 m only.
Now, the two prevoiously faulty drives were perfect (i have to say that my tests run until either a drive fails OR it runs for 24 hours under nearly full load. In the latter case, i consider the drive ok, so of course faults that take more than 24 hours to show up, i miss).
I am pretty confident that i have a situation that has more that one contributor:
– backplanes
– drives
– cables
– controller
My hypothesis is that these components together are operating at the brink of stability.
Push a little too far and your system breaks.So, to relax the strain, i use shorter cables.
Right now, i am re-building the RAID 5 (it will be ready in one hour), and then, i will rigorously test it – for more that 24 hours (obviously, i hope it does not fail in five minutes …).
After i will report back with results, be they good or bad.