Forum Replies Created
-
A little follow-up. The managed cluster seems … fragile. It seems to break … a lot.
-
Ian,
I now have the entire cluster online. Woot!
Essentially, I went node by node by node, beginning with the Cluster Controller and did the following:
1) Compressor: Reset Background Processing
2) Ran “Compressor Fix” from Digital Rebellion, to blow away caches etc. Note to the concerned, this did not blow away my compressor configurations.
3) Shut down C4 & Qa4. NOTE: I did not power cycle the boxes. I just closed the app.
4) I brought up C4/Qa4 on the nodes FIRST and the Cluster Controller last.
5) I rebuilt the cluster with all 4 nodes and the CC in the mix.
I offer a little more detail in the screencast linked below.
Thanks to Ian and Bob for their help.
Doug
https://screencast.com/t/217vGrGC
-
I added the box by IP both in Compressor a Qadmin Preferences. Both se te box as online, but it does not appear in the Qadminstrator queue to be added to the cluster.
Did I do it right?
-
We have success! I can submit to the cluster and it is tearing up the test file. Thanks for all of your help Ian.
Oddly, trying to submit to the local machine seems to hang. The local machine is part of the cluster and required to be part of managed services. So, maybe that is the issue? You would think Apple would put in an error message like “This machine is part of XWZ cluster and cannot process jobs independently”.
Incidentally, I went back and my storage mounts after seeing 0 machines mounted in sharing. I had just connected to machines over Bonjour, which should work. In any case, I specifically mounted the drive on all nodes via AFP://. I also created a specific cluster user and connected to the drive using that user, and saved the login creds to the keychain on each machine.
I am not sure that is a factor though, as I forgot to do it on one of the machines and the cluster is still using it. Weird huh?
So, one final quirk … one of my nodes (Green Leader) is not being seen by the cluster controller (Red Leader). GL sees RL and the rest of the nodes, but RL does not see it. So, it is not in the cluster. Again, weird.
Any ideas?
-
Resetting background processes let me build the cluster without error. Test render in process, but still appears to be hanging.
This process would be much easier if Apple incorporated more meaningful/useful error messages. 😀
-
Checked all machines … Firewall not on.
-
Ian – I really appreciate the help. So, don’t sweat it. 😀
1) When setting up the AFP share (in System Preferences >> Sharing), I gave Read/Write permissions to everyone. So, permissions should not be an issue, unless there is another way of setting them (CHMOD?) that I should be using.
2) No firewall of which I am aware. Is there something in OSX I should check?
3) As for resetting Compressor/QMaster, I have stopped and started Compressor, cleaned Compressor/QMaster 3.5 off of the systems and power cycled all of the boxes in the cluster. Is there is a “reset” switch in Compressor I am missing?
Thanks again.
Doug
-
Ian – All of the machine clocks are synced in Date/Time to the Apple time server. Is this what you meant?
DD
-
Ian,
I’ve include relevant bits of my console dump below. The following screencast explains what I THINK they are, given that they are out of context.
https://screencast.com/t/ak4cvrS8SK
Thanks,
Doug
5/30/12 12:34:04.012 AM [0x0-0x291291].com.apple.Compressor 2012-05-30 00:34:04.006 compressord[26945:6a03] In ‘__CFPasteboardCopyData’, file /SourceCache/CF/CF-635.21/AppServices.subproj/CFPasteboard.c, line 2372, during unlock, spin lock 0x6821a38 has value 0x0, which is not locked. The memory has been smashed or the lock is being unlocked when not locked.
*****************************
5/30/12 2:02:18.553 PM Apple Qadministrator exception (NSException raised by ‘NSPortTimeoutException’, reason = ‘Distributed objects message send timed out (timeout: 360104568.550303 at time: 360104538.550667) 1’) occured in connectAndCaptureService
*****************************
5/30/12 2:51:19.442 PM Apple Qadministrator ERROR: connect failed while updating service attribute info..
*****************************
5/30/12 3:32:33.484 PM qmasterd CDNSharedStorageServer::publishNotification: CException: SwampDaemon::NFSPublisher: authorization has expired
5/30/12 3:32:33.485 PM [0x0-0x14014].com.apple.Compressor 2012-05-30 15:32:33.483 qmasterd[293:1403] CDNSharedStorageServer::publishNotification: CException: SwampDaemon::NFSPublisher: authorization has expired*****************************
5/30/12 3:40:32.000 PM kernel IOSurface: buffer allocation size is zero
*****************************
5/30/12 3:41:29.034 PM Apple Qadministrator How did this happen!? JC is not captured!! Something is reeeeeaaaallly wrong! :-((
5/30/12 3:41:29.035 PM Apple Qadministrator JCName = RED LEADER (CC)*****************************
5/30/12 3:42:30.591 PM DiskImages UI Agent Could not find image named ‘background’.
5/30/12 3:42:30.591 PM [0x0-0x3d03d].com.apple.DiskImageMounter 2012-05-30 15:42:30.589 DiskImages UI Agent[641:707] Could not find image named ‘background’.
5/30/12 3:42:35.398 PM DiskImages UI Agent *** -[NSMachPort handlePortMessage:]: dropping incoming DO message because the connection is invalid
5/30/12 3:42:35.399 PM [0x0-0x3d03d].com.apple.DiskImageMounter 2012-05-30 15:42:35.397 DiskImages UI Agent[641:707] *** -[NSMachPort handlePortMessage:]: dropping incoming DO message because the connection is invalid