6.5 crashing / 6.0.3 stable

we’re running a cluster of 6.03 and 6.5 on Solaris 8. alteon load balancing. everything works fine… except…

the 6.5 box keeps crashing every day or so… fault recovery seems like it might actually be working (just turned it on), which is nice. but the crashing is not.

6.03 never crashes.

i am at a loss how to troubleshoot this… i’ve tried so many things… compact -c -i, tuning exactly like non-crashing machinem etc. there seems to be no logic to when or why and i dont know how to read NSDs :frowning:

anyone interested in helping? we are willing to pay ($$) :slight_smile:

suggestions are welcome too… should i redo all the FT indexes? i’ve seen that mentioned… but seems silly to me.

thanks guys,

Damian

damian@svothi.com

800-661-4269 x88

Subject: Post the stack trace which has the text “FATAL THREAD” from the NSD file.

Perhaps someone will recognize it and can give you some help

Subject: RE: Post the stack trace which has the text “FATAL THREAD” from the NSD file.

I will have one soon. It’s crashing more than once per day! I have been pulling my hair out trying to find the NSDs… until I JUST NOW found that I had set the execution time to 30 instead of 300… so they were never finishing when fault recovery restarted the box…

soon…

Subject: RE: Post the stack trace which has the text “FATAL THREAD” from the NSD file.

Alright I am getting annoyed. The server keeps crashing and there are no nsd logs being generated. I know it’s PANICing and running nsd… i can see it in the console and nsd running in the processes list.

However, I cannot find any logs. There are not in the specified NSD folder nor in the IBM_TECHNICAL_SUPPORT folder.

Fault recovery and NSD are enabled in server doc. 300 secs allowed for nsd to run… shouldnt that be enough?

UGH! Please help. I can’t begin to diagnose without getting an NSD… this is so annoying.

D

Subject: RE: Post the stack trace which has the text “FATAL THREAD” from the NSD file.

The server keeps crashing every few hours…

Set NSD execution time to 600 secs… still no NSD file…

Server is not overutilized. CPU < 50%. This is a SOLARIS box. Disk iowait not abnormal.

Maybe this will help - you can see that NSD runs… but no LOG FILE!!! Is it because it’s getting terminated by the fault recovery??? Maybe I should disable that altogether??

Stack base = 0xf741230c, Stack size = 7424 bytes

Fatal Error signal = 0x0000000b PID/TID = 10652/61

Thu Apr 22 20:48:16 Fault recovery is in progress

Thu Apr 22 20:48:16 Running NSD

NSD is in progress …

Thu Apr 22 20:53:37 Terminating tasks

INFO: Terminating NSD subprocess 8006

INFO: Terminating NSD subprocess 8052

Run /opt/dominor6/lotus/notes/latest/sunspa/nsd.sh -help for more info on new options/features

Script started at:

Script ended at: Thu Apr 22 20:53:38 EDT 2004

Generated Info/Warnings/Errors:

(1) INFO: Terminating NSD subprocess 8006

(2) INFO: Terminating NSD subprocess 8052

Thu Apr 22 20:53:55 Freeing resources

Thu Apr 22 20:53:55 Fault recovery completed

Lotus Domino (r) Server, Release 6.5, September 26, 2003

Copyright (c) IBM Corporation 1987, 2003. All Rights Reserved.

Restart Analysis (265 MB): 10% 20% 30% 40% 50% 60% 70% 80% 90% 100%

Recovery Manager: Recovery being performed for DB /dominor6/lotus/data/mail/rchadwic.nsf

Recove

Subject: RE: Post the stack trace which has the text “FATAL THREAD” from the NSD file.

OK, so it just crashed.

I looked, and the NSD was there… so i started looking at it… using the more command.

Before finishing, I closed it and was going to search it for “PANIC”…

but it had disappeared.

DOMINO IS DELETING THE NSD FILE for some reason… probably when the box comes back up.

How stupud is this? So now, I still havent been able to see it… its like a game of cat and mouse.

Subject: RE: Post the stack trace which has the text “FATAL THREAD” from the NSD file.

Hey guys. Well it doesnt look like anyone cares :frowning:

Anyway, the NSDs still delete themselves when the server restarts. However, I “snagged” one my copying it beforehand… got 99% of it…

there is no mention of Panic or Fatal or anything in it. However, in the console, it said that and just specified the SERVER ID.

the log is huge but too much info to digest.

help…

D

Subject: RE: Post the stack trace which has the text “FATAL THREAD” from the NSD file.

This is NOT IBM Technical Support – which is where you should be directing your questions. There might be an off chance that someone here could help you if you could provide fatal thread info, but you can’t. That means that the best anyone can do is take a wild-a$$ guess.

Subject: RE: Post the stack trace which has the text “FATAL THREAD” from the NSD file.

Thanks Stan. I posted the FATAL THREAD in another thread… but apparently it didn’t provide enough info.

We are in the processing of installing all the latest OS patches on the box… wasn’t as easy as I had hoped to accomplish. If THAT doesnt work, we will contact IBM Support.

I always try to sort out an issue on notes.net before calling support. It’s usually faster and more prudent to do so. That’s what this forum is for… isnt it?

Thanks again

Damian