Show stopper for R5 to R6 upgrades - Intermittent connection

Subject: Show stopper for R5 to R6 upgrades - Intermittent connection

John

We’ve also just hit this problem after upgrading our servers to R6.5.2 over the weekend.

Do your R5 clients also have the same problem? We’ve had several reports of the same behaviour from R5 clients as well as R6.5.2 clients. There have been no changes to the R5 clients at all in the period after the servers were upgraded.

I’m currently trying to to work out how long the period of inactivity is when the issue hits the user. One user who is regularly plagued by the problem yesterday left his machine for over 2 hours and didn’t have the problem when he went back to it. Later on it was left for a much shorter length of time (he can’t remember how long) and he got the problem again.

I’m currently experimenting with the DISABLE_TRIMWS=1 notes.ini setting (technote 1164456) after someone here suggested it in response to my posting a couple of days ago. I can confirm that this makes no difference on R5 clients, I haven’t come to any firm conclusions on R6.5.2 clients as of yet.

The only thing I’ve noticed in common with the problem machines is that the Notes client is not locked during the period of inactivity. I’ve currently got a spare machine setup on my desk experimenting whether locking the client makes any difference.

Like you I think I’ve probably got the problem where some users aren’t reporting the problem therefore working out any patterns to it is very difficult.

I’ve currently got a support call open with Lotus about this issue. They have suggested that it may be the power settings on the machine that is causing the problem. I’ve checked a few of the problem machines and have found that their power settings are set to Never. They do have a screen saver that kicks in after a certain amount of time, but so do I and I never have the problem.

I’ll keep you posted on anything else I find out.

Julie

Subject: Show stopper for R5 to R6 upgrades - Intermittent connection

[0] Are you 100% sure that if the server and the client are in the same subnet, it works fine, but otherwise no?

If so, here are some ideas:

[1] On the Clients, what is the default network port set to? Are more than one ports enabled?

[2] Is there a proxy setting in the Location document of the Client?

[3] Have you placed connection documents into the client PC with hard coded IP addresses to the server to see if the issue resolves?

[4] Are the server names in the Location document the NOTES Server names or are they hostname style? (i.e. Server1/Test vs server1.testworld.com)

[5] Have you tried using the following in the client to determine WHICH part of the Notes system is choking? (put in INI file)

Client_Clock=1

Debug_Outfile=SOME-FILENAME-TO-STORE-OUTPUT

[6] Simple Trace Route to the server (by name)- how long does it take between each hop?

[7] How about just copying a large Windows file to the server while on the same subnet and while not?

[8] Have you used network performance tools to see what results you get - comparing when they are on and off the same subnet?

Example:

http://dast.nlanr.net/Projects/Iperf/

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

In our environment we get the problems on machines that are on the same subnet as the servers and on machines that are on other subnets.

Some answers to your questions from my environment:

[1] TCPIP is the default and only port enabled in Notes.

[2] No proxy setting.

[3] The connection documents have always had IP addresses in them.

[4] The server names are Notes names.

[5] I haven’t tried this, but looking at other postings those who have tried it say that the report does not help.

Julie

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Your problem is then likely to be different than the original poster’s problem as yours does not go away when the networking and client versions change. Therefore, it has something to do with your new server.

You still ought to try the client clock - you will then be able to determine if the notes client is waiting on each network access, or is waiting for a view to update, or database to open, etc. This log will clearly point that out and it is important.

Here are some other things to try:

[1] Have you tried re-signing all the design elements in the discussion databases in question with an ID that is accepted by the users’ ECL?

[2] Where have you assigned View Indexing to occur on the server? Do you have Anti-Virus software running on the server?

[3] Very bizzare long shot on 6.5 clients: In the user preferences, turn off “Automatically Refresh Inbox” and see if that makes a difference.

[4] Is your Domino server on a SAN or SVA?

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Hi Julie, Hi Victor,[0] So far without exception, it always works on the same subnet. The intermittent failures occur only on a different local subnet and also users that connect via the Internet.

[1] TCP, only TCP enabled.

[2] No proxy

[3] All are IP and have been since R6

[4] Server1/test

[5] Still evaluating this, so far just get the same message logged as what appears in the pop-up.

[6] Tracert fails due FW

[7] No problems here, no hiccups even when Notes is stalled, eg. file copy continues, VNC session can be established and used normally while Notes is stalled.

[8] Subjectively there is no difference and the physical network does not stall. The failure is however very pronounced. Had a quick look at your reference seemed like an SDK, could not find any product downloads? Will get back to it however.

John

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

[8] I believe that there is a download link on the page. You download this piece of software and run one on the “server” and the other on the client. It will then test throughput. You can play with the size of the data and the amount of data.

After reading all of your additional posts, is it possible that the router is doing something naughty? You mentioned that you replaced the switches and NICs, but what about the router?

Also, have you checked full duplex vs half-duplex vs automatic on the cards and router? If you can set a specific port on the router to full duplex, for example, and then do the same for both the Domino server’s card and the client card, does it make a difference? (We have seen this scenerio, where the duplex setting is “automatic” on switches and the server NIC’s, but they do not seem to negociate properly - leading to reduced performance from the client’s perspective.)

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Hi Victor, and thanks for your interest.Have you read the plethora of similar posts that Julie highlighted which the same as mine:

http://www-10.lotus.com/ldd/nd6forum.nsf/55c38d716d632d9b8525689b005ba1c0/c143a0496a29428385256d5a00539c5f?OpenDocument

I have also experienced the issues you describe with the interface settings. The common problem is when each end is set to Auto negotiate. I have found that even some of the expensive switches seem to get confused here, with constant never ending traffic trying to setup each end. But that is not the issue here.

I will check out point [8] but I would be surprised if there is an issue as the network continues to work (eg. file copy, VNC etc.)the same even between the suspect client and server, while the Notes client is hourglassed.

You are correct about the router, it has to be suspect. I have upgraded the uCode to the latest level, factory reset and then carefully worked through the setup again. I have also contacted the manufacturer and discussed the problem a couple of times. I doubt if the router is implicated as the network still functions for other applications, while apparently hung for Notes.

The later router uCodes do not allow the MTU of the WAN interface to be setup. You can assign a value but an RFC change now says it must auto-set this value, so any manual entry gets overwritten. Old versions of uCode allowed it to be set and left the setting alone.

I think that means that I cannot rule out the possibility of it being an MTU/Defrag issue. Even replacing the router does not mean it is not an MTU issue unforunately.

From my workstation I have both R5 and R6 installed. I never have a problem with R5. I connect with a number of R6 servers all on the same network and each of them has the same problem. BUT the R5 servers are not on my Internet LAN, they are remote via the Internet.

What are your thoughts about Netbios over TCP being missing as described in a reply by me to this post? Why does it never fail when each end is on the same network? Is Netbios required?

Netbios cannot route, but Netbios over TCP can. Why does “Netbios over TCP” come up automatically on a server installation in addition to TCP that you select?

IBM support & Julie may be making headway with the SERVER_SESSION_TIMEOUT setting??

Regards, John

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Yes, I have read the earlier posts.

To your questions about NetBIOS and other items, perhaps you might find it valuable to know about our setup, and to know that we do not have the issues you describe.

We operate different subnets, run ND6.5 servers on United Linux, AIX, Windows 2000, and Windows 2003. We still have one 5.08 Server on Windows 2000. Our clients are a mix of 6.5 and 5.08. Most of the clients have NetBios over TCP/IP turned on. Some, like mine, do not (I personally feel that it is faster without it). The *nix server systems do not. Additionally, we have deliberately named some of our machine names differently than the Domino server name (i.e. Domino01/Test vs DTTST01 as the machine name) which would cause some problems with name resolution over plain old Windows NetBIOS.

We do not have the problem you describe.

Therefore, I will openly admit that I strongly lean against some default setting in the Domino servers or clients. We did nothing special in this regard when we set up our servers and clients.

If it is a problem with a server setting - like a “timeout” issue such as IBM is suggesting for Julie’s problem, then why don’t we experience it? Or, more importantly, why don’t your R5 clients experience it? — Why don’t both the R6 and R5 clients experience it when in the same subnet?

So, to me, the problem is not likely to be an incorrect server setting - although that might MASK the problem and make it “go away” it does not SOLVE the problem.

I think that there are now only two possible paths left:

  1. The client (or server) cannot “resolve” or “find” something it is looking for in your current network setup.

-OR-

  1. The client is sending unusual data to the server (or visa-versa) and some piece of hardware or cabling is choking it. (Which I have had happen before - with Windows appearing “fine”. It was a bad patch cable from the server to the switch.)

Here are some ideas to Check if it is #1:

[a] that the server can see the clients and that it isn’t doing something like reverse DNS (like newer versions of Linux do for telnet sessions)

[b] that the clients can always see the server

[c] it would be nice to know what the client is doing that takes so long (thus the suggestion for putting the debug line into your client’s INI). You could also try DEBUG_NAMELOOKUP=1 to see if it is choking on name resolution.

[d] things like DHCP and DNS are always available and DHCP lease refreshes are reasonable - connections aren’t being dropped somewhere (see my last paragraph below).

[d1] toward your NetBios idea - what if you put the client in the same subnet as the DNS server and leave the Domino server in a different subnet? Does it work then?

[e] there are no firewalls or packet filter devices that might be dropping packets. Perhaps the R6 clients are more sensitive than previous versions.

If it is #2 - well, that would be a pain to have to check every point of connection in the subnet configured network. And, like 1C above, it may be helpful to find out what Notes is doing (opening a database, opening a view, etc) that causes it so much trouble.

Lastly,

We do have one user who has a laptop that has to change between his office network and home network manually. For some reason, Windows gets confused, and DROPS the office connection for a few seconds (you can see the LAN connection with the X in it in the taskbar while this happens.). Then the Notes client experiences the same thing that you describe for all of your users. His solution is to shutdown Notes, and go back in, and all is well (until the next time).

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Hi Victor,I have been running with the client debug enabled for a week or so and trying to understand the log file.

I displayed the client’s performamce analyser while the “hang” was occuring and there is barely any load at all(Peak CPU <2%) and no change in memory requirements.

If a local replica exists then the client will not display an error message, instead it opens the local replica.

Following is two snips of log files, the first with a local replica and the second trying to open the same DB but after the local replica was removed.

Pops-up an error message “Network operation did not complete in a reasonable amount of time; please retry”

Note: Line [82]61 seconds

16 ms. [136+32=168] (Session Closed)

(51-7763 [63]) UPDATE_COLLECTION(REP4A256E22:0050840D-NT000011CE): 15 ms. [28+16=44]

(52-7763 [64]) GET_MODIFIED_NOTES(REP4A256E22:0050840D): 0 ms. [30+28=58] (No documents have been modified since specified time.)

(53-7763 [65]) SET_COLLATION: 0 ms. [18+16=34]

(54-7763 [66]) SET_COLLATION: 16 ms. [18+16=34]

(55-7763 [67]) READ_ENTRIES(REP4A256E22:0050840D-NT000011CE): 0 ms. [136+288=424]

(56-7763 [68]) GET_ALLFOLDERCHANGES_RQST: 15 ms. [66+50=116]

(57-7766 [69]) OPEN_NOTE(REP4A256E22:0050840D-NT00004BEA,06400136): 15 ms. [54+3492=3546]

(58-7788 [70]) OPEN_NOTE(REP4A256E22:0050840D-NT00004BEA,06400136): 16 ms. [54+3478=3532]

(59-7812 [71]) CLOSE_COLLECTION(REP4A256E22:0050840D-NT000011CE): 16 ms. [16+0=16]

(60-7813 [72]) SET_UNREAD_NOTE_TABLE: 94 ms. [104+42=146]

(61-7813 [73]) CLOSE_DB(REP4A256E22:0050840D): 0 ms. [18+0=18]

(62-8221 [74]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 16 ms) (OPEN_SESSION: 16 ms)

0 ms. [54+38=92]

(63-10369 [75]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 2953 ms) (OPEN_SESSION: 0 ms)

0 ms. [96+28=124]

(64-11274 [76]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 3016 ms) (OPEN_SESSION: 16 ms)

0 ms. [94+28=122]

(65-12179 [77]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 2938 ms) (OPEN_SESSION: 15 ms)

0 ms. [94+28=122]

(66-13084 [78]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 2969 ms) (OPEN_SESSION: 0 ms)

0 ms. [78+38=116]

(67-13989 [79]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 2953 ms) (OPEN_SESSION: 0 ms)

0 ms. [108+42=150]

(68-14894 [80]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 2954 ms) (OPEN_SESSION: 0 ms)

15 ms. [110+28=138]

(69-15797 [81]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 5969 ms) (OPEN_SESSION: 16 ms)

0 ms. [92+38=130]

(70-15816 [82]) OPEN_DB(CN=domino/O=ATECH!!mail\jaylmer.nsf): 61015 ms. [-2099±17918=-20017] (Network operation did not complete in a reasonable amount of time; please retry)

(71-15877 [83]) SERVER_AVAILABLE_LITE: (Connect to domino/ATECH: 15 ms) (OPEN_SESSION: 0 ms)

16 ms. [38+44=82]


With local replica:(no error message pop-up)

Note Row [63] 61 Seconds

(61-1610 [61]) CLOSE_DB(REP4A256E22:0050840D): 0 ms. [18+0=18]

(62-2507 [62]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 2953 ms) (OPEN_SESSION: 0 ms)

0 ms. [94+28=122]

(63-3145 [63]) OPEN_DB(CN=domino/O=ATECH!!mail\jaylmer.nsf): 61016 ms. [-3917±130706=-134623] (Network operation did not complete in a reasonable amount of time; please retry)

(64-3206 [64]) SERVER_AVAILABLE_LITE: (Connect to domino/ATECH: 15 ms) (OPEN_SESSION: 0 ms)

0 ms. [36+40=76]

(65-3247 [65]) OPEN_DB(CN=domino/O=ATECH!!mail\jaylmer.nsf): (Connect to domino/ATECH: 219 ms) (Exch names: 0 ms)(Authenticate: 0 ms.)

(OPEN_SESSION: 16 ms)

0 ms. [132+254=386]

(66-3247 [66]) GET_UNREAD_NOTE_TABLE: 0 ms. [76+122=198]

(67-3247 [67]) OPEN_NOTE(REP4A256E22:0050840D-NTFFFF0010,03000400): 0 ms. [56+754=810]

(68-3247 [68]) GET_NAMED_OBJECT_ID($profile_015calendarprofile_): 0 ms. [58+28=86]

(69-3247 [69]) OPEN_NOTE(REP4A256E22:0050840D-NT000008FA,00400020): 16 ms. [52+3580=3632]

(70-3247 [70]) GET_NAMED_OBJECT_ID($profile_024archive database profile_): 16 ms. [68+28=96]

(71-3247 [71]) OPEN_NOTE(REP4A256E22:0050840D-NT000008FE,00400020): 0 ms. [52+190=242]

(72-3248 [72]) OPEN_COLLECTION(REP4A256E22:0050840D-NT000011CE,0040,0000): 15 ms. [102+56=158]

(73-3248 [73]) GET_NOTE_INFO: 16 ms. [22+106=128]

(74-3248 [74]) SET_COLLATION: 15 ms. [18+16=34]

(75-3248 [75]) SET_COLLATION: 0 ms. [18+16=34]

(76-3248 [76]) READ_ENTRIES(REP4A256E22:0050840D-NT000011CE): 16 ms. [68+4576=4644]

(77-3248 [77]) GET_ALLFOLDERCHANGES_RQST: 0 ms. [66+52=118]

(78-3248 [78]) POLL_DEL_SEQNUM: (Connect to domino/ATECH: 16 ms) (OPEN_SESSION: 0 ms)

0 ms. [94+28=122]

(79-3249 [79]) CLOSE_COLLECTION(REP4A256E22:0050840D-NT000011CE): 0 ms. [16+0=16]

(80-3249 [80]) CLOSE_DB(REP4A256E22:0050840D): 0 ms. [18+0=18]

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Sorry I wasn’t here yesterday - I was on vacation.

Great log info. It looks like it trys to connect several times, then eventually DOES connect and then tries to open the database and then the connection is lost.

If you are still having the problem, here are some ideas.

You can try this at the server console at the same time you have a problem with the clients: SH TR to see the server’s perspective on similar stats.

Additionally, SH ST SEM may also be useful to see if anything is getting dropped due to overcrowding. (Doubtful, but still worth a look.)

Your log file indicates that it was connected fine and operating fine to a database (4A256E22:0050840D - is that your mail file, jaylmer.nsf?), closes the database, and then suddenly fails to connect again. Is that correct?

And, to confirm, this NEVER happens on an R5 client, right? And, if in the same subnet, it NEVER happens to R6 or R5?

Lastly, just curious, if you ping the server from your workstation when Notes is having these problems, with say 1000 packets, is there any packet loss?

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Hi Victor,Network:

In my immediate vicinity, I have an ADSL modem/router that IP forwards real-world IP addresses. Each Domino server is assigned one of these IP’s. Another IP is assigned to another router to create a workstation LAN.

I can ping either router; my immediate gateway or the ADSL router.

As you suggested, I pinged the ADSL router with the -t option (ping until stopped). Every five minutes or so there would be a packet loss but normally 100%. I left it running for a while and then tried opening a DB from the workstation lan. Notes hourglassed as usual for 60 seconds or so. The ping times stayed exactly the same and there was no packet loss over that period. The client performance monitor CPU and memory stayed flatlined. The Domino server performace monitor flatlined but the CPU kicked up to about 30% for a short time at the end of the error message period.

SH ST SEM Does not seem to be legal syntax?

Am I right that the SH TR keyin does not return real-time information, just a summary of accumlated counters? To be meaningful I would have to do this keyin just before and just after opening a DB?

Response: Normally I get the problem with my mail DB as I frequently try to open it. Normally the DB is closed, I doubleclick the icon and it hourglasses. Eventually 60/90 seconds it times out, but the DB never opens, at least visually on the client unless the operation is retryed at the client. If the client can find a local replica then it opens that instead of producing an error message. One problem I find with the “client error log” is that each row is not timestamped. This means you don’t know whether it is hours or milliseconds between rows.

Response: I have never had the problem with an R5 client talking to an R5 server. Only R6 to R6. I have never tried a crossover. If on the same subnet it never happens R6 to R6. I don’t have any R5 servers on this system. I can connect with R5 client to remote R5 servers via the Internet and through the same two routers mentioned above. I never have a problem with this system.

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Hi Victor,Tried pinging the Domino server machine from the Notes client machine. Same as for pinging the router, the Notes client is hourglasses but the ping period stays the same, no packets lost.

There is one other point I remember. When Domino was installed, it would not route mail as it could not resolve DNS. Every other product on the server worked fine. I could even Telnet out a mail message.

I had to add registry entries as Domino did not recognise the existing Win2000 server DNS entries.

Refer Technotes: 178915, 190865

I am not saying this is relevent, just mentioning it.

Any thoughts?

John

Time to get up on my soapbox:

There does not seem to be a hardware error. Eventually you get as far as your tools allow but that may not give the answer.

Domino/Notes has evolved as a DOS product which now has a Window’s wrapper. Various things interfer with its operation that do not affect any other product installed on a Window’s machine:

1/ You can only have a maximum of 250 fonts installed on a workstation.

2/ Intolerant of the MTU packet size being larger than the router pipe size.

3/ Domino does not sometimes be able to resolve DNS requiring a special registry fudge.

4/ The problem described here with the initial connection.

While I realise that not everyone is having this last issue, a significant number of organisations are experience it, both with R5 and R6. For each one of the common issues described above, not everyone experiences them with every organisation, but they are still Notes only issues.

My suspicion is that Notes and Domino are still very DOS based with their own drivers and although they have Windows wrappers, they do not actually hook into Windows in a functional sense still relying on their unique drivers and ini files etc.

I don’t really have a problem with that in concept as long as it works. But historically when it does not work we are told that it is not a Notes issue which would be true if we were still in a DOS world, but as every other installed product works fine in a Windows environment you have to wonder.

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Interesting results. And the technotes make for interesting reading. I’ve never been somewhere where the server and/or clients were upgraded from WinNT to Win2000 (always a clean install), so these technotes are good to know about.

I am not saying this is relevent, just mentioning it. Any thoughts?

Well, that actually sounds like a huge clue.

Communication has to be two way, and if as you state, Domino is in its own little DOS style world, the server may be attempting to contact the client and cannot find the client to respond to - so the client waits, and waits, and waits. Once it finds the client perhaps it cache’s that client’s connection info. This may also explain why changing the session timeout could help - it doesn’t dump the connection, and therefore the information on the client from memory.

That would also explain the success of the NetBIOS and of the same local subnet - if the server cannot find the client via IP, then it falls back upon NetBIOS name resolution. In the same subnet that should still work - if NetBIOS is on.

What if you put your client’s name and address in the Hosts file of the server AND do the same for the client? I wonder if that would make a difference? (Also, have you tried the DEBUG_NAMELOOKUP=1 debug parameter in your client’s INI?)

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

My Win2000 servers were clean installs and not upgrades from NT4. I have had that problem about every second win2000 Domino server I install. Plenty of Admins on the NewsGroup have also had the experience. However the Technote from memory does refer it to upgrades from NT4.I have taken another client dump this time with many more tcp debug entries in the client notes.ini

There are many TCP stack errors and many “check socket ready”.

Client_Clock=1

Debug_Outfile=C:\DebugFile.txt

CONSOLE_LOG_ENABLED=1

DEBUG_NETERRORS=1

DEBUG_TCP_ALL=1

DEBUG_TCP_ERRORS=1

DEBUG_TCP_INITTERM=1

DEBUG_TCP_NAMESERVER=1

DEBUG_TCP_OTHER=1

DEBUG_TCP_POLL=1

DEBUG_TCP_POST=1

DEBUG_TCP_QUEUEIO=1

DEBUG_TCP_QUIET=1

DEBUG_TCP_RECEIVE=1

DEBUG_TCP_RECEIVEDG=1

DEBUG_TCP_SEND=1

DEBUG_TCP_SENDDG=1

DEBUG_TCP_SERVICE=1

DEBUG_TCP_SESSION=1

DEBUG_TCP_TIMING=1

DEBUG_TCP_TMP=1

DEBUG_TCP_WAIT=1

DEBUG_TCP_WKP=1

20/08/2004 08:46:35.91 PM [056C:0002-02E0] TCPEndp_HandleErr> TCP/IP transport driver error function code: 1507h Stack Err: 10035, NTI Err: 000Ch

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> exit hEndp: 0EE00001h dwNtvErr = 00002733h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> Exit: hEndp = 0EE00001h dwNtvErr = 00002733h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] cmd_rcv> exit

16 ms. [68+28=96]

(9-12 [8]) OPEN_NOTE(REP4A256E22:0050840D-NT000008FE,00400020): 20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Snd> x_send completed: len:0036h, nStartByte:0000h, nBytes:0036h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Snd> hEndp: 0EE00001h dwNtvErr = 00000000h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_CheckSocketReady> Exit: iError = 0h, Events = 00000000h

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] TCPEndp_GetSktName> exit dwNtvErr = 00000000h

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] cmd_getprotaddr> UNlocking process

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] cmd_getprotaddr> exit

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] cmd_rcvconnect> Exit hEndp: 12540001h, iError = 0000h dwNtvErr = 0000h

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] TCPEndp_Snd> x_send completed: len:003Ch, nStartByte:0000h, nBytes:003Ch

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] TCPEndp_Snd> hEndp: 12540001h dwNtvErr = 00000000h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] cmd_poll> exit iError = 0h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_CheckSocketReady> Exit Events = 00000001h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_CheckSocketReady> Exit: iError = 0h, Events = 00000001h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> x_recv completed: len:0002h, nStartByte:0000h, nBytes:0002h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> exit hEndp: 0EE00001h dwNtvErr = 00000000h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> receive completed (NTI_RCV)

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> Exit: hEndp = 0EE00001h dwNtvErr = 00000000h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] cmd_poll> Main Poll Loop iError = 0h, dwResult = 1h

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] TCPEndp_Rcv> returned errno: 2733h

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] TCPEndp_HandleErr> TCP/IP transport driver error function code: 1507h Stack Err: 10035, NTI Err: 000Ch

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] TCPEndp_Rcv> exit hEndp: 12540001h dwNtvErr = 00002733h

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] TCPEndp_Rcv> Exit: hEndp = 12540001h dwNtvErr = 00002733h

20/08/2004 08:46:35.92 PM [056C:0004-1404:newmail] cmd_rcv> exit

20/08/2004 08:46:35.92 PM [056C:0002-02E0] cmd_poll> exit iError = 0h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> x_recv completed: len:00C0h, nStartByte:0000h, nBytes:00C0h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> exit hEndp: 0EE00001h dwNtvErr = 00000000h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> receive completed (NTI_RCV)

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> Exit: hEndp = 0EE00001h dwNtvErr = 00000000h

20/08/2004 08:46:35.92 PM [056C:0002-02E0] cmd_rcv> exit

20/08/2004 08:46:35.92 PM [056C:0002-02E0] TCPEndp_Rcv> returned errno: 2733h

20/08/2004 08:46:35.94 PM [056C:0002-02E0] TCPEndp_HandleErr> TCP/IP transport driver error function code: 1507h Stack Err: 10035, NTI Err: 000Ch

/08/2004 10:57:31.80 PM [056C:0002-02E0] cmd_poll> exit iError = 0h

20/08/2004 10:57:31.80 PM [056C:0002-02E0] TCPEndp_CheckSocketReady> Exit: iError = 0h, Events = 00000000h

20/08/2004 10:57:31.80 PM [056C:0002-02E0] cmd_poll> exit iError = 0h

20/08/2004 10:57:32.30 PM [056C:0002-02E0] TCPEndp_CheckSocketReady> Exit: iError = 0h, Events = 00000000h

20/08/2004 10:57:32.30 PM [056C:0002-02E0] cmd_poll> exit iError = 0h

20/08/2004 10:57:32.30 PM [056C:0002-02E0] TCPEndp_CheckSocketReady> Exit: iError = 0h, Events = 00000000h

20/08/2004 10:57:32.30 PM [056C:0002-02E0] cmd_poll> exit iError = 0h

20/08/2004 10:57:32.80 PM [056C:0002-02E0] TCPEndp_CheckSocketReady> Exit: iError = 0h, Events = 00000000h

20/08/2004 10:57:32.80 PM [056C:0002-02E0] cmd_poll> exit iError = 0h

20/08/2004 10:57:32.80 PM [056C:0002-02E0] TCPEndp_CheckSocketReady> Exit: iError = 0h, Events = 00000000h

20/08/2004 10:57:32.80 PM [056C:0002-02E0] cmd_poll> exit iError = 0h

20/08/2004 10:57:32.80 PM [056C:0002-02E0] cmd_close> x_shutdown(SendRecv) dwNtvErr = 0000h

20/08/2004 10:57:32.80 PM [056C:0002-02E0] TCPEndp_MakeUnbound> Post x_soclose: Skt: 00000290h, dwNtvErr = 0000h

20/08/2004 10:57:32.80 PM [056C:0002-02E0] TCPEndp_MakeUnbound> Global SocketCnt = 0

20/08/2004 10:57:32.80 PM [056C:0002-02E0] TCPEndp_MakeUnbound> exit hEndp: 0EE00001h dwNtvErr = 00000000h

20/08/2004 10:57:32.80 PM [056C:0002-02E0] cmd_close> exit hEndp: 0EE00001h dwNtvErr = 00000000h

61141 ms. [-1643±10412=-12055] (Network operation did not complete in a reasonable amount of time; please retry)

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Excellent.

So, the network is sending back 2733h which is 10035 which is the WinSock error WSAEWOULDBLOCK.

WSAEWOULDBLOCK (10035)

Resource temporarily unavailable

This error is returned from operations on non-blocking sockets that cannot be completed immediately, for example recv when no data is queued to be read from the socket. It is a non-fatal error, and the operation should be retried later. It is normal for WSAEWOULDBLOCK to be reported as the result from calling connect on a non-blocking SOCK_STREAM socket, since some time must elapse for the connection to be established.

Same problem in Pegasus Mail (They blame something on the server side.): http://ftp.uni-koeln.de/pc/mirrors/winsock-l/archive/1995/May.95

And, if I remember my propoganda correctly, in ND6 ansychronous communication was a new feature, permitting ASYNC replication, mail checking, etc. So, this fits.

Perhaps your Windows boxes are configured differently than ours. Especially the W2K Servers:

This is scary:

So, either there is no data to be received and the client is timing out - OR - the client isn’t correctly using Winsocks. If it were #2, then everybody’s client should be having that problem. We are not. So, let’s look at #1. Why isn’t any data being received? - Especially when all other Windows apps go across the network just fine.

It appears that for mail - either large attachments or errors on the server side the offending source is the server.

I think a closer look at your server setup, and especially finding more info about that MS Hotfix may be good. The problem is most likely not on the ND6.5 client. (Although perhaps you could mask it by upping the TCP/IP timeout value )

(BTW, Are you doing any encryption to and from the clients?)

Subject: Well…maybe not…

I wrote:

I think a closer look at your server setup, and especially finding more info about that MS Hotfix may be good. The problem is most likely not on the ND6.5 client. (Although perhaps you could mask it by upping the TCP/IP timeout value )

But, thinking further, that may not be correct. If the Server and Clients are in the same subnet, everything works fine.

So, SOMETHING is trashing the network communications (at least async network communications). Is there no way to swap that router with some other model?

Or, can you test ASYNC calls with some tool? (The only that come to mind now are not tools but multiplayer action games).

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Hi Victor,Interesting stuff. Tomorrow I plan to pick up a 50metre lan cable that will allow me to easily connect to either the private lan or the server lan from the same workstation. I think it only makes sense to use the same client in the two different locations to be 100% sure on that point, and also get some debug logs from both.

I’ll look at organising something with the routers in the next day or two after the private v’s server lan testing.

For two different Notes domains on two different servers, one has encryption and compression enabled and the other has neither, but both exhibit the problem.

I need a bit of digestion time as I don’t fully understand all of your comments as yet. I am an MCSE so have a basic understanding of TCP but not to any depth. Also my work rarely involves me in this area.

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

as I don’t fully understand all of your comments as yet

I apologize - I didn’t know that. And, you are correct that these socket errors would largely be unknown to an MSCE - they are for application developers. But, they provide a clue.

My point was that this particular socket error is due to either (1) The software (i.e. Notes Client in this case) requesting data when there isn’t any data available, or (2) The software is requesting network access, but (a) the network access is not available or (b) the request for network access was made incorrectly.

The error itself is not fatal, and the software should retry, which I think Notes does do, thus the long wait before it gives up.

If it were #2a, then I would expect you to have the problem with non-Notes software as well, which you do not. So, we can put this aside for the moment.

If it were #2b, then ALL Notes Clients should suffer from the problem you are experiencing as it would be due to a software programming error. But, we know that this is not the case as our ND6.5 clients operate just fine. So, we can put this aside as well.

Therefore, it is more probable that #1 is to blame. In this case, the Notes Client EXPECTS data, but none is available.

That would then mean some type of EXPECTED communication to the client has not occured. Since we know from your experience with this problem that this issue does not occur when the server and client are on the same subnet, it is therefore likely that:

(a) the Client THINKS it communicated to the server and then waits for a response. But, in reality, it never did. It may have communicated to something in front of the server (router maybe?), but not the server itself. Or, the server is experiencing problems with its sockets layer, received the communication, but the data was never sent to Domino. (Thus, the reference to the MS HotFix in my previous post).

-or-

(b) the server cannot communicate with the client regularly due to some network configuration on the server (e.g. missing DNS, etc) or error with the sockets layer (again the MS HotFix may apply here).

and/or

(c) there is some part of the network that is interfering with the asynchronous communication between the client and the server.

I think your test with the cable to bypass the router will shed great light on the matter.

Good Luck!

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

Hi Victor,I cabled in my workstation into the same subnet as Domino and had no problems. This bypassed the private LAN router.

I then setup one of the servers with an extra NIC card to use Internet Connection Sharing in place of the private LAN router. When I did that my problem immediately came back again.

This actually sounds like the good old “It’s not a Notes problem MTU issue”, that only affects Notes (and no other product), especially when I read your comments in your last response.

Typically with the MTU problem the packets are fragmentated by the client (or perhaps router) to fit through the pipe size. At the server port you see the message (with a protocal analyser)“Packets cannot be reassembled in a reasonable time”. That message never gets back to the client, it just sits there until a timeout occurs. This sounds similar to the problem we are having? What do you think?

In my case there are two classes of users:

1/ Like myself on the private LAN and

2/ Users who connect via the Internet. They connect through an ADSL router which port forwards a bunch of IP’s.

Both of these sets of users are affected in the same way. The problem with the ADSL router is that I cannot change the MTU of the WAN port as the latest uCode complies with the latest RFC which says it must be auto-set so it overrides my manual setting.

The only difference between your server setup and mine should be the region preference, you select US and I select AU.

I wonder if router uCode compliance changes are incompatible with the Notes “do-it-yourself” TCP driver?

In the dark of night routers have their uCode upgraded and never any problems reported the next day.

That might explain why we are suddenly seeing mass problems with this issue accross R5 and R6 whereas it did not exist before.

John

Subject: RE: Show stopper for R5 to R6 upgrades - Intermittent connection

This sounds similar to the problem we are having? What do you think?

I believe that you have found it! But, I have to wonder why Notes doesn’t have the same issue with our routers?

The only difference between your server setup and mine should be the region preference, you select US and I select AU. I wonder if router uCode compliance changes are incompatible with the Notes “do-it-yourself” TCP driver?

I will check with our Network guys to see (a) what kind of routers we are using and (b) what these settings are so that we can compare.