Commit f706d830 authored by Gang He's avatar Gang He Committed by David Teigland

dlm: make sctp_connect_to_sock() return in specified time

When the user setup a two-ring cluster, DLM kernel module
will automatically selects to use SCTP protocol to communicate
between each node. There will be about 5 minute hang in DLM
kernel module, in case one ring is broken before switching to
another ring, this will potentially affect the dependent upper
applications, e.g. ocfs2, gfs2, clvm and clustered-MD, etc.
Unfortunately, if the user setup a two-ring cluster, we can not
specify DLM communication protocol with TCP explicitly, since
DLM kernel module only supports SCTP protocol for multiple
ring cluster.
Base on my investigation, the time is spent in sock->ops->connect()
function before returns ETIMEDOUT(-110) error, since O_NONBLOCK
argument in connect() function does not work here, then we should
make sock->ops->connect() function return in specified time via
setting socket SO_SNDTIMEO atrribute.
Signed-off-by: default avatarGang He <ghe@suse.com>
Signed-off-by: default avatarDavid Teigland <teigland@redhat.com>
parent b09c603c
...@@ -1037,6 +1037,7 @@ static void sctp_connect_to_sock(struct connection *con) ...@@ -1037,6 +1037,7 @@ static void sctp_connect_to_sock(struct connection *con)
int result; int result;
int addr_len; int addr_len;
struct socket *sock; struct socket *sock;
struct timeval tv = { .tv_sec = 5, .tv_usec = 0 };
if (con->nodeid == 0) { if (con->nodeid == 0) {
log_print("attempt to connect sock 0 foiled"); log_print("attempt to connect sock 0 foiled");
...@@ -1083,8 +1084,19 @@ static void sctp_connect_to_sock(struct connection *con) ...@@ -1083,8 +1084,19 @@ static void sctp_connect_to_sock(struct connection *con)
kernel_setsockopt(sock, SOL_SCTP, SCTP_NODELAY, (char *)&one, kernel_setsockopt(sock, SOL_SCTP, SCTP_NODELAY, (char *)&one,
sizeof(one)); sizeof(one));
/*
* Make sock->ops->connect() function return in specified time,
* since O_NONBLOCK argument in connect() function does not work here,
* then, we should restore the default value of this attribute.
*/
kernel_setsockopt(sock, SOL_SOCKET, SO_SNDTIMEO, (char *)&tv,
sizeof(tv));
result = sock->ops->connect(sock, (struct sockaddr *)&daddr, addr_len, result = sock->ops->connect(sock, (struct sockaddr *)&daddr, addr_len,
O_NONBLOCK); O_NONBLOCK);
memset(&tv, 0, sizeof(tv));
kernel_setsockopt(sock, SOL_SOCKET, SO_SNDTIMEO, (char *)&tv,
sizeof(tv));
if (result == -EINPROGRESS) if (result == -EINPROGRESS)
result = 0; result = 0;
if (result == 0) if (result == 0)
......
Markdown is supported
0%
or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment