UDP sendto()函数偶发耗时过长问题排查问询
UDP sendto()偶发高耗时原因咨询
问题现象
使用UDP的sendto()函数,按UDP协议设计,该函数应在发送/丢弃数据报后立即返回。但在局域网环境下连续发送200M数据包的测试中发现,多数情况下耗时低于1ms,偶尔会超过10ms,甚至达到20~30ms。
测试代码(client.c)
#include <stdio.h> #include <stdlib.h> #include <unistd.h> #include <string.h> #include <arpa/inet.h> #include <time.h> #include <sys/time.h> #define SIZE 1024 * 1024 * 200 int main(int argc, char **argv) { /* if(argc < 2) { printf("needs params\n"); exit(0); } int num = atoi(argv[1]); */ int fd = socket(AF_INET, SOCK_DGRAM, 0); if(fd == -1) { perror("socket"); exit(0); } /* int bufsize = 0; int len = sizeof(int); getsockopt(fd, SOL_SOCKET, SO_SNDBUF, (void *)&bufsize, &len); printf("default: send buff = %d\n", bufsize); bufsize = bufsize * num; printf("bufsize change to %d\n", bufsize); setsockopt(fd, SOL_SOCKET, SO_SNDBUF, (const void *)&bufsize, sizeof(int)); bufsize = -1; //reset len = sizeof(int); getsockopt(fd, SOL_SOCKET, SO_SNDBUF, (void *)&bufsize, &len); printf("after set, send buff = %d\n", bufsize); */ struct sockaddr_in seraddr; seraddr.sin_family = AF_INET; unsigned short port = 49244; seraddr.sin_port = htons(port); inet_pton(AF_INET, "43.82.153.211", &seraddr.sin_addr.s_addr); char buf[1400]; memset(buf, 'g', sizeof(buf)); long long size = SIZE; struct timeval before, after; double mseconds = 0; unsigned int count = 1; while(size > 0) { gettimeofday(&before, NULL); int err = sendto(fd, buf, 1400, 0, (struct sockaddr*)&seraddr, sizeof(seraddr)); gettimeofday(&after, NULL); if (-1 == err) { perror("sendto"); exit(0); } mseconds = 1000 * (after.tv_sec - before.tv_sec) + (double)(after.tv_usec - before.tv_usec) / 1000; printf("%d: %-0.6f mseconds\n", count, mseconds); memset(buf, 0, sizeof(buf)); size -= 1400; count++; } close(fd); return 0; }
测试日志片段
... 546: 0.074000 mseconds 547: 0.083000 mseconds 548: 0.041000 mseconds 549: 0.067000 mseconds 550: 0.072000 mseconds 551: 0.048000 mseconds 552: 7.541000 mseconds(Much higher than other values) 553: 0.082000 mseconds 554: 0.084000 mseconds ...
参数调整尝试(运行环境:ARM架构,内核版本linux-3.0.27)
1. 默认行为
# time ./client > client.log real 0m 18.19s user 0m 2.35s sys 0m 4.78s
仅打印耗时>10ms的记录:
# time ./client_print_largethan_10ms > client_print_largethan_10ms.log real 0m 17.88s user 0m 0.50s sys 0m 4.44s
2. 扩大sendbuf至8倍(内核会翻倍该值以预留管理开销)
先调整系统允许的最大sendbuf值:
# sysctl -a | grep "net.core.wmem_" net.core.wmem_max = 110592 net.core.wmem_default = 110592 # sysctl -w net.core.wmem_max=2211840 net.core.wmem_max = 2211840 # sysctl -a | grep "net.core.wmem_" net.core.wmem_max = 2211840 net.core.wmem_default = 110592
再运行测试:
# time ./client_sndbuf_x 10 > client_sndbuf.log real 0m 9.34s user 0m 0.62s sys 0m 4.63s
总耗时减少,但日志中仍存在耗时超10ms的记录。
3. 调整SIOCSIFTXQLEN参数
通过ioctl()将SIOCSIFTXQLEN从1000调整为10000后测试:
# time ./client_print_largethan_10ms 49244 > client_print_largethan_10ms_afterset_SIOCSIFTXQLEN.log real 0m 17.82s user 0m 0.49s sys 0m 4.32s
无明显变化;同时调整sendbuf和SIOCSIFTXQLEN时,总耗时反而增加,恢复SIOCSIFTXQLEN后总耗时回到9秒。
诉求
不需要完全消除耗时,仅希望了解该偶发耗时的产生原因。
内容的提问来源于stack exchange,提问作者Drake Wu
相关产品推荐
相关产品推荐

