SmartOS/illumos/OpenSolaris环境下NTP卡在INIT阶段且服务器时钟漂移问题求助
问题背景
我之前在SmartOS邮件列表问过同样的问题但没得到回复,最近发现其中一台服务器时钟差了好几分钟,排查NTP时发现了异常:
运行ntpq -c peers输出如下:
remote refid st t when poll reach delay offset jitter beyond.work.de .INIT. 16 u - 1024 0 0.000 0.000 0.000 ntp1.m-online.n .INIT. 16 u - 1024 0 0.000 0.000 0.000
所有服务器都卡在.INIT.状态,看起来根本没完成初始化。
重启ntpd后的日志如下(之后就再没新日志了):
6 Aug 14:10:47 ntpd[53519]: ntpd exiting on signal 15 (Terminated) 6 Aug 14:10:47 ntpd[53519]: 213.238.32.2 local addr 212.12.41.94 -> <null> 6 Aug 14:10:47 ntpd[53519]: 144.76.43.40 local addr 212.12.41.94 -> <null> 6 Aug 14:10:47 ntpd[55024]: Listen and drop on 0 v6wildcard [::]:123 6 Aug 14:10:47 ntpd[55024]: Listen and drop on 1 v4wildcard 0.0.0.0:123 6 Aug 14:10:47 ntpd[55024]: Listen normally on 2 lo0 [::1]:123 6 Aug 14:10:47 ntpd[55024]: Listen normally on 3 lo0 127.0.0.1:123 6 Aug 14:10:47 ntpd[55024]: Listen normally on 4 igb0 212.12.41.94:123 6 Aug 14:10:47 ntpd[55024]: Listening on routing socket on fd #53 for interface updates 6 Aug 14:10:47 ntpd[55024]: kernel reports TIME_ERROR: 0x41: Clock Unsynchronized 6 Aug 14:10:47 ntpd[55024]: kernel reports TIME_ERROR: 0x41: Clock Unsynchronized
我的ntp.conf配置:
driftfile /var/ntp/ntp.drift logfile /var/log/ntp.log # Ignore all network traffic by default restrict default ignore restrict -6 default ignore # Allow localhost to manage ntpd restrict 127.0.0.1 restrict -6 ::1 # Allow servers to reply to our queries restrict source nomodify noquery notrap # Allow connections from trusted NTP servers restrict pool.ntp.org mask 255.255.255.255 nomodify notrap noquery restrict 213.238.32.2 mask 255.255.255.255 nomodify notrap noquery # Time Servers server 213.238.32.2 burst iburst minpoll 4 server pool.ntp.org iburst
目前没有防火墙规则,ntpdate能正常同步时钟,但这样长期下来肯定会再次漂移,想知道问题出在哪?
排查与解决建议
我之前在SmartOS环境碰过类似的NTP同步问题,给你几个实用的排查方向:
检查restrict规则是否阻碍了同步
你配置里的restrict pool.ntp.org mask 255.255.255.255 ...可能有问题——pool.ntp.org会解析成多个动态IP,用255.255.255.255的子网掩码根本覆盖不到所有节点。建议把这条改成restrict pool.ntp.org nomodify notrap,同时去掉noquery参数(这个参数会阻止外部服务器向你的ntpd发送响应数据包)。另外,restrict source那行也可以去掉noquery,改成restrict source nomodify notrap试试。验证网络数据包是否正常收发
虽然你说没有防火墙,但还是可以抓包确认ntpd的请求是否真的发出去、有没有收到响应。在服务器上运行snoop port 123(或者tcpdump udp port 123),观察是否有从本地IP到NTP服务器的请求包,以及服务器返回的响应包。如果只有请求没有响应,那可能是上游网络的问题;如果有响应但ntpd没处理,那大概率是配置的问题。先手动同步时钟再重启ntpd
日志里的TIME_ERROR: 0x41: Clock Unsynchronized提示时钟未同步,如果初始时钟偏差太大(比如超过1000秒),ntpd默认不会自动调整。你已经用ntpdate同步过了,这时候重启ntpd,应该能让它顺利进入同步状态,而不是卡在INIT。检查driftfile的权限
driftfile /var/ntp/ntp.drift需要让ntpd进程有写入权限,否则它无法保存时钟漂移数据,可能导致同步异常。运行ls -l /var/ntp/ntp.drift看看文件所有者和权限,如果不是ntpd运行用户(一般是ntp用户),用chown ntp:ntp /var/ntp/ntp.drift调整一下。调整minpoll参数
你给213.238.32.2设置了minpoll 4(也就是16秒一次查询),这个频率可能有些NTP服务器会拒绝,建议改成默认的minpoll 6(64秒)试试,避免被服务器限流。
备注:内容来源于stack exchange,提问作者Adrian Gschwend

