Nebula 3.8.0集群重启后所有服务触发段错误求助
3节点Nebula 3.8.0集群重启后主节点服务持续段错误排查求助
环境信息
- 平台:搭载NVMe的SuperMicro刀片Proxmox VE HA集群
- 操作系统:Ubuntu 20.04.6 LTS
- 虚拟机配置:8核、17GB内存、virtio-scsi-single
- 存储:NVMe存储池
问题现象
系统重启后主节点的metad、graphd、storaged三个Nebula Graph服务均持续触发段错误,服务启动后立即退出,无常规服务日志,仅日志文件夹内存在.dmp文件。此前已出现过一次相同问题,重装二进制文件无效,只能销毁虚拟机重建但丢失所有数据。
内核软锁导致重启的系统日志
Jul 01 07:23:10 nebulaprod1 kernel: watchdog: BUG: soft lockup - CPU#1 stuck for 23s! [kworker/1:1:818312] Jul 01 07:23:47 nebulaprod1 kernel: Modules linked in: tcp_diag inet_diag dm_multipath... Jul 01 07:23:48 nebulaprod1 kernel: Workqueue: events drm_fb_helper_dirty_work [drm_kms_helper]
服务持续段错误的系统日志
Jul 01 15:31:33 nebulaprod1 kernel: nebula-metad[6088]: segfault at 0 ip 0000000001abca71 sp 00007ffc24bbed20 error 4 in nebula-metad[fbf000+1c09000] Jul 01 15:31:33 nebulaprod1 kernel: nebula-graphd[6100]: segfault at 0 ip 0000000001d634a1 sp 00007ffd38358330 error 4 in nebula-graphd[f4d000+195d000] Jul 01 15:31:33 nebulaprod1 kernel: nebula-storaged[6112]: segfault at 0 ip 0000000001bd5591 sp 00007ffce64e89f0 error 4 in nebula-storaged[1004000+1cde000]
Metad崩溃前的服务日志
I20250701 07:23:52.686851 5484 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9779, role = STORAGE I20250701 07:23:58.039000 5484 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.52":9669, role = GRAPH I20250701 07:24:02.326965 5484 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9669, role = GRAPH I20250701 07:24:02.458894 5483 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.51":9669, role = GRAPH I20250701 07:24:02.569334 5485 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.51":9779, role = STORAGE I20250701 07:24:02.690387 5481 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9779, role = STORAGE I20250701 07:24:11.225418 5481 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.52":9669, role = GRAPH I20250701 07:24:13.905009 5481 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9669, role = GRAPH I20250701 07:24:13.905243 5485 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.51":9669, role = GRAPH I20250701 07:24:13.916110 5483 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9779, role = STORAGE I20250701 07:24:18.133842 5482 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.51":9779, role = STORAGE I20250701 07:25:56.450991 5484 AuthenticationProcessor.cpp:268] List User failed, error: E_LEADER_CHANGED I20250701 07:25:56.450992 5482 AuthenticationProcessor.cpp:268] List User failed, error: E_LEADER_CHANGED I20250701 07:25:56.450992 5483 AuthenticationProcessor.cpp:268] List User failed, error: E_LEADER_CHANGED W20250701 07:25:57.479228 5251 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response W20250701 07:25:57.479349 5343 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response W20250701 07:25:57.479182 5345 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response W20250701 07:25:57.479182 5342 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response W20250701 07:25:57.479182 5344 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response W20250701 07:25:23.667162 5340 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response
服务启动测试结果
apocalypse0@nebulaprod1:/usr/local/nebula/scripts$ sudo ./nebula.service status all [INFO] nebula-metad(7458486): Exited [INFO] nebula-graphd(7458486): Exited [INFO] nebula-storaged(7458486): Exited apocalypse0@nebulaprod1:/usr/local/nebula/scripts$ sudo ./nebula.service start all [INFO] Starting nebula-metad... [INFO] Done [INFO] Starting nebula-graphd... [INFO] Done [INFO] Starting nebula-storaged... [INFO] Done apocalypse0@nebulaprod1:/usr/local/nebula/scripts$ sudo ./nebula.service status all [INFO] nebula-metad(7458486): Exited [INFO] nebula-graphd(7458486): Exited [INFO] nebula-storaged(7458486): Exited
调试步骤建议
1. 分析Core Dump文件
- 安装调试工具:
sudo apt update && sudo apt install gdb - 针对metad分析崩溃栈:
gdb /usr/local/nebula/bin/nebula-metad /path/to/nebula-metad.dmp(替换为实际dmp文件路径) - 在gdb中执行
bt命令,输出完整调用栈,定位崩溃代码位置 - 对graphd和storaged重复上述操作,确认是否为同一原因导致崩溃
2. 检查存储层完整性
- 检查NVMe磁盘健康状态:
sudo smartctl -a /dev/nvme0n1(根据实际磁盘设备调整) - 检查文件系统一致性:先卸载对应存储挂载点,再执行
sudo fsck /dev/mapper/pve-data(Proxmox存储池对应设备) - 验证Nebula数据目录权限:
ls -ld /usr/local/nebula/data/*,确保目录及文件归属为nebula用户
3. 排查虚拟化与宿主机问题
- 查看Proxmox宿主机系统日志,确认是否仍有CPU软锁问题
- 临时关闭Proxmox HA功能,单独启动故障虚拟机,测试服务是否能正常运行
- 修改虚拟机磁盘控制器为
virtio-scsi(而非virtio-scsi-single),重启后测试服务启动情况
4. 验证集群数据一致性
- 从其他正常节点备份metad数据目录,覆盖故障节点的对应目录(先备份故障节点原数据),尝试启动服务
- 检查Nebula配置文件(
/usr/local/nebula/etc/nebula-metad.conf、nebula-graphd.conf、nebula-storaged.conf),确认节点地址、端口、数据目录等配置正确
5. 版本验证测试
- 尝试升级到Nebula 3.8.x系列最新补丁版本,或回退到3.7.x稳定版本,排除版本特定bug
内容的提问来源于stack exchange,提问作者hazzardousmonk
相关产品推荐
相关产品推荐

