You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Nebula 3.8.0集群重启后所有服务触发段错误求助

3节点Nebula 3.8.0集群重启后主节点服务持续段错误排查求助

环境信息

  • 平台:搭载NVMe的SuperMicro刀片Proxmox VE HA集群
  • 操作系统:Ubuntu 20.04.6 LTS
  • 虚拟机配置:8核、17GB内存、virtio-scsi-single
  • 存储:NVMe存储池

问题现象

系统重启后主节点的metad、graphd、storaged三个Nebula Graph服务均持续触发段错误,服务启动后立即退出,无常规服务日志,仅日志文件夹内存在.dmp文件。此前已出现过一次相同问题,重装二进制文件无效,只能销毁虚拟机重建但丢失所有数据。

内核软锁导致重启的系统日志

Jul 01 07:23:10 nebulaprod1 kernel: watchdog: BUG: soft lockup - CPU#1 stuck for 23s! [kworker/1:1:818312]
Jul 01 07:23:47 nebulaprod1 kernel: Modules linked in: tcp_diag inet_diag dm_multipath...
Jul 01 07:23:48 nebulaprod1 kernel: Workqueue: events drm_fb_helper_dirty_work [drm_kms_helper]

服务持续段错误的系统日志

Jul 01 15:31:33 nebulaprod1 kernel: nebula-metad[6088]: segfault at 0 ip 0000000001abca71 sp 00007ffc24bbed20 error 4 in nebula-metad[fbf000+1c09000]
Jul 01 15:31:33 nebulaprod1 kernel: nebula-graphd[6100]: segfault at 0 ip 0000000001d634a1 sp 00007ffd38358330 error 4 in nebula-graphd[f4d000+195d000]
Jul 01 15:31:33 nebulaprod1 kernel: nebula-storaged[6112]: segfault at 0 ip 0000000001bd5591 sp 00007ffce64e89f0 error 4 in nebula-storaged[1004000+1cde000]

Metad崩溃前的服务日志

I20250701 07:23:52.686851  5484 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9779, role = STORAGE
I20250701 07:23:58.039000  5484 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.52":9669, role = GRAPH
I20250701 07:24:02.326965  5484 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9669, role = GRAPH
I20250701 07:24:02.458894  5483 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.51":9669, role = GRAPH
I20250701 07:24:02.569334  5485 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.51":9779, role = STORAGE
I20250701 07:24:02.690387  5481 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9779, role = STORAGE
I20250701 07:24:11.225418  5481 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.52":9669, role = GRAPH
I20250701 07:24:13.905009  5481 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9669, role = GRAPH
I20250701 07:24:13.905243  5485 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.51":9669, role = GRAPH
I20250701 07:24:13.916110  5483 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.53":9779, role = STORAGE
I20250701 07:24:18.133842  5482 HBProcessor.cpp:33] Receive heartbeat from "192.168.5.51":9779, role = STORAGE
I20250701 07:25:56.450991  5484 AuthenticationProcessor.cpp:268] List User failed, error: E_LEADER_CHANGED
I20250701 07:25:56.450992  5482 AuthenticationProcessor.cpp:268] List User failed, error: E_LEADER_CHANGED
I20250701 07:25:56.450992  5483 AuthenticationProcessor.cpp:268] List User failed, error: E_LEADER_CHANGED
W20250701 07:25:57.479228  5251 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response
W20250701 07:25:57.479349  5343 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response
W20250701 07:25:57.479182  5345 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response
W20250701 07:25:57.479182  5342 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response
W20250701 07:25:57.479182  5344 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response
W20250701 07:25:23.667162  5340 Host.cpp:439] [Port: 9560, Space: 0, Part: 0] [Host: 192.168.5.52:9560] Pasue this host because long time no heartbeat response

服务启动测试结果

apocalypse0@nebulaprod1:/usr/local/nebula/scripts$ sudo ./nebula.service status all
[INFO] nebula-metad(7458486): Exited
[INFO] nebula-graphd(7458486): Exited
[INFO] nebula-storaged(7458486): Exited

apocalypse0@nebulaprod1:/usr/local/nebula/scripts$ sudo ./nebula.service start all
[INFO] Starting nebula-metad...
[INFO] Done
[INFO] Starting nebula-graphd...
[INFO] Done
[INFO] Starting nebula-storaged...
[INFO] Done

apocalypse0@nebulaprod1:/usr/local/nebula/scripts$ sudo ./nebula.service status all
[INFO] nebula-metad(7458486): Exited
[INFO] nebula-graphd(7458486): Exited
[INFO] nebula-storaged(7458486): Exited

调试步骤建议

1. 分析Core Dump文件

  • 安装调试工具:sudo apt update && sudo apt install gdb
  • 针对metad分析崩溃栈:gdb /usr/local/nebula/bin/nebula-metad /path/to/nebula-metad.dmp(替换为实际dmp文件路径)
  • 在gdb中执行bt命令,输出完整调用栈,定位崩溃代码位置
  • 对graphd和storaged重复上述操作,确认是否为同一原因导致崩溃

2. 检查存储层完整性

  • 检查NVMe磁盘健康状态:sudo smartctl -a /dev/nvme0n1(根据实际磁盘设备调整)
  • 检查文件系统一致性:先卸载对应存储挂载点,再执行sudo fsck /dev/mapper/pve-data(Proxmox存储池对应设备)
  • 验证Nebula数据目录权限:ls -ld /usr/local/nebula/data/*,确保目录及文件归属为nebula用户

3. 排查虚拟化与宿主机问题

  • 查看Proxmox宿主机系统日志,确认是否仍有CPU软锁问题
  • 临时关闭Proxmox HA功能,单独启动故障虚拟机,测试服务是否能正常运行
  • 修改虚拟机磁盘控制器为virtio-scsi(而非virtio-scsi-single),重启后测试服务启动情况

4. 验证集群数据一致性

  • 从其他正常节点备份metad数据目录,覆盖故障节点的对应目录(先备份故障节点原数据),尝试启动服务
  • 检查Nebula配置文件(/usr/local/nebula/etc/nebula-metad.conf、nebula-graphd.conf、nebula-storaged.conf),确认节点地址、端口、数据目录等配置正确

5. 版本验证测试

  • 尝试升级到Nebula 3.8.x系列最新补丁版本,或回退到3.7.x稳定版本,排除版本特定bug

内容的提问来源于stack exchange,提问作者hazzardousmonk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 18:54:53