Ceph集群故障求助:迁移节点后第三节点离线,OSD/MON异常及数据冗余降级
Hey there, let’s walk through fixing your Ceph cluster issue step by step. Moving a node to a new physical spot almost always triggers network or connectivity hiccups first, so we’ll start there and work our way through to restoring redundancy.
1. First, Diagnose Network Connectivity Issues
Physical location changes often break cluster communication—this is the most common root cause:
- From an online node, test basic connectivity to the offline node using both IP and hostname:
ping <offline-node-ip> ping <offline-node-hostname> - Check overall cluster status to confirm which components are offline:
Look forceph -s ceph osd treedown/outstatuses next to the migrated node’s MON or OSDs. - Verify firewall rules on the migrated node and its new network environment. Ceph requires these ports to be open:
- 6789 for MON communication
- 6800–7300 for OSD replication and heartbeat
- 3300 if you’re running MDS services
- Double-check time synchronization: Ceph is extremely sensitive to clock drift (anything over 0.5s can cause nodes to be marked offline). Use these commands to verify:
chronyc sources # For chrony ntpq -p # For NTP
2. Bring Offline OSD/MON Services Back Online
If network checks pass, focus on the Ceph services themselves:
- Log into the migrated node and check if MON/OSD services are running:
# For MON services systemctl status ceph-mon@<your-node-hostname> # For OSD services (replace <osd-id> with actual IDs) systemctl status ceph-osd@<osd-id> - If services aren’t running, start them manually:
systemctl start ceph-mon@<your-node-hostname> systemctl start ceph-osd@<osd-id> - If services fail to start, check logs for specific errors (this will tell you exactly what’s broken—disk mounts, permissions, or config issues):
journalctl -u ceph-mon@<your-node-hostname> -f journalctl -u ceph-osd@<osd-id> -f - For MON nodes with a changed IP after migration:
- First, remove the old MON entry from the cluster (ensure at least 2 MONs are online first):
ceph mon remove <old-mon-name> - Add the MON with its new IP:
ceph mon add <new-mon-name> <new-node-ip>:6789
- First, remove the old MON entry from the cluster (ensure at least 2 MONs are online first):
- For OSDs, confirm their associated disks are mounted correctly:
If a disk isn’t mounted, fix the mount point and restart the OSD service.df -h ceph-disk list
3. Restore Data Redundancy
Once the node is back online, Ceph will automatically start rebalancing data, but you’ll need to monitor and assist the process:
- Check detailed health status to track degraded objects and recovery progress:
ceph health detail - Monitor pool-level recovery stats:
ceph osd pool stats - If recovery is too slow, temporarily increase the maximum active recovery threads (adjust based on your cluster’s performance—don’t set this too high):
Remember to reset this to the default (3) once recovery completes to avoid impacting production traffic.ceph osd set recovery_max_active 10 - If some objects remain degraded, check for disk errors on the migrated node using
smartctl:
For persistent object corruption, usesmartctl -a /dev/<osd-disk>ceph-objectstore-tool(proceed with caution—back up data first before attempting repairs).
4. Prevent This Issue in Future Migrations
To avoid repeating this headache next time:
- Before moving a node, mark its OSDs as
outto let Ceph migrate data to other nodes first:
Wait until the cluster reportsceph osd out <osd-id>HEALTH_OKbefore stopping services and moving the node. - After migration, validate network connectivity, firewall rules, and time sync before starting Ceph services.
- Regularly back up your cluster maps for quick recovery:
ceph mon getmap > monmap.bin ceph osd getmap > osdmap.bin
内容的提问来源于stack exchange,提问作者Juan Sebastian Leal Rojas
相关产品推荐
相关产品推荐

