You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ceph集群故障求助:迁移节点后第三节点离线,OSD/MON异常及数据冗余降级

Hey there, let’s walk through fixing your Ceph cluster issue step by step. Moving a node to a new physical spot almost always triggers network or connectivity hiccups first, so we’ll start there and work our way through to restoring redundancy.

1. First, Diagnose Network Connectivity Issues

Physical location changes often break cluster communication—this is the most common root cause:

  • From an online node, test basic connectivity to the offline node using both IP and hostname:
    ping <offline-node-ip>
    ping <offline-node-hostname>
    
  • Check overall cluster status to confirm which components are offline:
    ceph -s
    ceph osd tree
    
    Look for down/out statuses next to the migrated node’s MON or OSDs.
  • Verify firewall rules on the migrated node and its new network environment. Ceph requires these ports to be open:
    • 6789 for MON communication
    • 6800–7300 for OSD replication and heartbeat
    • 3300 if you’re running MDS services
  • Double-check time synchronization: Ceph is extremely sensitive to clock drift (anything over 0.5s can cause nodes to be marked offline). Use these commands to verify:
    chronyc sources  # For chrony
    ntpq -p          # For NTP
    
2. Bring Offline OSD/MON Services Back Online

If network checks pass, focus on the Ceph services themselves:

  • Log into the migrated node and check if MON/OSD services are running:
    # For MON services
    systemctl status ceph-mon@<your-node-hostname>
    # For OSD services (replace <osd-id> with actual IDs)
    systemctl status ceph-osd@<osd-id>
    
  • If services aren’t running, start them manually:
    systemctl start ceph-mon@<your-node-hostname>
    systemctl start ceph-osd@<osd-id>
    
  • If services fail to start, check logs for specific errors (this will tell you exactly what’s broken—disk mounts, permissions, or config issues):
    journalctl -u ceph-mon@<your-node-hostname> -f
    journalctl -u ceph-osd@<osd-id> -f
    
  • For MON nodes with a changed IP after migration:
    1. First, remove the old MON entry from the cluster (ensure at least 2 MONs are online first):
      ceph mon remove <old-mon-name>
      
    2. Add the MON with its new IP:
      ceph mon add <new-mon-name> <new-node-ip>:6789
      
  • For OSDs, confirm their associated disks are mounted correctly:
    df -h
    ceph-disk list
    
    If a disk isn’t mounted, fix the mount point and restart the OSD service.
3. Restore Data Redundancy

Once the node is back online, Ceph will automatically start rebalancing data, but you’ll need to monitor and assist the process:

  • Check detailed health status to track degraded objects and recovery progress:
    ceph health detail
    
  • Monitor pool-level recovery stats:
    ceph osd pool stats
    
  • If recovery is too slow, temporarily increase the maximum active recovery threads (adjust based on your cluster’s performance—don’t set this too high):
    ceph osd set recovery_max_active 10
    
    Remember to reset this to the default (3) once recovery completes to avoid impacting production traffic.
  • If some objects remain degraded, check for disk errors on the migrated node using smartctl:
    smartctl -a /dev/<osd-disk>
    
    For persistent object corruption, use ceph-objectstore-tool (proceed with caution—back up data first before attempting repairs).
4. Prevent This Issue in Future Migrations

To avoid repeating this headache next time:

  • Before moving a node, mark its OSDs as out to let Ceph migrate data to other nodes first:
    ceph osd out <osd-id>
    
    Wait until the cluster reports HEALTH_OK before stopping services and moving the node.
  • After migration, validate network connectivity, firewall rules, and time sync before starting Ceph services.
  • Regularly back up your cluster maps for quick recovery:
    ceph mon getmap > monmap.bin
    ceph osd getmap > osdmap.bin
    

内容的提问来源于stack exchange,提问作者Juan Sebastian Leal Rojas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:54:16