You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

K8s集群MongoDB副本集Pod周期性故障:DBPathInUse只读文件系统问题

问题分析与解决方案

核心关联确认

节点内核存储I/O错误是导致MongoDB Pod崩溃循环的直接原因。当节点底层存储出现IO异常时,Linux内核会自动将挂载的文件系统置为只读模式(防止数据进一步损坏),此时MongoDB无法写入/bitnami/mongodb/data/db/mongod.lock锁文件,触发DBPathInUse错误并进入崩溃循环。


一、排查验证步骤

  1. 节点内核日志检查
    查看Pod崩溃时所在节点的内核日志(dmesg或/var/log/messages),确认是否存在以下类错误:

    I/O error, dev sdX, sector XXXXX
    Buffer I/O error on device dm-X, logical block X
    Remounting filesystem read-only
    这类日志直接证明节点存储栈(本地磁盘、RAID控制器、Longhorn组件)存在故障。

  2. Longhorn组件状态排查

    • 查看Longhorn Manager和Instance Manager的日志,定位卷挂载或IO超时问题:
      kubectl logs -n longhorn-system -l app=longhorn-manager
      kubectl logs -n longhorn-system -l app=longhorn-instance-manager
      
    • 检查MongoDB对应PVC的Longhorn卷状态,确认是否出现Degraded或Faulted:
      kubectl get volumes.longhorn.io -n longhorn-system
      

二、修复与优化措施

1. 节点底层存储故障修复

  • 若为本地磁盘损坏:替换故障磁盘,重新配置RAID(如有),确保节点存储硬件稳定。
  • 若为控制器驱动问题:更新节点磁盘控制器驱动至官方最新稳定版,解决兼容性导致的IO异常。

2. Longhorn存储层优化

  • 启用自动故障转移:确保节点故障时,Longhorn自动将卷切换到健康节点的副本,避免Pod绑定在故障节点。
  • 提升卷副本数:将Longhorn卷的副本数设置为3(匹配集群节点数,至少2个以上),增强存储冗余。修改MongoDB Helm的values.yaml:
    persistence:
      storageClass: longhorn
      volumeAttributes:
        numberOfReplicas: "3"
        staleReplicaTimeout: "30"
    
  • 调整IO超时阈值:在Longhorn UI的Settings -> General中,调大Replica Timeout和Instance Manager Timeout,避免短暂IO波动导致卷被标记为故障。

3. MongoDB Pod稳定性优化

  • 配置Pod反亲和性:避免MongoDB副本集Pod调度到同一节点,降低单点存储故障的影响:
    affinity:
      podAntiAffinity:
        requiredDuringSchedulingIgnoredDuringExecution:
        - labelSelector:
            matchExpressions:
            - key: app.kubernetes.io/name
              operator: In
              values:
              - mongodb
          topologyKey: kubernetes.io/hostname
    
  • 优化健康检查:调整livenessProbe和readinessProbe的阈值,避免误判Pod状态:
    livenessProbe:
      exec:
        command: ["mongo", "--disableImplicitSessions", "--eval", "db.adminCommand('ping')"]
      initialDelaySeconds: 30
      periodSeconds: 10
      timeoutSeconds: 5
      failureThreshold: 6
    readinessProbe:
      exec:
        command: ["mongo", "--disableImplicitSessions", "--eval", "db.adminCommand('ping')"]
      initialDelaySeconds: 5
      periodSeconds: 10
      timeoutSeconds: 5
      failureThreshold: 3
    

三、临时恢复方法

当Pod出现崩溃循环时:

  1. 删除故障Pod,Kubernetes会将其重新调度到健康节点,Longhorn自动挂载卷的健康副本:
    kubectl delete pod <mongodb-pod-name> -n <namespace>
    
  2. 若卷处于Degraded状态,手动触发Longhorn卷重建:
    kubectl patch volumes.longhorn.io <volume-name> -n longhorn-system --type merge -p '{"spec":{"restoreRequested":true}}'
    

内容的提问来源于stack exchange,提问作者user3142695

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 15:20:32