CAM运行数天后cam-mongo Pod宕机,请求排查解决
Looks like your cam-mongo pod is stuck starting because a stale mongod.lock file in your persistent storage is blocking the new mongod instance from grabbing the necessary lock. Here's a step-by-step solution to get your service back up:
Step 1: Access the Persistent Volume and Clean Up the Lock File
The lock file lives in a persistent volume claim (PVC) attached to your pod, so we'll use a temporary container to mount this volume and fix the issue:
First, find the PVC name linked to your cam-mongo deployment:
kubectl get pvc -n servicesLook for the PVC associated with your cam-mongo pod (it might have a name like
cam-mongo-pvcor similar).Spin up a temporary container using the official mongo image (so we have access to mongod repair tools) and mount the same PVC:
kubectl run -n services temp-mongo-fix --image=mongo:latest --rm -it --volume claimName=<YOUR_CAM_MONGO_PVC_NAME>,mountPath=/data/db -- bashReplace
<YOUR_CAM_MONGO_PVC_NAME>with the actual PVC name from step 1.Inside the temporary container, delete the stale lock file:
rm /data/db/mongod.lockRun a database repair to fix any potential corruption from the abnormal shutdown:
mongod --repair --dbpath /data/dbWait for the repair to finish, then type
exitto leave the container. The temporary pod will auto-delete thanks to the--rmflag.
Step 2: Restart the Faulty cam-mongo Pod
With the lock file gone and database repaired, delete the failing pod to let your deployment create a fresh one:
kubectl delete pod -n services cam-mongo-5c89fcccbd-r2hv4
Your deployment will immediately spin up a new pod, which should start successfully without the lock file conflict.
Preventive Tips to Avoid This Issue Again
- Set resource limits: Configure CPU/memory requests and limits for your cam-mongo pod to prevent OOM kills (a top cause of unexpected mongod shutdowns).
- Use MongoDB Replica Sets: Deploying a replica set adds redundancy, handles node failures more gracefully, and reduces the risk of stale lock files.
- Add startup checks (optional): You can tweak your mongo container's startup script to check for
mongod.lockand verify no active mongod process exists before starting. If the lock file is stale, auto-clean it (use this cautiously to avoid accidental data loss).
内容的提问来源于stack exchange,提问作者Gian Filippo Maniscalco

