Node.js项目部署DigitalOcean遇SIGTERM错误,Pod部分故障求助
Node.js项目部署到DigitalOcean Kubernetes集群出现
SIGTERM故障排查请求 我们将Node.js项目部署到DigitalOcean的Kubernetes集群时,遇到command failed signal SIGTERM错误:7个Pod中有4个正常运行,另外3个启动失败;重启所有Pod后,起初正常运行20-25分钟,但之后又有3个Pod再次故障。相同配置在开发集群可正常运行,请求协助排查。
错误日志
npm notice npm notice New major version of npm available! 8.19.4 -> 10.4.0 npm notice Changelog: <https://github.com/npm/cli/releases/tag/v10.4.0> npm notice Run `npm install -g npm@10.4.0` to update! npm notice npm ERR! path /app npm ERR! command failed npm ERR! signal SIGTERM npm ERR! command sh -c -- node dist/index.js npm ERR! A complete log of this run can be found in: npm ERR!
Kubernetes YAML配置
apiVersion: v1 kind: Service metadata: name: gateway-consumer-node labels: app: gateway-consumer-node annotations: externalTrafficPolicy: Local nginx.ingress.kubernetes.io/enable-cors: "true" spec: type: LoadBalancer ports: - name: http protocol: TCP port: 80 targetPort: 8080 selector: app: gateway-consumer-node --- apiVersion: apps/v1 kind: Deployment metadata: name: gateway-consumer-node spec: selector: matchLabels: app: gateway-consumer-node revisionHistoryLimit: 2 replicas: 7 strategy: rollingUpdate: maxSurge: 8 maxUnavailable: 1 minReadySeconds: 5 template: metadata: labels: app: gateway-consumer-node spec: containers: - name: gateway-consumer-node image: <IMAGE> imagePullPolicy: Always livenessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 10 periodSeconds: 30 timeoutSeconds: 5 ports: - containerPort: 8080 resources: limits: cpu: 100m env:
package.json内容
{ "name": "gateway-consumer-node", "version": "1.0.0", "description": "Client gateway for all the services", "main": "index.js", "scripts": { "dev": "ts-node-dev src/index.ts", "start": "node dist/index.js", "test": "echo \"Error: no test specified\" && exit 1", "build": "tsc -p .", "postbuild": "cp ./src/public/* ./dist/" }, "author": "", "license": "ISC", "dependencies": { "@aws-sdk/client-s3": "^3.137.0", "@farcaster/hub-nodejs": "^0.11.0", "@lighthouse-web3/sdk": "^0.2.8", "aws-sdk": "^2.1183.0", "axios": "^0.27.2", "bs58": "^5.0.0", "connect-timeout": "^1.9.0", "cookie-parser": "^1.4.6", "cookie-session": "^2.0.0", "cors": "^2.8.5", "crypto-js": "^4.1.1", "dotenv": "^16.0.1", "elastic-apm-node": "^3.38.0", "ethers": "^5.5.1", "express": "^4.18.1", "express-async-errors": "^3.1.1", "express-fileupload": "^1.4.0", "express-validator": "^6.14.1", "googleapis": "^108.0.0", "ioredis": "^5.2.3", "jsonwebtoken": "^9.0.0", "morgan": "^1.10.0", "multer": "^1.4.5-lts.1", "multer-s3": "^3.0.1", "node-cron": "^3.0.2", "notion-types": "^6.13.4", "npm": "^8.13.1", "open-graph-scraper": "^5.0.5", "siwe": "^1.1.6", "socket.io": "^4.5.3", "stripe": "^14.14.0", "ts-node": "^10.8.2", "ts-node-dev": "^2.0.0", "tweetnacl": "^1.0.3", "typescript": "^4.7.3", "uuid": "^8.3.2", "web3.storage": "^4.3.0" }, "devDependencies": { "@types/connect-timeout": "^0.0.36", "@types/cookie-parser": "^1.4.3", "@types/cookie-session": "^2.0.44", "@types/cors": "^2.8.12", "@types/ejs": "^3.1.3", "@types/express": "^4.17.13", "@types/express-fileupload": "^1.2.2", "@types/jsonwebtoken": "^8.5.8", "@types/morgan": "^1.9.3", "@types/multer": "^1.4.7", "@types/multer-s3": "^3.0.0", "@types/node": "^18.7.14", "@types/node-cron": "^3.0.7", "@types/open-graph-scraper": "^4.8.1", "@types/uuid": "^8.3.4" } }
排查方向及解决方案
- 资源限制缺失:当前仅设置CPU上限100m,无资源请求配置和内存限制。生产集群负载更高,Pod可能因内存不足被内核发送SIGTERM终止。建议补充资源配置:
resources: requests: cpu: 50m memory: 256Mi limits: cpu: 100m memory: 512Mi - 进程信号未正确处理:Node.js进程未监听SIGTERM信号,导致收到终止信号时异常退出。在应用代码中添加优雅关闭逻辑:
process.on('SIGTERM', () => { console.log('Received SIGTERM, initiating graceful shutdown'); // 关闭HTTP服务器、释放数据库连接等资源 server.close(() => { console.log('Server closed, exiting process'); process.exit(0); }); // 设置超时强制退出,避免无限等待 setTimeout(() => { console.error('Graceful shutdown timed out, forcing exit'); process.exit(1); }, 10000); }); - 健康检查配置不合理:存活探针初始延迟10秒可能不足以让应用完成启动,导致Pod被误判为故障重启。建议延长初始延迟,并添加就绪探针确保Pod就绪后再接收流量:
livenessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 30 periodSeconds: 30 timeoutSeconds: 5 readinessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 15 periodSeconds: 10 timeoutSeconds: 5 - 镜像完整性验证:检查生产环境镜像是否包含完整的
dist目录及依赖。可进入故障Pod执行ls /app/dist和npm list确认文件和依赖状态。 - 环境变量缺失:YAML中
env配置不完整,生产集群可能缺少应用必需的环境变量(如数据库地址、密钥等),导致启动失败。核对所有必填环境变量是否已配置。 - 节点资源不足:检查DigitalOcean集群节点的资源使用情况,执行
kubectl describe nodes查看节点CPU、内存剩余量,若节点资源耗尽,需扩容节点或调整Pod资源配置。
内容的提问来源于stack exchange,提问作者Venkatesh
相关产品推荐
相关产品推荐

