Kubernetes中Jenkins Slave离线问题排查求助
使用Kubernetes插件为Jenkins Slave配置云环境,Pod模板采用jenkins/inbound-agent镜像。运行流水线时系统提示jenkins-slave is offline,执行kubectl -n jenkins get pods -A --watch可见Pod状态依次为ContainerCreating、Running、Terminating,随后重复创建新Pod进入循环。
相关配置
service.yaml内容
apiVersion: v1 kind: Service metadata: name: jenkins-svc namespace: jenkins labels: app: jenkins-svc annotations: prometheus.io/scrape: 'true' prometheus.io/path: / prometheus.io/port: '8080' spec: selector: app: jenkins-deployment type: NodePort ports: - port: 8080 targetPort: 8080 nodePort: 32100 --- apiVersion: v1 kind: Service metadata: name: jenkins-jnlp namespace: jenkins labels: app: jenkins-jnlp spec: selector: app: jenkins-deployment type: ClusterIP ports: - port: 50000 targetPort: 50000
Kubernetes云配置详情
- Kubernetes URL:
https://kubernetes.docker.internal:6443 - Kubernetes Namespace: jenkins
- Credentials: Kubernetes Service Account(已连通Kubernetes v1.25.4)
- Jenkins URL:
http://localhost:32100 - Jenkins tunnel:
localhost:50000
已尝试操作:更换Slave镜像、重启Jenkins、修改服务类型,问题仍存在。
1. 修正Jenkins URL和Tunnel地址(核心问题)
jenkins/inbound-agent镜像的Slave Pod需从K8s集群内部访问Jenkins主节点,当前配置的localhost:32100和localhost:50000在Pod内部指向的是Slave自身,而非Jenkins主节点。
修改为集群内部可访问的服务地址:
- Jenkins URL改为:
http://jenkins-svc.jenkins.svc.cluster.local:8080 - Jenkins tunnel改为:
jenkins-jnlp.jenkins.svc.cluster.local:50000
2. 校验JNLP服务与主节点Pod的匹配性
确认jenkins-jnlp Service的selector(app: jenkins-deployment)能精准匹配Jenkins主节点Pod的标签,同时主节点Deployment已正确暴露50000端口。
3. 查看Slave Pod日志定位具体错误
执行命令获取终止Pod的日志,直接排查连接失败、权限不足等具体问题:
kubectl -n jenkins logs <slave-pod-name>
4. 验证Service Account权限
确保配置的Service Account拥有管理Slave Pod的必要权限,需绑定包含以下权限的Role/RoleCluster:
- pods的create、delete、get、list、watch权限
- pods/exec权限
- secrets的get权限(若使用私有镜像)
5. 检查K8s节点资源状态
Slave Pod启动后立即终止,可能是节点CPU/内存不足导致被驱逐,执行命令查看节点资源压力:
kubectl describe nodes
查看Conditions字段是否存在MemoryPressure或CPUPressure状态,若有需扩容节点或调整Slave Pod的资源请求/限制。
内容的提问来源于stack exchange,提问作者DanOps

