集群升级后OKD openshift-controller-manager卡在Progressing状态如何解决?
排查与修复步骤
确认deployment详细状态:
执行oc describe deployment/route-controller-manager -n openshift-controller-manager,重点查看Conditions和Revision History板块,检查是否存在旧版本Pod残留、更新记录异常的情况。
同时验证所有Pod的版本一致性:oc get pods -n openshift-controller-manager -l deployment=route-controller-manager -o jsonpath='{.items[*].metadata.labels.openshift\.io/revision}',如果输出的版本号不统一,说明更新未彻底完成。强制触发滚动更新:
若确认所有Pod正常运行但状态未同步,手动触发滚动更新让控制器重新校验状态:oc rollout restart deployment/route-controller-manager -n openshift-controller-manager
等待滚动完成后,再次查看集群Operator状态:oc get clusteroperator openshift-controller-manager检查Operator日志与状态缓存:
查看openshift-controller-manager Operator的日志,排查状态判断异常:oc logs deployment/openshift-controller-manager-operator -n openshift-controller-manager-operator -f,搜索RouteControllerManagerProgressing相关内容,确认是否有错误或延迟判断的问题。
若日志显示Operator未正确识别更新后的副本数,可尝试删除其状态缓存(操作前确认无其他异常):oc delete configmap/controller-manager-operator-status -n openshift-controller-manager-operator,Operator会重新生成状态缓存。手动修正ClusterOperator状态(仅上述方法无效时使用):
编辑ClusterOperator的状态配置:oc edit clusteroperator openshift-controller-manager,找到status.conditions中类型为Progressing的条目,将status修改为False,reason设为RouteControllerManagerAvailable,message改为All replicas updated and available。注意:此操作直接修改集群核心状态,必须确保所有Route Controller Manager Pod确实正常运行后再执行。
内容的提问来源于stack exchange,提问作者lweller

