OpenShift Route被拒,如何配置ArgoCD/OpenShift自动重试重建?
解决ArgoCD部署中Route因依赖缺失导致Rejected的自动重试问题
针对你遇到的灾难恢复场景下Route因Service未就绪而永久处于Rejected状态的问题,以下是几种可行的解决方案:
1. 配置ArgoCD自定义健康检查+自愈机制
通过给OpenShift Route配置自定义健康规则,让ArgoCD识别Rejected状态为资源不健康,再利用ArgoCD的自愈功能自动重建资源:
- 在ArgoCD Application B的配置中添加Route的健康检查规则:
spec: resourceCustomizations: route.openshift.io/Route: health.lua: | hs = {} hs.status = "Healthy" hs.message = "" if obj.status ~= nil and obj.status.conditions ~= nil then for _, condition in ipairs(obj.status.conditions) do if condition.type == "Ready" and condition.status == "False" and condition.reason == "Rejected" then hs.status = "Degraded" hs.message = "Route rejected: " .. condition.message end end end return hs - 同时确保Application B开启自愈和自动同步:
spec: automated: prune: true selfHeal: true
当Route处于Rejected状态时,ArgoCD会将其标记为Degraded,并自动触发同步重建,直到Route能正常关联到Service。
2. 建立ArgoCD应用间依赖关系
通过给Application B添加依赖规则,强制ArgoCD在Application A完全就绪后再部署Application B,从根源避免竞态条件:
# Application B的配置片段 spec: dependsOn: - name: ApplicationA namespace: argocd
这个方案需要评估是否符合你对两个应用生命周期独立的要求,但在灾难恢复场景下,能彻底消除资源创建顺序问题。
3. 用CronJob定期清理Rejected Route
如果上述方案无法适配你的流程,可以通过OpenShift CronJob定期检查并删除Rejected状态的Route,让ArgoCD自动重建:
apiVersion: batch/v1 kind: CronJob metadata: name: route-retry-cleanup namespace: your-target-namespace spec: schedule: "*/3 * * * *" # 每3分钟执行一次 jobTemplate: spec: template: spec: containers: - name: oc-cli image: image-registry.openshift-image-registry.svc:5000/openshift/cli:latest command: - /bin/sh - -c - | oc get routes -n your-target-namespace -o json | \ jq -r '.items[] | select(.status.conditions[]? | .type=="Ready" and .status=="False" and .reason=="Rejected") | .metadata.name' | \ xargs -r oc delete route -n your-target-namespace restartPolicy: OnFailure serviceAccountName: route-cleanup-sa # 需要创建具有Route删除权限的SA
注意需要提前创建一个拥有目标命名空间Route删除权限的ServiceAccount,并绑定对应的Role。
内容的提问来源于stack exchange,提问作者Fabry
相关产品推荐
相关产品推荐

