Kubernetes Deployment重启策略:能否为其配置失败触发的重启策略?
Great question! Let’s break this down clearly so you understand exactly how this works in Kubernetes, and what your options are here.
First off, a quick clarification: Deployments themselves don’t have a direct restart policy setting. Instead, the restart behavior is controlled by the restartPolicy field defined in the Pod template inside the Deployment’s spec. But there’s a key catch with how this interacts with Deployment’s core functionality that you need to know about.
1. The Three restartPolicy Options for Pods
Every Pod has three possible restart policies:
Always: This is the default for Deployments. No matter if your container exits successfully (exit code 0) or fails (non-zero exit code), Kubernetes will restart the container. This aligns with Deployment’s main job: keeping a fixed number of replica Pods running at all times.OnFailure: This is the one you’re asking about—Kubernetes will only restart the container if it terminates with a non-zero exit code (i.e., it fails). If the container exits successfully, it won’t be restarted.Never: Kubernetes won’t restart the container under any circumstances, whether it succeeds or fails.
2. Using OnFailure with Deployments: The Catch
You can set restartPolicy: OnFailure in your Deployment’s Pod template, but here’s the gotcha: Deployment’s underlying ReplicaSet controller is designed to maintain the exact number of replicas you specify.
Let’s say you have a Deployment with replicas: 1 and restartPolicy: OnFailure, and your container runs a task that exits successfully. The Pod will move to a Completed state—but the ReplicaSet will immediately spin up a new Pod to replace it, because it needs to keep that 1 replica count active. That’s probably not what you’re expecting when you want a "failure-only" restart policy.
Here’s a quick example YAML to illustrate this:
apiVersion: apps/v1 kind: Deployment metadata: name: test-failure-restart spec: replicas: 1 selector: matchLabels: app: test template: metadata: labels: app: test spec: restartPolicy: OnFailure containers: - name: test-container image: busybox command: ["sh", "-c", "echo 'Task finished successfully!' && exit 0"]
Run this, and you’ll see the first Pod complete, then a new one start right away—all because the ReplicaSet is doing its job to maintain the replica count.
3. Better Tools for "Failure-Only Restart" Needs
If your core requirement is "restart only when the container fails, and leave it alone if it succeeds," a Deployment isn’t the best fit. Instead, use these Kubernetes resources:
- Jobs: Built for one-time tasks. When a Job’s container completes successfully, it won’t restart. If it fails, you can configure how many times it retries with the
backoffLimitfield. - CronJobs: For tasks that run on a schedule (like hourly/daily jobs). They work just like Jobs but add scheduling capabilities, and also support failure retries.
If you absolutely have to use a Deployment for some reason, you’d have to manually adjust the replica count to 0 after the container succeeds—but that’s not an automated or reliable solution for most use cases.
Wrap-Up
- You can set
restartPolicy: OnFailurein a Deployment’s Pod template, but the ReplicaSet’s replica maintenance will replace any successfully completed Pods immediately. - For true "restart on failure, stop on success" behavior, use a Job or CronJob instead—they’re purpose-built for this kind of workload.
内容的提问来源于stack exchange,提问作者Wassim Ben Salem

