AirFlow Flower Pod存活/就绪探针失败原因及极简部署配置需求
AirFlow Flower Pod探针失败排查及极简部署配置
一、Flower Pod探针失败排查步骤
1. 验证Flower服务状态
- 查看Pod日志,确认服务是否正常启动:
重点检查是否有绑定5555端口的日志,或启动阶段的报错信息(如依赖缺失、配置错误)。kubectl logs <airflow-flower-pod-name> - 手动在Pod内执行探针命令,验证连通性:
如果同样无法连通,说明Flower服务未正常启动,需聚焦服务本身问题而非探针配置。kubectl exec -it <airflow-flower-pod-name> -- curl localhost:5555
2. 检查资源限制合理性
你已设置资源限制,需确认是否因资源不足导致服务启动失败:
- 查看Pod资源使用情况:
若CPU/内存接近或达到限制值,需调高资源请求或限制阈值。kubectl top pod <airflow-flower-pod-name>
3. 确认ServiceAccount权限
你禁用了ServiceAccount自动创建,需验证手动创建的账号权限:
- 查看Pod关联的ServiceAccount:
kubectl describe pod <airflow-flower-pod-name> | grep ServiceAccount - 检查该ServiceAccount的权限配置,确保其拥有Flower所需的Kubernetes API访问权限(如获取Worker状态)。
4. 调整探针参数
默认探针可能因启动延迟导致误判,可调整探针的启动延迟和超时时间:
flower: livenessProbe: initialDelaySeconds: 60 timeoutSeconds: 10 exec: command: ["curl", "localhost:5555"] readinessProbe: initialDelaySeconds: 30 timeoutSeconds: 5 exec: command: ["curl", "localhost:5555"]
5. 检查镜像内工具存在性
确认AirFlow镜像中是否包含curl命令:
kubectl exec -it <airflow-flower-pod-name> -- which curl
若镜像无curl,需更换探针方式(如使用wget或端口检查命令)。
二、AirFlow极简部署values.yaml
以下配置仅保留核心组件,关闭非必要功能,适合快速测试:
# 禁用内置中间件(需自行准备外部PostgreSQL/Redis) redis: enabled: false postgresql: enabled: false airflow: config: AIRFLOW__CORE__EXECUTOR: CeleryExecutor # 替换为你的外部数据库地址 AIRFLOW__DATABASE__SQL_ALCHEMY_CONN: postgresql+psycopg2://airflow:airflow@postgres-host:5432/airflow AIRFLOW__CELERY__RESULT_BACKEND: db+postgresql://airflow:airflow@postgres-host:5432/airflow # 替换为你的外部Redis地址 AIRFLOW__CELERY__BROKER_URL: redis://redis-host:6379/0 # 禁用ServiceAccount自动创建,指定已存在的账号 serviceAccount: create: false name: your-existing-service-account # 基础资源配置(按需调整) resources: requests: cpu: 100m memory: 256Mi limits: cpu: 500m memory: 512Mi # 核心组件启用配置 webserver: enabled: true scheduler: enabled: true workers: enabled: true replicas: 1 flower: enabled: true livenessProbe: initialDelaySeconds: 60 timeoutSeconds: 10 exec: command: ["curl", "-f", "localhost:5555"] readinessProbe: initialDelaySeconds: 30 timeoutSeconds: 5 exec: command: ["curl", "-f", "localhost:5555"] # 关闭非核心组件 triggerer: enabled: false logs: persistence: enabled: false
若无需分布式执行,可切换为SequentialExecutor(单节点测试用):
airflow: config: AIRFLOW__CORE__EXECUTOR: SequentialExecutor AIRFLOW__DATABASE__SQL_ALCHEMY_CONN: sqlite:////opt/airflow/airflow.db
内容的提问来源于stack exchange,提问作者Makrushin Evgenii
相关产品推荐
相关产品推荐

