You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CeleryExecutor时Apache Airflow中DAG的存储位置咨询

Understanding DAG Folder Placement for CeleryExecutor in Apache Airflow

Let's break this down clearly to answer your questions about DAG storage and why you're seeing those sample DAGs right now.

First: Why You See Sample DAGs Without a DAG Folder

Those sample DAGs are built into Airflow by default. If you check your airflow.cfg, you’ll probably find load_examples = True enabled. Even if your configured dags_folder doesn’t exist, Airflow will still load these pre-packaged examples to help new users get familiar with the tool. If you don’t want them cluttering your UI, just set load_examples = False and restart your Airflow services.

Where to Deploy Your DAG Folder with CeleryExecutor

In a CeleryExecutor cluster, your DAG files need to live on specific nodes because different Airflow components rely on them for core functions:

  • Scheduler + Flower Node: The Scheduler’s main job is scanning DAG files, parsing their logic, and queuing task instances. It needs direct access to your DAG folder to do this work. Flower (the Celery monitoring tool) doesn’t need the DAG folder, but since it’s co-located with the Scheduler, you can keep the DAG folder here alongside the Scheduler.
  • WebServer Node: The WebServer reads DAG files to render the UI—displaying DAG graphs, task statuses, and letting you trigger runs. Without access to the DAG folder, the UI won’t show your custom DAGs correctly.
  • Celery Worker Nodes: This is critical! Workers need the exact same DAG files as the Scheduler. When a worker picks up a task from the Celery queue, it needs to load the DAG’s code (like your Python task functions, dependencies, and business logic) to execute the task. If a worker doesn’t have the DAG file, it’ll throw errors about missing DAGs or tasks.
  • MySQL Server Node: No need for the DAG folder here—your database only stores Airflow metadata (task run statuses, DAG configuration details), not the actual DAG code files.

To keep DAGs consistent across all required nodes, use one of these reliable approaches:

  • Shared Storage: Mount a network file system (like NFS) to the same /home/airflow/dags path on the Scheduler, WebServer, and Worker nodes. Any update to DAGs in the shared storage will be instantly available to all components.
  • Version Control + Automation: Store your DAGs in a Git repo, then use tools like Ansible, Puppet, or a CI/CD pipeline to pull the latest DAG code to all required nodes on a regular basis.
  • Dockerized Deployment: If you’re running Airflow in containers, package your DAGs into a shared Docker image, or use a volume mount to sync DAGs across all container instances.

Quick Action Steps for Your Setup

  1. Create the configured DAG folder on all required nodes: run mkdir -p /home/airflow/dags and ensure the user running Airflow has read/write permissions for this directory.
  2. Copy your custom DAG files into this folder on the Scheduler, WebServer, and both Worker nodes.
  3. Restart the Airflow Scheduler, WebServer, and Worker services to pick up the new DAGs.
  4. (Optional) Disable sample DAGs by setting load_examples = False in airflow.cfg and restarting your services.

内容的提问来源于stack exchange,提问作者gunj_desai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:28:22