如何禁用Airflow重启时所有任务自动启动?解决服务器负载过高问题
It’s super frustrating when restarting Airflow leads to a flood of tasks firing up and cranking your server load—let’s break down the fixes based on your setup.
First, let’s confirm what you already have right: your default_args includes catchup=False, which is great because it stops backfilling past scheduled runs. But there are a few other common culprits here:
1. Clear Pending/Queued Task Instances Before Restart
When you restart Airflow, the scheduler will pick up any tasks that were in a queued, running, or failed state before the restart. To avoid these re-triggering:
- Use the Airflow UI: Go to your DAG, click "Tree View" or "Graph View", select all tasks that aren’t in a "success" state, and click "Clear" (uncheck "Recurse" if you don’t want to clear downstream tasks).
- Or use the CLI command:
airflow tasks clear start_data_collect --state queued --state running --state failed
This removes pending tasks so the scheduler won’t reprocess them on restart.
2. Pause the DAG Before Restarting
A quick, low-effort fix: Pause your DAG in the Airflow UI before restarting the scheduler/webserver. Once the restart is complete, you can unpause it again. This prevents the scheduler from triggering any new runs while it’s booting up.
3. Verify catchup is Properly Applied
Double-check that your catchup setting is actually being honored by the DAG. Even though you set it in default_args, explicitly defining it in the DAG ensures there’s no accidental override:
dag = DAG( 'start_data_collect', default_args=default_args, schedule_interval='@daily', # or your desired interval catchup=False # Explicitly set here to confirm )
Typos or incorrect default_args passing can sometimes cause catchup to fail silently, so this extra step eliminates that risk.
4. Block Missed Schedule Triggers
If your scheduler was down for a while before restarting, Airflow might trigger the latest missed run even with catchup=False (since catchup=False stops backfilling all past runs, not the most recent one). To prevent this, add a LatestOnlyOperator at the start of your DAG:
from airflow.operators.latest_only import LatestOnlyOperator latest_only = LatestOnlyOperator(task_id='latest_only', dag=dag) # Make your other tasks depend on this operator your_data_collect_task.set_upstream(latest_only)
This ensures only the latest scheduled run executes, and any older missed runs are skipped entirely.
5. Tweak Scheduler Configuration (Airflow 2.x+)
If you’re running Airflow 2.x, adjust these airflow.cfg settings to reduce mass task triggering on startup:
scheduler_max_threads: Lower this value if the scheduler spawning too many threads at once is causing the load spike.min_file_process_interval: Increase this to reduce how often the scheduler scans for DAG changes, which slows down rapid task triggering during startup.
Just remember to test these changes in a staging environment first—they affect overall scheduler performance.
Content of the question originates from Stack Exchange, question author Sheridan

