AirFlow 1.10.2并行任务延迟60秒执行问题求助
It’s frustrating when your DAGs show as running in the UI but take a full minute to actually start executing—especially since this worked smoothly in Airflow 1.9. Let’s break down the most likely culprits based on your setup and configuration, along with actionable fixes:
1. Scheduler Thread Count Bottleneck
Your max_threads is set to 2, which controls how many threads the scheduler uses to process DAGs (parse them, check for ready tasks, and assign work). With frequent concurrent DAG triggers and 7 total DAGs, this low thread count can easily overwhelm the scheduler—meaning it can’t keep up with checking for new dag runs and assigning tasks quickly enough.
Fix: Increase max_threads to a value that matches your server’s CPU capacity (try 4 or 8 to start). Update your airflow.cfg:
max_threads = 8
Restart the scheduler service after making this change.
2. Scheduler Polling Interval for New Dag Runs
Airflow 1.10.x introduced changes to the scheduler’s loop logic compared to 1.9. Even if the UI shows a dag run as "running", the scheduler might not pick up the associated task instances immediately if its polling interval is too long.
Check & Adjust: Verify the scheduler_heartbeat_sec parameter (default is 5 seconds, but if modified, it could be longer). Ensure it’s set to a low value to make the scheduler poll for new tasks more frequently:
scheduler_heartbeat_sec = 5
Your dag_dir_list_interval (300s) is fine for your use case since triggers are code-based, not file-change driven.
3. Metadata Database Bottlenecks
You’re using MySQL for the metadata store, which is a common source of delays if the scheduler is making slow or frequent queries. Bottlenecks here can prevent the scheduler from quickly reading/writing task state.
Fixes:
- Enable SQLAlchemy connection pooling to reduce DB connection overhead. Add these lines to
airflow.cfg:sql_alchemy_pool_size = 10 sql_alchemy_max_overflow = 20 - Check your MySQL server’s resource usage (CPU, memory, disk I/O). Slow queries or locked tables can cripple the scheduler’s ability to process task state changes.
4. DAG Parsing Overhead
Your trigger_cleansing.py DAG takes 0.87s to parse—nearly 70% of your total DagBag parsing time. With max_threads=2, frequent triggers of this DAG can tie up scheduler threads, leaving less capacity to handle new task assignments.
Optimization:
- Move heavy imports or initialization code outside the DAG definition in
trigger_cleansing.py—this ensures it only runs once during parsing, not every time the scheduler scans the DAG. - Explicitly set
schedule_interval=Nonefor your non-scheduled DAGs to avoid the scheduler wasting cycles checking for scheduled runs.
5. Scheduler Process Health
Your systemctl output shows many child processes under the scheduler service—these are likely LocalExecutor worker processes. If the main scheduler process is stuck or inefficient, it can cause delays in task scheduling.
Check:
- Restart the scheduler to rule out stuck processes:
systemctl restart airflow-scheduler - Monitor the scheduler logs (typically
/var/log/airflow/scheduler.log) for errors, warnings, or repeated messages about delayed task processing.
Additional Testing to Isolate the Issue
- Trigger a single DAG instance manually and track the time between the UI showing "running" and the Kubernetes pod starting. If the delay persists, it’s likely a scheduler/DB issue rather than concurrent load.
- Temporarily pause file monitoring to reduce load—if the delay disappears, this confirms the problem is tied to high concurrency and scheduler capacity.
Since this worked seamlessly in Airflow 1.9, the most impactful fixes will likely be adjusting the scheduler thread count, optimizing DB connections, or tweaking polling intervals. Start with increasing max_threads and monitoring logs to see if that reduces the 60s delay.
内容的提问来源于stack exchange,提问作者Shanit

