运行时设置Airflow环境变量:与预设置的效果是否一致?
No, the effects are not the same—it all depends on when Airflow loads its configuration versus when you set the environment variables.
Let's break this down based on how Airflow handles configuration and DAG loading:
1. Airflow Core Configuration Loads First
When you run an Airflow binary (like airflow scheduler, airflow webserver, or even airflow dags list), the first thing Airflow does is initialize its core configuration system. This process:
- Pulls values from multiple sources (in order of priority: command-line args > environment variables >
airflow.cfgfile > defaults) - Caches these values in memory for the lifetime of the running process
- Happens before any DAGs are loaded into memory
So if you set an Airflow-specific environment variable (like AIRFLOW__CORE__SQL_ALCHEMY_CONN or AIRFLOW__EXECUTOR__WORKER_CONCURRENCY) after starting the binary but during DAG loading, Airflow will never pick up this new value. The core configuration was already finalized before DAGs started loading.
2. Exception: DAG Code That Directly Reads Env Vars
The only exception here is if your DAG code explicitly reads environment variables using Python's os.getenv() or similar methods. For example:
import os from airflow import DAG dag = DAG( "example_dag", default_args={"owner": os.getenv("DAG_OWNER", "default_owner")}, schedule_interval="@daily" )
In this case, if you set os.environ["DAG_OWNER"] = "my_team" during DAG loading (e.g., in a top-level DAG file), this value will be picked up by the DAG code. But this has nothing to do with Airflow's core configuration system—it's just regular Python code reading the current process's environment.
3. Example Comparison
Let's use a concrete example to highlight the difference:
Scenario 1 (Set before running binary):
export AIRFLOW__CORE__EXECUTOR=CeleryExecutor airflow schedulerAirflow will use the Celery Executor, since the environment variable was present when the configuration was initialized.
Scenario 2 (Set during DAG loading):
Start the scheduler first:airflow schedulerThen, in a DAG file, add:
import os os.environ["AIRFLOW__CORE__EXECUTOR"] = "CeleryExecutor"Airflow's scheduler will continue using the Executor it was started with (e.g., SequentialExecutor by default), because the core configuration was already loaded. The DAG code will see the new value if it reads it directly, but Airflow's core behavior won't change.
- For Airflow core configuration variables (prefixed with
AIRFLOW__), you must set them before running the Airflow binary—changes made after startup (even during DAG loading) won't affect Airflow's behavior. - For custom environment variables used directly in DAG code, setting them during DAG loading will work, but this is independent of Airflow's built-in configuration system.
内容的提问来源于stack exchange,提问作者Peter Berg

