如何在CentOS 7及以上系统搭建Airflow 2.0及更高版本高可用集群——Airflow 2.3.1基于RabbitMQ与PostgreSQL的HA部署方法咨询
Alright, let's walk through setting up a highly available Airflow 2.3.1 cluster on CentOS 7, using RabbitMQ as the message broker and PostgreSQL as the metadata database. I've deployed this exact setup a few times for production environments, so I'll break it down into actionable steps to make this as smooth as possible.
First, let's cover the basics you'll need:
- At least 3 CentOS 7 servers (1 for PostgreSQL, 1 for RabbitMQ, 2+ for Airflow nodes to enable high availability)
- Sudo privileges on all servers
- Network connectivity between all nodes (ensure ports 5432, 5672, 8080 are open, or adjust firewall rules accordingly)
- Disable SELinux temporarily (or configure it properly for Airflow):
setenforce 0
Step 1: Set Up PostgreSQL (Metadata Database)
Airflow relies on a robust metadata database, and PostgreSQL is the recommended choice. For true high availability, you can later set up PostgreSQL streaming replication, but we'll start with a single instance for simplicity.
Install PostgreSQL and dependencies:
yum install -y epel-release yum install -y postgresql-server postgresql-contribInitialize and start the service:
postgresql-setup initdb systemctl start postgresql systemctl enable postgresqlCreate an Airflow database and user:
Switch to the postgres user first:su - postgresThen run these commands:
createdb airflow createuser -P airflow # You'll be prompted to set a secure password GRANT ALL PRIVILEGES ON DATABASE airflow TO airflow; exitAllow remote access to PostgreSQL:
Editvar/lib/pgsql/data/pg_hba.confand add this line to allow connections from your Airflow nodes (replace192.168.0.0/24with your cluster subnet):host airflow airflow 192.168.0.0/24 scram-sha-256Then edit
var/lib/pgsql/data/postgresql.confto set:listen_addresses = '*'Restart PostgreSQL to apply changes:
systemctl restart postgresql
Step 2: Install and Configure RabbitMQ (Message Broker)
RabbitMQ will handle task queuing between Airflow schedulers and workers.
Install Erlang (required for RabbitMQ):
yum install -y erlangInstall RabbitMQ:
rpm --import https://github.com/rabbitmq/signing-keys/releases/download/2.0/rabbitmq-release-signing-key.asc curl -s https://packagecloud.io/install/repositories/rabbitmq/rabbitmq-server/script.rpm.sh | sudo bash yum install -y rabbitmq-serverStart and enable the service:
systemctl start rabbitmq-server systemctl enable rabbitmq-serverCreate an Airflow user and set permissions:
rabbitmqctl add_user airflow your_secure_rabbit_password rabbitmqctl set_permissions -p / airflow ".*" ".*" ".*" rabbitmqctl set_user_tags airflow administrator # Optional, for management UI access(Optional) Enable the RabbitMQ management UI for monitoring:
rabbitmq-plugins enable rabbitmq_management
Step 3: Install Airflow 2.3.1 on All Airflow Nodes
Repeat these steps on every node that will run Airflow webserver, scheduler, worker, or triggerer.
Install Python 3 and dependencies:
yum install -y python3 python3-pip gcc python3-devel openldap-devel(Optional) Set up a faster PyPI mirror to speed up installs:
pip3 config set global.index-url https://pypi.tuna.tsinghua.edu.cn/simpleInstall Airflow 2.3.1 with required providers:
pip3 install apache-airflow==2.3.1 apache-airflow-providers-postgres apache-airflow-providers-rabbitmqConfigure Airflow home and database connection:
Set the Airflow home directory (we'll use/opt/airflow):export AIRFLOW_HOME=/opt/airflow echo "export AIRFLOW_HOME=/opt/airflow" >> ~/.bashrcInitialize the Airflow database:
airflow db initUpdate Airflow config (
$AIRFLOW_HOME/airflow.cfg):
Edit these key settings:# Metadata database connection sql_alchemy_conn = postgresql+psycopg2://airflow:your_postgres_password@postgres_server_ip:5432/airflow # Message broker (RabbitMQ) broker_url = amqp://airflow:your_rabbit_password@rabbitmq_server_ip:5672// result_backend = db+postgresql://airflow:your_postgres_password@postgres_server_ip:5432/airflow # Use CeleryExecutor for distributed tasks executor = CeleryExecutor # Enable scheduler HA scheduler_health_check_threshold = 30Create an Airflow admin user:
airflow users create --username admin --firstname Admin --lastname Ops --role Admin --email admin@yourdomain.com
Step 4: Configure High Availability Components
Webserver HA
Deploy the Airflow webserver on multiple nodes, then use a load balancer (like Nginx) to route traffic.
Start the webserver on each node (run in background with
-D):airflow webserver -D --port 8080Set up Nginx as a load balancer (on a separate node or one of the web nodes):
Install Nginx:yum install -y nginxCreate a config file
/etc/nginx/conf.d/airflow.conf:upstream airflow_web { server web_node_1_ip:8080; server web_node_2_ip:8080; # Add more nodes as needed } server { listen 80; server_name airflow.yourdomain.com; location / { proxy_pass http://airflow_web; proxy_set_header Host $host; proxy_set_header X-Real-IP $remote_addr; proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; } }Start and enable Nginx:
systemctl start nginx systemctl enable nginx
Scheduler HA
Airflow 2.x natively supports multiple schedulers. Just start the scheduler on 2+ nodes:
airflow scheduler -D
Airflow will automatically handle leader election and task distribution between schedulers.
Worker Nodes
Start workers on as many nodes as you need for task capacity:
airflow celery worker -D
(Optional) Assign workers to specific queues:
airflow celery worker -D -q default,data_processing
Triggerer (Optional but Recommended)
For async tasks (like sensor operators), run the triggerer on multiple nodes:
airflow triggerer -D
Step 5: Verify the Cluster
- Access the Airflow UI via your Nginx load balancer URL, log in with the admin user.
- Navigate to Admin > Cluster Activity to confirm multiple schedulers, workers, and triggerers are online.
- Run a test DAG to validate task execution:
airflow dags test example_bash_operator 2024-01-01 - Test failover: Stop one scheduler/webserver/worker node and confirm the cluster continues operating normally.
Production Best Practices
Use systemd services to manage Airflow components (instead of
-Dflag). For example, create/etc/systemd/system/airflow-scheduler.service:[Unit] Description=Airflow Scheduler After=network.target postgresql.service rabbitmq-server.service [Service] User=root Environment=AIRFLOW_HOME=/opt/airflow ExecStart=/usr/local/bin/airflow scheduler Restart=always [Install] WantedBy=multi-user.targetThen enable it:
systemctl daemon-reload && systemctl enable airflow-scheduler && systemctl start airflow-schedulerEnable remote logging (e.g., to NFS or S3) to avoid losing logs when nodes go down.
Regularly back up your PostgreSQL database with
pg_dump.Set up monitoring (Prometheus + Grafana) to track Airflow metrics, RabbitMQ queue lengths, and PostgreSQL health.
内容的提问来源于stack exchange,提问作者Tushar

