Airflow中如何固定DAG标签在UI中的显示顺序?
Airflow DAG标签自定义显示顺序问题
问题描述
我有两个DAG:
- example1:标签配置为
['toy', 'umbrella'],Airflow UI中显示顺序为toyumbrella,与配置一致 - example2:标签配置为
['toy', 'ball'],但UI中显示顺序为balltoy,与配置顺序不符
需要让example2的标签在UI中按toy ball的顺序显示。
原因分析
Airflow UI默认会对DAG标签按字母顺序排序展示。example1中toy(首字母t)的字母顺序晚于umbrella(首字母u),所以显示顺序和配置一致;而example2中ball(首字母b)早于toy(首字母t),因此UI自动调整了顺序。
解决方案
Airflow 2.5及以上版本支持通过Tag对象和tag_order参数强制指定标签显示顺序,具体步骤如下:
1. 导入Tag类
在代码顶部添加导入语句:
from airflow.models import Tag
2. 显式创建标签对象并指定顺序
修改DAG定义中的tags参数,用Tag对象列表替代字符串列表,并通过tag_order参数指定显示顺序:
with DAG( dag_id=DAG_ID, start_date=datetime(2023, 2, 12), max_active_runs=1, schedule=None, default_args=default_dag_args, params={ "tables_to_ingest": Param( [], type="array", ), }, catchup=False, # 显式创建标签对象,并指定显示顺序 tags=[Tag("toy"), Tag("ball")], tag_order=["toy", "ball"] # 指定UI中标签的展示顺序 ) as trigger_dag: # 后续任务代码保持不变 ...
低版本兼容方案(Airflow < 2.5)
如果无法升级Airflow版本,可以通过给标签名称添加排序前缀的方式实现(比如0_toy、1_ball),然后在UI中通过自定义CSS隐藏前缀,但这种方法属于临时hack,不推荐长期使用。
修改后的完整代码示例
import os import logging from datetime import datetime import yaml from airflow import DAG from airflow.decorators import task from airflow.models import Tag # 新增导入Tag类 from airflow.models.param import Param from lib.utils.bigquery import load_json_into_table, full_table_name from lib.utils.cloud_storage import load_config, read_files_in_bucket from lib.utils.slack import create_notification_operators # Default parameters for the task ENV = os.environ.get("ENV") DAG_ID = f"toy_data__ingestion__{ENV}" PROJECT_ID = f"cat-{ENV}-toy" # Defining the bucket to read the configuration files CONFIG_BUCKET = f"cat-{ENV}-toy/products/ref_data" CONFIG_FILENAME = "toy_ref_data.yaml" # Load Config CONFIG = load_config(bucket_path=CONFIG_BUCKET, filename=CONFIG_FILENAME) logging.info(f"##### The overall config loaded for {DAG_ID} has been loaded") REGION = CONFIG["region"] DATASET = CONFIG["dataset_id"] default_dag_args = { "owner": "fancybear", "retries": 1, "project_id": PROJECT_ID, "region": REGION, } with DAG( dag_id=DAG_ID, start_date=datetime(2023, 2, 12), max_active_runs=1, schedule=None, default_args=default_dag_args, params={ "tables_to_ingest": Param( [], type="array", ), }, catchup=False, tags=[Tag("toy"), Tag("ball")], tag_order=["toy", "ball"] # 指定标签显示顺序 ) as trigger_dag: @task def load_data(**context): file_names = [ f"{file_name}.yaml" for file_name in context["params"]["tables_to_ingest"] ] TABLE_CONFIGS = [ yaml.safe_load( read_files_in_bucket( bucket_path=(CONFIG_BUCKET + "/schemas"), matching_filename=file )[0] ) for file in file_names ] for config in TABLE_CONFIGS: logging.info( f"####### Loading data into table {config['table_id']} #######" ) config[ "raw_file_path" ] = f"gs://{CONFIG_BUCKET}/data/{config['raw_file_name']}" load_json_into_table( config=config, table_id=full_table_name(PROJECT_ID, DATASET, config["table_id"]), region=REGION, ) notify_operators = create_notification_operators(DAG_ID) load_data() >> notify_operators
内容的提问来源于stack exchange,提问作者fancybear
相关产品推荐
相关产品推荐

