如何用Shell脚本解析含列表的YAML文件并实现循环调用
解析YAML列表到Shell变量并实现循环调用
原始YAML内容
configuration: account: account1 warehouse: warehouse1 database: database1 object_type: schema: schema1 functions: funtion1 tables: - table: table1 sql_file_loc: some_path/some_file.sql - table: table2 sql_file_loc: some_path/some_file.sql
需求
- 将
account、warehouse、database的值存入Shell变量供后续使用 - 遍历
tables列表,获取每个表名(table1/table2)和对应的sql_file_loc,实现循环处理
现有问题
使用自定义的parse_yaml函数解析后,存在两个问题:
- 未生成
tables列表中table字段对应的变量(如configuration_object_type_tables__table="table1") - 列表项的变量名出现双下划线(
__),与其他层级的单下划线格式不一致
解决方案
方案1:改进纯Shell解析脚本
修改parse_yaml函数,支持YAML列表项的解析,同时修复变量名格式问题:
function parse_yaml { local prefix="${2:-}" local s='[[:space:]]*' local w='[a-zA-Z0-9_]*' local fs=$(echo -e "\034") # 用不可见字符作为字段分隔符 sed -ne "s|^\($s\)-$s|\1-|p" \ -e "s|^\($s\)\(-\)\($s\)\($w\)$s:$s[\"']\(.*\)[\"']$s\$|\1$fs\4$fs\5|p" \ -e "s|^\($s\)\(-\)\($s\)\($w\)$s:$s\(.*\)$s\$|\1$fs\4$fs\5|p" \ -e "s|^\($s\)\($w\)$s:$s[\"']\(.*\)[\"']$s\$|\1$fs\2$fs\3|p" \ -e "s|^\($s\)\($w\)$s:$s\(.*\)$s\$|\1$fs\2$fs\3|p" "$1" | awk -F"$fs" -v prefix="$prefix" '{ indent = length($1)/2; # 处理列表项(以-开头的行) if ($1 ~ /-$/) { indent -= 1; list_key = vname[indent] "_item"; # 维护列表项计数 if (!list_count[list_key]) list_count[list_key] = 0; list_count[list_key]++; vname[indent+1] = list_key "_" list_count[list_key]; } else { vname[indent] = $2; } # 删除超出当前层级的键 for (i in vname) { if (i > indent + ($1 ~ /-$/ ? 1 : 0)) { delete vname[i]; } } if (length($3) > 0) { vn = ""; for (i=0; i < indent + ($1 ~ /-$/ ? 1 : 0); i++) { if (vn != "") vn = vn "_"; vn = vn vname[i]; } printf("%s%s_%s=\"%s\"\n", prefix, vn, $2, $3); } }' }
使用方法
- 解析YAML并加载变量:
parse_yaml config.yaml > parsed_vars.sh source parsed_vars.sh
生成的变量格式如下:
configuration_account="account1" configuration_warehouse="warehouse1" configuration_database="database1" configuration_object_type_schema="schema1" configuration_object_type_functions="funtion1" configuration_object_type_tables_item_1_table="table1" configuration_object_type_tables_item_1_sql_file_loc="some_path/some_file.sql" configuration_object_type_tables_item_2_table="table2" configuration_object_type_tables_item_2_sql_file_loc="some_path/some_file.sql"
- 循环处理表数据:
table_count=1 while [ -n "${configuration_object_type_tables_item_${table_count}_table}" ]; do current_table="${configuration_object_type_tables_item_${table_count}_table}" current_sql="${configuration_object_type_tables_item_${table_count}_sql_file_loc}" echo "处理表: $current_table,SQL文件路径: $current_sql" # 在这里添加你的业务逻辑(如执行SQL脚本) ((table_count++)) done
方案2:使用yq工具(更简洁可靠)
如果可以安装yq(YAML处理工具),推荐用这种方式,避免复杂的Shell正则解析:
安装yq
(根据系统选择安装方式,例如Ubuntu:sudo apt install yq;macOS:brew install yq)
提取变量并循环处理
# 提取基础配置变量 export ACCOUNT=$(yq '.configuration.account' config.yaml) export WAREHOUSE=$(yq '.configuration.warehouse' config.yaml) export DATABASE=$(yq '.configuration.database' config.yaml) # 遍历tables列表 yq -r '.configuration.object_type.tables[] | "\(.table)|\(.sql_file_loc)"' config.yaml | while IFS='|' read -r table sql_path; do echo "处理表: $table,SQL文件: $sql_path" # 执行你的业务逻辑,例如: # psql -d "$DATABASE" -U "$ACCOUNT" -f "$sql_path" done
内容的提问来源于stack exchange,提问作者Chuck
相关产品推荐
相关产品推荐

