如何用Python将百万行YAML文件转换为指定结构的CSV文件
将百万行YAML转换为指定格式CSV的解决方案
问题分析
需要将包含嵌套结构的百万行YAML文件转换为指定列的CSV,核心是把多层嵌套的categories→properties→values结构扁平化,同时保留所有要求的字段。由于CSV不允许重复列名,我们会对原需求中的重复字段(如id、ru)重命名以保证文件有效性。
最终CSV列名(修正重复后)
categories, category_id, depth, active, name, name_ru, is_merged, sort, properties, property_id, filter, top, filter_sort, display_type, values, value_key, value_xmlid, value_rank, value, value_ru
Python实现代码
import yaml import csv # 定义CSV字段(修正重复名称,完全对应原需求的所有字段) CSV_FIELDS = [ "categories", "category_id", "depth", "active", "name", "name_ru", "is_merged", "sort", "properties", "property_id", "filter", "top", "filter_sort", "display_type", "values", "value_key", "value_xmlid", "value_rank", "value", "value_ru" ] def yaml_to_csv(yaml_path, csv_path): with open(yaml_path, "r", encoding="utf-8") as yaml_file, \ open(csv_path, "w", newline="", encoding="utf-8") as csv_file: writer = csv.DictWriter(csv_file, fieldnames=CSV_FIELDS) writer.writeheader() # 加载YAML数据(超大文件可改用流式解析,此处为标准场景实现) yaml_data = yaml.safe_load(yaml_file) categories = yaml_data.get("categories", []) for category in categories: # 填充category层级的固定标识和基础数据 category_row = { "categories": "categories", "category_id": category.get("id"), "depth": category.get("depth"), "active": category.get("active"), "name": "name", "name_ru": category.get("name", {}).get("ru"), "is_merged": category.get("is_merged"), "sort": category.get("sort"), "properties": "properties" } properties = category.get("properties", []) if not properties: # 无properties时填充空值 empty_prop = {k: "" for k in CSV_FIELDS[9:]} writer.writerow({**category_row, **empty_prop}) continue for prop in properties: # 填充property层级数据 prop_row = { "property_id": prop.get("id"), "filter": prop.get("filter", ""), "top": prop.get("top", ""), "filter_sort": prop.get("filter_sort", ""), "display_type": prop.get("display_type", ""), "values": "values" if prop.get("values") else "" } values = prop.get("values", []) if not values: # 无values时填充空值 empty_val = {k: "" for k in CSV_FIELDS[14:]} writer.writerow({**category_row, **prop_row, **empty_val}) continue for val in values: # 填充value层级数据并写入行 val_row = { "value_key": val.get("key"), "value_xmlid": val.get("xmlid"), "value_rank": val.get("rank"), "value": "value", "value_ru": val.get("value", {}).get("ru") } writer.writerow({**category_row, **prop_row, **val_row}) # 执行转换(替换为你的输入输出文件路径) yaml_to_csv("input.yaml", "output.csv")
关键说明
- 重复字段处理:把原需求中重复的
id、ru分别重命名为category_id/property_id、name_ru/value_ru,避免CSV列名冲突。 - 空值填充:对缺失的字段(如部分property没有
top、部分没有values)用空字符串填充,符合CSV规范。 - 大文件适配:如果YAML文件超出内存承载上限,可改用
yaml.parse实现流式解析,避免一次性加载全部数据。 - 编码设置:全程使用UTF-8编码,确保俄文等非ASCII字符正常显示。
依赖安装
执行前需安装PyYAML库:
pip install pyyaml
内容的提问来源于stack exchange,提问作者Gerzzog
相关产品推荐
相关产品推荐

