You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将百万行YAML文件转换为指定结构的CSV文件

将百万行YAML转换为指定格式CSV的解决方案

问题分析

需要将包含嵌套结构的百万行YAML文件转换为指定列的CSV,核心是把多层嵌套的categories→properties→values结构扁平化,同时保留所有要求的字段。由于CSV不允许重复列名,我们会对原需求中的重复字段(如id、ru)重命名以保证文件有效性。

最终CSV列名(修正重复后)

categories, category_id, depth, active, name, name_ru, is_merged, sort, properties, property_id, filter, top, filter_sort, display_type, values, value_key, value_xmlid, value_rank, value, value_ru

Python实现代码

import yaml
import csv

# 定义CSV字段(修正重复名称,完全对应原需求的所有字段)
CSV_FIELDS = [
    "categories", "category_id", "depth", "active", "name", "name_ru",
    "is_merged", "sort", "properties", "property_id", "filter", "top",
    "filter_sort", "display_type", "values", "value_key", "value_xmlid",
    "value_rank", "value", "value_ru"
]

def yaml_to_csv(yaml_path, csv_path):
    with open(yaml_path, "r", encoding="utf-8") as yaml_file, \
         open(csv_path, "w", newline="", encoding="utf-8") as csv_file:
        
        writer = csv.DictWriter(csv_file, fieldnames=CSV_FIELDS)
        writer.writeheader()
        
        # 加载YAML数据(超大文件可改用流式解析,此处为标准场景实现)
        yaml_data = yaml.safe_load(yaml_file)
        categories = yaml_data.get("categories", [])
        
        for category in categories:
            # 填充category层级的固定标识和基础数据
            category_row = {
                "categories": "categories",
                "category_id": category.get("id"),
                "depth": category.get("depth"),
                "active": category.get("active"),
                "name": "name",
                "name_ru": category.get("name", {}).get("ru"),
                "is_merged": category.get("is_merged"),
                "sort": category.get("sort"),
                "properties": "properties"
            }
            
            properties = category.get("properties", [])
            if not properties:
                # 无properties时填充空值
                empty_prop = {k: "" for k in CSV_FIELDS[9:]}
                writer.writerow({**category_row, **empty_prop})
                continue
            
            for prop in properties:
                # 填充property层级数据
                prop_row = {
                    "property_id": prop.get("id"),
                    "filter": prop.get("filter", ""),
                    "top": prop.get("top", ""),
                    "filter_sort": prop.get("filter_sort", ""),
                    "display_type": prop.get("display_type", ""),
                    "values": "values" if prop.get("values") else ""
                }
                
                values = prop.get("values", [])
                if not values:
                    # 无values时填充空值
                    empty_val = {k: "" for k in CSV_FIELDS[14:]}
                    writer.writerow({**category_row, **prop_row, **empty_val})
                    continue
                
                for val in values:
                    # 填充value层级数据并写入行
                    val_row = {
                        "value_key": val.get("key"),
                        "value_xmlid": val.get("xmlid"),
                        "value_rank": val.get("rank"),
                        "value": "value",
                        "value_ru": val.get("value", {}).get("ru")
                    }
                    writer.writerow({**category_row, **prop_row, **val_row})

# 执行转换(替换为你的输入输出文件路径)
yaml_to_csv("input.yaml", "output.csv")

关键说明

  1. 重复字段处理:把原需求中重复的id、ru分别重命名为category_id/property_id、name_ru/value_ru,避免CSV列名冲突。
  2. 空值填充:对缺失的字段(如部分property没有top、部分没有values)用空字符串填充,符合CSV规范。
  3. 大文件适配:如果YAML文件超出内存承载上限,可改用yaml.parse实现流式解析,避免一次性加载全部数据。
  4. 编码设置:全程使用UTF-8编码,确保俄文等非ASCII字符正常显示。

依赖安装

执行前需安装PyYAML库:

pip install pyyaml

内容的提问来源于stack exchange,提问作者Gerzzog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 15:27:20