You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pandas groupby实现CSV文件到指定格式XML的转换

可行实现方案

我们可以配合pandas做数据分组+Python标准库xml.etree.ElementTree生成XML节点完成转换,全程不需要额外安装复杂的XML处理库。

前置依赖

先确保已安装pandas:

pip install pandas

完整可运行代码

import pandas as pd
import xml.etree.ElementTree as ET
from xml.dom import minidom

# 1. 读取CSV文件,替换为你本地的csv文件路径
df = pd.read_csv("coupon.csv")

# 2. 按Amount分组,把同组的Code存入列表
# 分组后得到的结构为:key是Amount值(如CODE50),value是对应Code的列表
grouped = df.groupby("Amount")["Code"].apply(list).to_dict()

# 3. 生成XML结构
# 先创建临时根节点(符合XML规范要求,后续会去除)
root = ET.Element("tmp-root")

for coupon_id, code_list in grouped.items():
    # 创建coupon-codes节点,设置coupon-id属性
    coupon_node = ET.SubElement(root, "coupon-codes", attrib={"coupon-id": coupon_id})
    # 遍历code列表创建子节点
    for code in code_list:
        code_node = ET.SubElement(coupon_node, "code")
        code_node.text = str(code)

# 4. 格式化XML输出,调整缩进匹配要求
rough_string = ET.tostring(root, "utf-8")
reparsed = minidom.parseString(rough_string)
pretty_xml = reparsed.toprettyxml(indent="    ", encoding="utf-8").decode("utf-8")

# 去掉XML声明、临时根节点标签,得到你需要的最终格式
xml_lines = [
    line for line in pretty_xml.splitlines() 
    if line.strip() not in ["<?xml version=\"1.0\" encoding=\"utf-8\"?>", "<tmp-root>", "</tmp-root>"]
]
final_output = "\n".join(xml_lines)

# 5. 写入文件
with open("output.xml", "w", encoding="utf-8") as f:
    f.write(final_output)

关键逻辑说明

  • 分组处理:groupby("Amount")["Code"].apply(list)会把相同Amount的所有Code聚合为一个列表,转成字典后可以直接遍历键值对,不需要处理复杂的groupby原生对象结构
  • XML生成:ET.SubElement直接创建子节点,attrib参数快速设置节点属性,text属性填充节点内容
  • 格式适配:用minidom的toprettyxml指定缩进为4个空格,和你要求的输出格式完全对齐

内容的提问来源于stack exchange,提问作者Jeremy Altman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 05:36:01