You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python/Pandas将类透视表数据转换为JSON格式

解决CSV按profile分组并转换为指定JSON结构的问题

问题分析

你之前的代码用groupby.apply(lambda x: x.to_json(orient='records'))会把每个分组的整行数据转成JSON字符串,而不是提取id组成数组,这就是无效的原因。针对你的需求,直接用pandas的聚合方法更高效准确,一万行数据完全能轻松处理。

完整解决方案代码

import pandas as pd
import json

# 读取CSV文件(替换成你的文件路径)
df = pd.read_csv("your_input.csv")

# 按profile分组,将对应id聚合为数组
grouped_data = df.groupby("profile")["id"].agg(list).reset_index()

# 转换为{profile: [id1, id2,...]}的字典结构
result_dict = grouped_data.set_index("profile")["id"].to_dict()

# 生成格式化的JSON字符串(可选,方便查看或保存)
result_json = json.dumps(result_dict, indent=4, ensure_ascii=False)

# 输出结果或保存到文件
print(result_json)
# 保存到JSON文件
with open("output.json", "w", encoding="utf-8") as f:
    f.write(result_json)

关键说明

  • groupby("profile")["id"].agg(list):直接对每个profile分组下的id列进行聚合,生成数组,这一步是向量化操作,比lambda遍历高效得多,适合大数据量。
  • 如果你的CSV里还有其他需要保留的字段,可以调整聚合逻辑,比如agg({"id": list, "other_col": "first"}),但根据你的需求,只聚合id就足够。
  • json.dumps的indent=4是为了生成格式化的JSON,方便阅读;ensure_ascii=False处理中文等非ASCII字符。

内容的提问来源于stack exchange,提问作者Siddharth Karkhanis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 22:10:06