如何用Python将CSV列中ML模型输出的关系数据拆分到多列
问题
我有一个存储机器学习模型输出的CSV文件,理想状态下该文件应包含Source、Relation type、Target三列,但目前每行的输出内容都存储在单个单元格中。我不需要entities数据,仅需将relations中的内容拆分至对应的独立列中。
当前CSV中每行的原始数据示例:
{'entities': [{'title': 'WarnerMedia', 'wikild': 'Q191715', 'label': 'Organization'}, {'title': 'Time (magazine)', 'wikild': 'Q43297', 'label': 'Organization'}, {'title': 'AOL', 'wikild': 'Q27585', 'label': 'Organization'}, {'title': 'Google', 'wikild': 'Q95', 'label': 'Organization'}, {'title': 'Warner Bros.', 'wikild': 'Q126399', 'label': 'Organization'}, {'title': 'U.S. Securities and Exchange Commission', 'wikild': 'Q953944', 'label': 'Organization'}], 'relations': [{'source': 'Time (magazine)', 'target': 'WarnerMedia', 'type': 'owned by'}, {'source': 'WarnerMedia', 'target': 'Time (magazine)', 'type': 'subsidiary'}, {'source': 'WarnerMedia', 'target': 'Time (magazine)', 'type': 'owned by'}, {'source': 'WarnerMedia', 'target': 'U.S. Securities and Exchange Commission', 'type': 'subsidiary'}, {'source': 'U.S. Securities and Exchange Commission', 'target': 'WarnerMedia', 'type': 'subsidiary'}, {'source': 'WarnerMedia', 'target': 'AOL', 'type': 'subsidiary'}, {'source': 'AOL', 'target': 'WarnerMedia', 'type': 'subsidiary'}]} {'entities': [{'title': 'Europe', 'wikild': 'Q46', 'label': 'Location'}, {'title': 'London', 'wikild': 'Q84', 'label': 'Organization'}, {'title': 'Federal Reserve', 'wikild': 'Q53536', 'label': 'Organization'}, {'title': 'United States', 'wikild': 'Q30', 'label': 'Organization'}, {'title': 'Federal government of the United States', 'wikild': 'Q48525', 'label': 'Organization'}, {'title': 'Bank of America', 'wikild': 'Q487907', 'label': 'Organization'}, {'title': 'Group of Seven', 'wikild': 'Q1764511', 'label': 'Organization'}, {'title': 'United States dollar', 'wikild': 'Q4917', 'label': 'Organization'}, {'title': 'New York (state)', 'wikild': 'Q1384', 'label': 'Organization'}, {'title': 'Alan Greenspan', 'wikild': 'Q193635', 'label': 'Person'}, {'title': 'Euro', 'wikild': 'Q4916', 'label': 'Organization'}, {'title': 'Germany', 'wikild': 'Q183', 'label': 'Organization'}], 'relations': [{'source': 'Federal Reserve', 'target': 'London', 'type': 'headquarters location'}, {'source': 'Bank of America', 'target': 'New York (state)', 'type': 'headquarters location'}, {'source': 'London', 'target': 'Federal Reserve', 'type': 'headquarters location'}, {'source': 'New York (state)', 'target': 'Bank of America', 'type': 'headquarters location'}]}
期望输出为包含三列的结构化表格:
| Source | Relation type | Target |
|---|---|---|
| Time (magazine) | owned by | WarnerMedia |
| WarnerMedia | subsidiary | Time (magazine) |
| WarnerMedia | owned by | Time (magazine) |
| WarnerMedia | subsidiary | U.S. Securities and Exchange Commission |
| U.S. Securities and Exchange Commission | subsidiary | WarnerMedia |
| WarnerMedia | subsidiary | AOL |
| AOL | subsidiary | WarnerMedia |
| Federal Reserve | headquarters location | London |
| Bank of America | headquarters location | New York (state) |
| London | headquarters location | Federal Reserve |
| New York (state) | headquarters location | Bank of America |
解决方案
可以用Python的pandas和ast模块实现拆分,步骤如下:
完整代码
import pandas as pd import ast # 读取原始CSV文件,替换为你的文件路径 df = pd.read_csv('input.csv', header=None, names=['raw_data']) # 存储最终结果的空列表 result_rows = [] # 遍历每行原始数据 for raw_str in df['raw_data']: # 将字符串格式的字典解析为Python字典 data_dict = ast.literal_eval(raw_str) # 提取relations列表,无数据则返回空列表 relations = data_dict.get('relations', []) # 拆分每个关系条目到目标列 for rel in relations: result_rows.append({ 'Source': rel['source'], 'Relation type': rel['type'], 'Target': rel['target'] }) # 转换为DataFrame并输出到新CSV result_df = pd.DataFrame(result_rows) result_df.to_csv('output.csv', index=False) # 打印前5行预览 print(result_df.head())
代码说明
ast.literal_eval():安全解析字符串格式的字典,避免直接使用eval()的安全风险- 完全忽略
entities字段,只提取并处理relations中的数据 - 每个关系条目会被拆分为单独一行,对应到Source、Relation type、Target三列
- 最终生成的
output.csv即为符合要求的结构化文件
内容的提问来源于stack exchange,提问作者S.Sushmi
相关产品推荐
相关产品推荐

