如何用Pandas拆分含字典列表的列并移除version键?
解决Pandas拆分含字典列表的列并移除指定键的问题
直接上可运行的解决方案,分两步处理:先清理每个字典里的version键,再把列表拆分成多列。
步骤1:构造样例数据(方便你测试)
import pandas as pd # 模拟你的原始DataFrame data = { 'URL': ['https://example.com', 'https://test.org'], 'report': [ [{'detected': True, 'name': 'Python', 'version': '3.9'}, {'detected': False, 'name': 'Java'}], [{'detected': True, 'name': 'Node.js', 'version': '18'}] ] } df = pd.DataFrame(data)
步骤2:移除字典中的version键
遍历report列的每个列表,对里面的字典做键过滤:
# 清理每个字典,去掉version键 df['report'] = df['report'].apply( lambda lst: [{k: v for k, v in d.items() if k != 'version'} for d in lst] )
步骤3:拆分列表为多列
把处理后的列表拆成report_1、report_2这样的列,自动适配每行的列表长度:
# 将列表拆分为多列 report_cols = df['report'].apply(pd.Series) # 重命名列名 report_cols.columns = [f'report_{i+1}' for i in report_cols.columns] # 合并原DataFrame和拆分后的列,去掉原report列 final_df = pd.concat([df.drop('report', axis=1), report_cols], axis=1)
最终效果
运行后final_df的结构如下:
| URL | report_1 | report_2 |
|---|---|---|
| https://example.com | {'detected': True, 'name': 'Python'} | {'detected': False, 'name': 'Java'} |
| https://test.org | {'detected': True, 'name': 'Node.js'} | NaN |
如果某行的report列表长度更长,会自动生成report_3、report_4等列,长度不足的行对应位置填充NaN,完全适配你的需求。
内容的提问来源于stack exchange,提问作者morfic001
相关产品推荐
相关产品推荐

