如何利用不同长度字典输出生成Pandas DataFrame并对齐同键列
将元组列表的列表转换为对齐列的Pandas DataFrame(缺失值补0)
你可以通过把每个元组列表转换成字典,再直接传入Pandas的DataFrame构造器来实现需求,Pandas会自动识别所有唯一键作为列,缺失的键对应值填充NaN,最后把NaN替换为0并转为整数类型即可。
修改后的完整代码
from collections import Counter import pandas as pd # 从原始数据集处理的代码(替换你原有的循环部分) liwc = [] for item in data.tokens: counts = Counter(category for token in item for category in parse(token)) liwc.append(dict(counts)) # 将统计结果直接转为字典存入列表 # 生成DataFrame并处理缺失值 df = pd.DataFrame(liwc).fillna(0).astype(int) print(df)
示例测试(用你给出的样本数据)
如果用你提供的4行样本数据模拟:
# 模拟你的liwc列表 liwc = [ dict([('social', 8), ('drives', 3), ('prep', 12), ('function', 5)]), dict([('social', 10), ('conj', 8), ('ppron', 7), ('i', 7)]), dict([('pronoun', 12), ('ppron', 10), ('i', 7), ('drives', 8)]), dict([('affiliation', 6), ('function', 55), ('pronoun', 19), ('ppron', 16)]) ] df = pd.DataFrame(liwc).fillna(0).astype(int) print(df)
输出结果
social drives prep function conj ppron i pronoun affiliation 0 8 3 12 5 0 0 0 0 0 1 10 0 0 0 8 7 7 0 0 2 0 8 0 0 0 10 7 12 0 3 0 0 0 55 0 16 0 19 6
关键步骤说明
- 转字典:把
Counter生成的元组列表转为字典,让Pandas能直接识别键值对关系。 - 自动列对齐:Pandas会自动收集所有字典中的键作为列名,每个字典对应一行数据,未出现的键自动填充
NaN。 - 缺失值处理:用
fillna(0)把空值替换为0,astype(int)确保所有值为整数类型(避免缺失值填充后变成浮点数)。
内容的提问来源于stack exchange,提问作者Todd
相关产品推荐
相关产品推荐

