You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用不同长度字典输出生成Pandas DataFrame并对齐同键列

将元组列表的列表转换为对齐列的Pandas DataFrame(缺失值补0)

你可以通过把每个元组列表转换成字典,再直接传入Pandas的DataFrame构造器来实现需求,Pandas会自动识别所有唯一键作为列,缺失的键对应值填充NaN,最后把NaN替换为0并转为整数类型即可。

修改后的完整代码

from collections import Counter
import pandas as pd

# 从原始数据集处理的代码(替换你原有的循环部分)
liwc = [] 
for item in data.tokens:
    counts = Counter(category for token in item for category in parse(token))
    liwc.append(dict(counts))  # 将统计结果直接转为字典存入列表

# 生成DataFrame并处理缺失值
df = pd.DataFrame(liwc).fillna(0).astype(int)
print(df)

示例测试(用你给出的样本数据)

如果用你提供的4行样本数据模拟:

# 模拟你的liwc列表
liwc = [
    dict([('social', 8), ('drives', 3), ('prep', 12), ('function', 5)]),
    dict([('social', 10), ('conj', 8), ('ppron', 7), ('i', 7)]),
    dict([('pronoun', 12), ('ppron', 10), ('i', 7), ('drives', 8)]),
    dict([('affiliation', 6), ('function', 55), ('pronoun', 19), ('ppron', 16)])
]

df = pd.DataFrame(liwc).fillna(0).astype(int)
print(df)

输出结果

social  drives  prep  function  conj  ppron  i  pronoun  affiliation
0       8       3    12         5     0      0  0        0            0
1      10       0     0         0     8      7  7        0            0
2       0       8     0         0     0     10  7       12            0
3       0       0     0        55     0     16  0       19            6

关键步骤说明

  • 转字典:把Counter生成的元组列表转为字典,让Pandas能直接识别键值对关系。
  • 自动列对齐:Pandas会自动收集所有字典中的键作为列名,每个字典对应一行数据,未出现的键自动填充NaN。
  • 缺失值处理:用fillna(0)把空值替换为0,astype(int)确保所有值为整数类型(避免缺失值填充后变成浮点数)。

内容的提问来源于stack exchange,提问作者Todd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 11:12:28