You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从含不同长度列表的字典生成汇总DataFrame表格?

问题描述

给定如下字典:

response = {
    'A': ['CATEGORY 2'],
    'B': ['CATEGORY 1', 'CATEGORY 2'],
    'C': [],
    'D': ['CATEGORY 3'],
}

希望生成如下格式的DataFrame:

|  ITEM  |  CATEGORY 1  |  CATEGORY 2  |  CATEGORY 3  |
|   A    |              |      x       |              |
|   B    |      x       |      x       |              |
|   C    |              |              |              |
|   D    |              |              |      x       |

尝试了以下代码,但结果不符合预期:

df = pd.DataFrame.from_dict(response, orient='index').fillna('x')

df = df.reset_index()

df = df.rename(columns={'index': 'ITEM'})

print(df)

输出结果:

ITEM           0           1
0    A  CATEGORY 2           x
1    B  CATEGORY 1  CATEGORY 2
2    C           x           x
3    D  CATEGORY 3           x
解决方案

方法一:使用MultiLabelBinarizer(推荐)

借助sklearn的MultiLabelBinarizer快速将多标签列表转换为二进制矩阵,再替换标记值并整理格式:

import pandas as pd
from sklearn.preprocessing import MultiLabelBinarizer

response = {
    'A': ['CATEGORY 2'],
    'B': ['CATEGORY 1', 'CATEGORY 2'],
    'C': [],
    'D': ['CATEGORY 3'],
}

# 初始化多标签二值化器
mlb = MultiLabelBinarizer()
# 将字典值转换为二进制矩阵
binary_matrix = mlb.fit_transform(response.values())
# 以类别为列、字典键为索引创建DataFrame
df = pd.DataFrame(binary_matrix, index=response.keys(), columns=mlb.classes_)
# 替换标记:1改为'x',0改为空字符串
df = df.replace({1: 'x', 0: ''})
# 重置索引并命名为ITEM
df = df.reset_index().rename(columns={'index': 'ITEM'})

# 输出为markdown格式表格(可直接打印或导出)
print(df.to_markdown(index=False))

输出结果:

| ITEM   | CATEGORY 1   | CATEGORY 2   | CATEGORY 3   |
|:-------|:-------------|:-------------|:-------------|
| A      |              | x            |              |
| B      | x            | x            |              |
| C      |              |              |              |
| D      |              |              | x            |

方法二:手动构建列(无需额外依赖)

如果不想引入sklearn,可以手动遍历所有类别,逐个判断每个条目是否包含该类别:

import pandas as pd

response = {
    'A': ['CATEGORY 2'],
    'B': ['CATEGORY 1', 'CATEGORY 2'],
    'C': [],
    'D': ['CATEGORY 3'],
}

# 提取所有唯一类别并排序
all_categories = set()
for cats in response.values():
    all_categories.update(cats)
all_categories = sorted(all_categories)

# 构建数据字典
data = {'ITEM': list(response.keys())}
for cat in all_categories:
    data[cat] = ['x' if cat in cats else '' for cats in response.values()]

# 创建目标DataFrame
df = pd.DataFrame(data)

print(df.to_markdown(index=False))

输出结果与方法一完全一致。

原代码问题说明

你之前的代码逻辑是把每个列表的元素按顺序放到0、1这类序列列中,而不是把类别名称作为列来标记是否存在。from_dict(orient='index')的行为是将每个列表元素映射到不同的列,这和你需要的“类别作为列、标记存在性”的需求不匹配,因此需要先完成多标签到类别列的转换,再构建DataFrame。

内容的提问来源于stack exchange,提问作者VERBOSE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 16:03:32