如何将DataFrame中含元组列表的列展开并保留分类列?
解决方法
你的代码问题在于,当把仅2行的category列和6行的元组数据按列拼接时,Pandas会按索引对齐,导致后4行的category值缺失(变为NaN),无法正确关联每个词汇对应的分类。
正确思路是先将包含元组的列表展开,让每个元组成为单独一行,同时保留对应分类,再拆分元组为两列。
步骤1:展开列表列
使用explode()方法将column_with_tuples中的每个元组拆分为独立行,此时category会自动重复对应次数:
df_exploded = df.explode('column_with_tuples')
执行后df_exploded的结构:
| column_with_tuples | category | |
|---|---|---|
| 0 | ('word1', 10) | category1 |
| 0 | ('word2', 20) | category1 |
| 0 | ('word3', 30) | category1 |
| 1 | ('word4', 40) | category2 |
| 1 | ('word5', 50) | category2 |
| 1 | ('word6', 60) | category2 |
步骤2:拆分元组为独立列
通过str访问器提取元组的第一个和第二个元素,分别作为word和frequency列:
df_new = df_exploded.assign( word=lambda x: x['column_with_tuples'].str[0], frequency=lambda x: x['column_with_tuples'].str[1] ).drop('column_with_tuples', axis=1)
完整代码
import pandas as pd df = pd.DataFrame({'column_with_tuples': [[('word1', 10), ('word2', 20), ('word3', 30)], [('word4', 40), ('word5', 50), ('word6', 60)]], 'category':['category1','category2']}) # 展开列表并拆分元组 df_new = df.explode('column_with_tuples').assign( word=lambda x: x['column_with_tuples'].str[0], frequency=lambda x: x['column_with_tuples'].str[1] ).drop('column_with_tuples', axis=1).reset_index(drop=True) print(df_new)
执行后输出完全符合预期:
word frequency category 0 word1 10 category1 1 word2 20 category1 2 word3 30 category1 3 word4 40 category2 4 word5 50 category2 5 word6 60 category2
额外说明
explode()是Pandas 0.25+版本支持的方法,若使用旧版本,可通过列表推导式实现类似效果:pd.DataFrame([(cat, tup) for cat, tups in zip(df['category'], df['column_with_tuples']) for tup in tups], columns=['category', 'tuple_col'])。- 拆分元组也可使用
pd.Series方法:df_exploded['column_with_tuples'].apply(pd.Series).rename(columns={0:'word', 1:'frequency'}),但str访问器效率更高。
内容的提问来源于stack exchange,提问作者Yana
相关产品推荐
相关产品推荐

