如何基于DataFrame列中的单词拆分并展开数据?
解决DataFrame拆分标题并关联对应Score和num_comments的问题
原始DataFrame
Score num_comments titles 0 134 518 Uhaul implement nicotine-free hiring policy 1 28 43 Orangutan saves child from a giant volcano 2 30 114 Swimmer dies in a horrific shark attack in harbour 3 745 298 More teenagers than ever are addicted to glue 4 40 67 Lebanese lawyers union accuse Al Capone of fraud ... 9366 345 32 City of Louisville closed off this summer 9367 1200 234 New york rats "stronger than ever", reports say 9368 432 123 Congolese militia shipwrecked in Norway 9369 594 203 Scientists now agree on how to use ice in drinks 9370 611 153 Historic drought hits Atlantis
期望输出
需要生成每个单词对应原行Score和num_comments的新DataFrame:
Word score num_comments Uhaul 134 518 implement 134 518 nicotine-free 134 518 hiring 134 518 policy 134 518 Orangutan 28 43 saves 28 43 child 28 43 from 28 43 a 28 43 giant 28 43 volcano 28 43 ...
尝试过的错误方法及报错
- 错误使用
explode:
df3.explode(df3.assign(titles_split=df3.titles_split.str.split(',')), 'titles_split')
报错:
ValueError: column must be a scalar, tuple, or list thereof
- 尝试用
apply生成重复Score列:
df3['score_repeat'] = df3.apply(lambda x: [x.score] * len(x.titles_split) , axis =1)
报错:
TypeError: object of type 'float' has no len()
- 尝试用列表推导式生成重复列:
df3['score_repeat'] = [[y] * x for x, y in zip(df3['titles_split'].str.len(),df['score'])]
报错:
TypeError: can't multiply sequence by non-int of type 'float'
正确解决方案
不需要手动创建重复列,利用pandas的explode方法可以直接实现需求,步骤如下:
# 假设原始DataFrame名为df # 1. 将titles列按空格拆分为单词列表,生成Word列 df['Word'] = df['titles'].str.split() # 2. 展开Word列,自动复制对应行的Score和num_comments result_df = df.explode('Word')[['Word', 'Score', 'num_comments']] # 可选:重置索引 result_df = result_df.reset_index(drop=True)
错误原因说明
- 第一个错误:
explode的参数只需指定要展开的列名,不需要嵌套assign操作,语法使用错误。 - 第二个错误:
titles_split可能存在空值(float类型的NaN),导致len()无法调用;且完全不需要手动生成重复列,explode会自动处理其他列的复制。 - 第三个错误:
titles_split.str.len()可能返回float类型(存在空值时),无法用来乘以列表,同样这个思路冗余,explode是更高效的解决方案。
内容的提问来源于stack exchange,提问作者DavidReynisson
相关产品推荐
相关产品推荐

