You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于DataFrame列中的单词拆分并展开数据?

解决DataFrame拆分标题并关联对应Score和num_comments的问题

原始DataFrame

Score       num_comments     titles
0       134         518              Uhaul implement nicotine-free hiring policy
1       28          43               Orangutan saves child from a giant volcano
2       30          114              Swimmer dies in a horrific shark attack in harbour
3       745         298              More teenagers than ever are addicted to glue 
4       40          67               Lebanese lawyers union accuse Al Capone of fraud
...
9366    345         32               City of Louisville closed off this summer
9367    1200        234              New york rats "stronger than ever", reports say
9368    432         123              Congolese militia shipwrecked in Norway
9369    594         203              Scientists now agree on how to use ice in drinks
9370    611         153              Historic drought hits Atlantis

期望输出

需要生成每个单词对应原行Score和num_comments的新DataFrame:

Word           score    num_comments
Uhaul          134      518
implement      134      518
nicotine-free  134      518
hiring         134      518
policy         134      518
Orangutan      28       43
saves          28       43
child          28       43
from           28       43
a              28       43
giant          28       43
volcano        28       43
...

尝试过的错误方法及报错

  1. 错误使用explode:
df3.explode(df3.assign(titles_split=df3.titles_split.str.split(',')), 'titles_split')

报错:

ValueError: column must be a scalar, tuple, or list thereof
  1. 尝试用apply生成重复Score列:
df3['score_repeat'] = df3.apply(lambda x: [x.score] * len(x.titles_split) , axis =1)

报错:

TypeError: object of type 'float' has no len()
  1. 尝试用列表推导式生成重复列:
df3['score_repeat'] = [[y] * x for x, y in zip(df3['titles_split'].str.len(),df['score'])]

报错:

TypeError: can't multiply sequence by non-int of type 'float'

正确解决方案

不需要手动创建重复列,利用pandas的explode方法可以直接实现需求,步骤如下:

# 假设原始DataFrame名为df
# 1. 将titles列按空格拆分为单词列表,生成Word列
df['Word'] = df['titles'].str.split()

# 2. 展开Word列,自动复制对应行的Score和num_comments
result_df = df.explode('Word')[['Word', 'Score', 'num_comments']]

# 可选:重置索引
result_df = result_df.reset_index(drop=True)

错误原因说明

  • 第一个错误:explode的参数只需指定要展开的列名,不需要嵌套assign操作,语法使用错误。
  • 第二个错误:titles_split可能存在空值(float类型的NaN),导致len()无法调用;且完全不需要手动生成重复列,explode会自动处理其他列的复制。
  • 第三个错误:titles_split.str.len()可能返回float类型(存在空值时),无法用来乘以列表,同样这个思路冗余,explode是更高效的解决方案。

内容的提问来源于stack exchange,提问作者DavidReynisson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 09:27:39