You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按指定单词数拆分DataFrame文本列并生成同内容新行?

按指定单词数拆分DataFrame文本列并生成新行

我来帮你搞定这个按单词拆分文本列的需求!字符拆分确实没法满足按单词计数的要求,咱们用R的tidyverse工具包就能轻松实现,下面是具体的步骤和代码:

核心思路

先把文本拆成单个单词的列表,再按每30个单词分组拼接成新的文本片段,最后将这些片段展开成新行,同时原行的其他所有列数据会自动复制到对应新行中。

完整代码示例

先加载所需工具包,再处理你的示例数据:

# 加载tidyverse包(整合了dplyr、stringr、purrr等实用工具)
library(tidyverse)

# 你的示例数据(补全了部分文本方便测试拆分效果)
df1 <- tibble(
  V1 = c(1, 2, 3),
  V2 = c('Red', 'Blue', 'Red'),
  text = c(
    'Folly words widow one downs few age every seven. If miss part by fact he park just shew. Discovered had get considered projection so literature. Was admiration unreserved discovered projecting inquietude. Way necessary had intention happiness but september delighted his. Are gay head need down draw. Misery wonder enable mutual get set oppose the uneasy. End why melancholy estimating her had indulgence middletons. Say ferrars demands besides her address. Blind going you merit few fancy their.',
    'Blue text example with more words to test the splitting. Let us add enough words here so that we can see how the 30-word split works. Words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words words.',
    'Short text that won\'t need splitting since it has less than 30 words.'
  )
)

# 执行拆分操作
df_split <- df1 %>%
  # 把文本拆成单个单词的列表(匹配任意数量空格,处理连续空格的情况)
  mutate(word_list = str_split(text, "\\s+")) %>%
  # 把每个单词列表按每30个单词分组,生成子列表
  mutate(split_text = map(word_list, ~ split(., ceiling(seq_along(.) / 30)))) %>%
  # 把每个子列表的单词重新拼接成完整文本片段
  mutate(split_text = map(split_text, ~ map_chr(., paste, collapse = " "))) %>%
  # 移除临时的word_list列,展开split_text成新行
  select(-word_list) %>%
  unnest(split_text)

# 查看拆分后的结果
print(df_split)

代码细节解释

  • str_split(text, "\\s+"):用正则表达式匹配任意数量的空格拆分文本,避免连续空格导致的空元素问题。
  • map(word_list, ~ split(., ceiling(seq_along(.) / 30))):对每个单词列表,用ceiling(seq_along(.)/30)生成分组标识(每30个单词一组),再按标识拆分列表。
  • map(split_text, ~ map_chr(., paste, collapse = " ")):把每个分组的单词重新拼接成连贯的文本片段。
  • unnest(split_text):把列表格式的拆分结果展开成独立行,原行的V1、V2列会自动同步到每一行。

可选优化:处理标点与单词粘连

如果你的文本里有大量seven.这类标点和单词粘连的情况,想要拆分标点或排除标点计数,可以在拆分前先清洗文本:

mutate(text_clean = str_replace_all(text, "([[:punct:]])", " \\1 ")) %>%
mutate(word_list = str_split(text_clean, "\\s+"))

这样会在标点前后添加空格,拆分时标点会被当成单独元素;如果不想把标点算成单词,后续可以用filter过滤掉标点元素。

内容的提问来源于stack exchange,提问作者Nuria

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:20:54