You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R按id与spk列将每行单词的DataFrame合并为语句行

用R按id和spk分组合并单词为语句的解决方案

我来帮你搞定这个需求!你想要把每行仅含单个单词的DataFrame,按id和spk两列分组,将每组内的单词拼接成完整语句,用R可以轻松实现,下面给你两种常用方法:

方法一:使用dplyr包(推荐,代码更直观)

首先我们先还原你提供的示例数据:

library(tibble)
df <- tibble(
  id = c(rep(123, 14), rep(789, 9)),
  word = c("hi", "how", "are", "you", "good", "thank", "you", "Anything", "I", "can", "do", "nothing", "for", "now", 
           "what", "is", "the", "issue", "there", "is", "no", "issue", "now", "thank", "you", "for", "contacting", "us..."),
  spk = c(rep(1,4), rep(2,3), rep(1,4), rep(2,3), rep(1,4), rep(2,5), rep(1,4))
)

接下来用dplyr的分组和汇总功能实现拼接:

library(dplyr)

# 按id和spk分组,拼接单词为语句
result_df <- df %>%
  group_by(id, spk) %>%
  summarize(statement = paste(word, collapse = " "), .groups = "drop")

# 查看结果
print(result_df)

这段代码的核心逻辑:

  • group_by(id, spk):指定按id和spk的组合进行分组
  • paste(word, collapse = " "):将每组内的word列元素用空格连接成字符串
  • .groups = "drop":分组完成后取消分组状态,避免后续操作出现警告

方法二:使用Base R的aggregate函数

如果你不想加载额外包,用Base R的aggregate也能实现:

# 按id和spk分组拼接
result_base <- aggregate(word ~ id + spk, data = df, FUN = function(x) paste(x, collapse = " "))

# 查看结果
print(result_base)

两种方法得到的结果是一致的,最终的DataFrame会每行对应一组id+spk的完整语句,比如:

id spk                     statement

1 123 1 hi how are you
2 789 1 what is the issue
3 123 2 good thank you
4 789 2 there is no issue now
5 123 1 Anything I can do
6 789 1 thank you for contacting us...
7 123 2 nothing for now

内容的提问来源于stack exchange,提问作者K. Am

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:59:06