如何用R按id与spk列将每行单词的DataFrame合并为语句行
用R按id和spk分组合并单词为语句的解决方案
我来帮你搞定这个需求!你想要把每行仅含单个单词的DataFrame,按id和spk两列分组,将每组内的单词拼接成完整语句,用R可以轻松实现,下面给你两种常用方法:
方法一:使用dplyr包(推荐,代码更直观)
首先我们先还原你提供的示例数据:
library(tibble) df <- tibble( id = c(rep(123, 14), rep(789, 9)), word = c("hi", "how", "are", "you", "good", "thank", "you", "Anything", "I", "can", "do", "nothing", "for", "now", "what", "is", "the", "issue", "there", "is", "no", "issue", "now", "thank", "you", "for", "contacting", "us..."), spk = c(rep(1,4), rep(2,3), rep(1,4), rep(2,3), rep(1,4), rep(2,5), rep(1,4)) )
接下来用dplyr的分组和汇总功能实现拼接:
library(dplyr) # 按id和spk分组,拼接单词为语句 result_df <- df %>% group_by(id, spk) %>% summarize(statement = paste(word, collapse = " "), .groups = "drop") # 查看结果 print(result_df)
这段代码的核心逻辑:
group_by(id, spk):指定按id和spk的组合进行分组paste(word, collapse = " "):将每组内的word列元素用空格连接成字符串.groups = "drop":分组完成后取消分组状态,避免后续操作出现警告
方法二:使用Base R的aggregate函数
如果你不想加载额外包,用Base R的aggregate也能实现:
# 按id和spk分组拼接 result_base <- aggregate(word ~ id + spk, data = df, FUN = function(x) paste(x, collapse = " ")) # 查看结果 print(result_base)
两种方法得到的结果是一致的,最终的DataFrame会每行对应一组id+spk的完整语句,比如:
id spk statement1 123 1 hi how are you
2 789 1 what is the issue
3 123 2 good thank you
4 789 2 there is no issue now
5 123 1 Anything I can do
6 789 1 thank you for contacting us...
7 123 2 nothing for now
内容的提问来源于stack exchange,提问作者K. Am
相关产品推荐
相关产品推荐

