在R中为DataFrame(Tibble)实现条件循环生成词对列
为Tibble生成名词-关联词配对列
问题背景
现有存储句子信息的Tibble结构如下:
| word | position | category | related_word | sentence |
|---|---|---|---|---|
| a | 1 | det | 2 | 1 |
| man | 2 | noun | 3 | 1 |
| sees | 3 | verb | 0 | 1 |
| a | 4 | det | 5 | 1 |
| horse | 5 | noun | 3 | 1 |
| and | 6 | conj | 7 | 1 |
| a | 7 | det | 8 | 1 |
| dog | 8 | noun | 3 | 1 |
需求:遍历每个句子,当行的category为"noun"时,用该行的related_word(对应关联词的position)找到关联词,新增pair列存储"单词 关联词"格式的内容。例如"man"的related_word是3,对应position=3的"sees",所以pair值为"man sees"。
解决方案
方法一:向量化操作(推荐,高效简洁)
利用dplyr和向量匹配实现,无需循环:
library(tibble) library(dplyr) # 构建示例数据 df <- tibble( word = c("a", "man", "sees", "a", "horse", "and", "a", "dog"), position = 1:8, category = c("det", "noun", "verb", "det", "noun", "conj", "det", "noun"), related_word = c(2, 3, 0, 5, 3, 7, 8, 3), sentence = rep(1, 8) ) # 生成pair列 df <- df %>% mutate( pair = if_else( category == "noun", paste(word, word[match(related_word, position)]), NA_character_ ) )
核心逻辑:
match(related_word, position):为每个related_word找到其在position列中的索引位置word[索引]:通过索引提取对应的关联词if_else:仅对category="noun"的行生成配对内容,其余行设为NA
运行后结果:
| word | position | category | related_word | sentence | pair |
|---|---|---|---|---|---|
| a | 1 | det | 2 | 1 | NA |
| man | 2 | noun | 3 | 1 | man sees |
| sees | 3 | verb | 0 | 1 | NA |
| a | 4 | det | 5 | 1 | NA |
| horse | 5 | noun | 3 | 1 | horse sees |
| and | 6 | conj | 7 | 1 | NA |
| a | 7 | det | 8 | 1 | NA |
| dog | 8 | noun | 3 | 1 | dog sees |
方法二:循环实现(按句子遍历)
如果需要严格按句子遍历的循环逻辑,可参考以下代码:
library(tibble) # 构建示例数据 df <- tibble( word = c("a", "man", "sees", "a", "horse", "and", "a", "dog"), position = 1:8, category = c("det", "noun", "verb", "det", "noun", "conj", "det", "noun"), related_word = c(2, 3, 0, 5, 3, 7, 8, 3), sentence = rep(1, 8) ) # 初始化pair列 df$pair <- NA_character_ # 遍历每个唯一句子 for (s in unique(df$sentence)) { # 提取当前句子的子集 sentence_data <- df[df$sentence == s, ] # 遍历子集的每一行 for (i in seq(nrow(sentence_data))) { current_row <- sentence_data[i, ] if (current_row$category == "noun") { # 找到对应position的关联词 related_term <- sentence_data$word[sentence_data$position == current_row$related_word] # 赋值到原数据框对应位置 df$pair[df$sentence == s & df$position == current_row$position] <- paste(current_row$word, related_term) } } }
核心逻辑:
- 先初始化
pair列为NA - 遍历每个句子,提取对应子集
- 对子集内的行逐个判断,若为名词则匹配
position找到关联词,拼接后赋值到原数据框的对应位置
内容的提问来源于stack exchange,提问作者RNewbie
相关产品推荐
相关产品推荐

