You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中为DataFrame(Tibble)实现条件循环生成词对列

为Tibble生成名词-关联词配对列

问题背景

现有存储句子信息的Tibble结构如下:

wordpositioncategoryrelated_wordsentence
a1det21
man2noun31
sees3verb01
a4det51
horse5noun31
and6conj71
a7det81
dog8noun31

需求:遍历每个句子,当行的category为"noun"时,用该行的related_word(对应关联词的position)找到关联词,新增pair列存储"单词 关联词"格式的内容。例如"man"的related_word是3,对应position=3的"sees",所以pair值为"man sees"。

解决方案

方法一:向量化操作(推荐,高效简洁)

利用dplyr和向量匹配实现,无需循环:

library(tibble)
library(dplyr)

# 构建示例数据
df <- tibble(
  word = c("a", "man", "sees", "a", "horse", "and", "a", "dog"),
  position = 1:8,
  category = c("det", "noun", "verb", "det", "noun", "conj", "det", "noun"),
  related_word = c(2, 3, 0, 5, 3, 7, 8, 3),
  sentence = rep(1, 8)
)

# 生成pair列
df <- df %>%
  mutate(
    pair = if_else(
      category == "noun",
      paste(word, word[match(related_word, position)]),
      NA_character_
    )
  )

核心逻辑:

  • match(related_word, position):为每个related_word找到其在position列中的索引位置
  • word[索引]:通过索引提取对应的关联词
  • if_else:仅对category="noun"的行生成配对内容,其余行设为NA

运行后结果:

wordpositioncategoryrelated_wordsentencepair
a1det21NA
man2noun31man sees
sees3verb01NA
a4det51NA
horse5noun31horse sees
and6conj71NA
a7det81NA
dog8noun31dog sees

方法二:循环实现(按句子遍历)

如果需要严格按句子遍历的循环逻辑,可参考以下代码:

library(tibble)

# 构建示例数据
df <- tibble(
  word = c("a", "man", "sees", "a", "horse", "and", "a", "dog"),
  position = 1:8,
  category = c("det", "noun", "verb", "det", "noun", "conj", "det", "noun"),
  related_word = c(2, 3, 0, 5, 3, 7, 8, 3),
  sentence = rep(1, 8)
)

# 初始化pair列
df$pair <- NA_character_

# 遍历每个唯一句子
for (s in unique(df$sentence)) {
  # 提取当前句子的子集
  sentence_data <- df[df$sentence == s, ]
  # 遍历子集的每一行
  for (i in seq(nrow(sentence_data))) {
    current_row <- sentence_data[i, ]
    if (current_row$category == "noun") {
      # 找到对应position的关联词
      related_term <- sentence_data$word[sentence_data$position == current_row$related_word]
      # 赋值到原数据框对应位置
      df$pair[df$sentence == s & df$position == current_row$position] <- paste(current_row$word, related_term)
    }
  }
}

核心逻辑:

  1. 先初始化pair列为NA
  2. 遍历每个句子,提取对应子集
  3. 对子集内的行逐个判断,若为名词则匹配position找到关联词,拼接后赋值到原数据框的对应位置

内容的提问来源于stack exchange,提问作者RNewbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 02:10:18