You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中获取句子中指定词汇的出现顺序?

R语言获取句子中指定词汇的出现顺序

需求说明

需要从给定句子集合中,提取目标词汇的出现顺序位次,输出类似如下格式的结果:

sentence cat dog cow
1          1   2  NA
2          2   1   3 

解决方案

使用tidyverse工具链(stringr、dplyr、tidyr)可以高效实现需求,步骤如下:

1. 准备示例数据

library(tidyverse)

# 示例句子集合
sentences <- tibble(
  sentence = c(
    "The cat chased the dog",
    "The dog barked at the cat, then the cow mooed"
  )
)

# 目标词汇集合
animals <- c("cat", "dog", "cow")

2. 核心处理代码

result <- sentences %>%
  # 给句子添加唯一编号
  mutate(id = row_number()) %>%
  # 去除标点并拆分句子为单个词汇
  mutate(word = str_split(str_remove_all(sentence, "[^a-zA-Z\\s]"), "\\s+")) %>%
  # 展开为每行一个词汇的长格式
  unnest(word) %>%
  # 仅保留目标动物词汇
  filter(word %in% animals) %>%
  # 按句子分组,标记每个词汇的出现顺序
  group_by(id) %>%
  mutate(rank = row_number()) %>%
  ungroup() %>%
  # 转换为宽格式,未出现的词汇填充NA
  pivot_wider(
    id_cols = id,
    names_from = word,
    values_from = rank,
    values_fill = NA_integer_
  ) %>%
  # 合并原句子内容
  left_join(sentences %>% mutate(id = row_number()), by = "id") %>%
  # 调整列顺序
  select(id, sentence, all_of(animals))

# 查看结果
print(result)

3. 最终输出

运行代码后会得到符合需求的结果:

# A tibble: 2 × 5
     id sentence                                   cat   dog   cow
  <int> <chr>                                   <int> <int> <int>
1     1 The cat chased the dog                      1     2    NA
2     2 The dog barked at the cat, then the cow mooed     2     1     3

关于str_locate_all的局限性

str_locate_all返回的是词汇在句子中的字符位置索引,而非出现顺序位次。要转换为目标结果需要额外排序、去重等操作,不如直接拆分词汇后标记顺序的方法直接高效。

内容的提问来源于stack exchange,提问作者agbarnett

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 12:21:00