如何在R语言中获取句子中指定词汇的出现顺序?
R语言获取句子中指定词汇的出现顺序
需求说明
需要从给定句子集合中,提取目标词汇的出现顺序位次,输出类似如下格式的结果:
sentence cat dog cow 1 1 2 NA 2 2 1 3
解决方案
使用tidyverse工具链(stringr、dplyr、tidyr)可以高效实现需求,步骤如下:
1. 准备示例数据
library(tidyverse) # 示例句子集合 sentences <- tibble( sentence = c( "The cat chased the dog", "The dog barked at the cat, then the cow mooed" ) ) # 目标词汇集合 animals <- c("cat", "dog", "cow")
2. 核心处理代码
result <- sentences %>% # 给句子添加唯一编号 mutate(id = row_number()) %>% # 去除标点并拆分句子为单个词汇 mutate(word = str_split(str_remove_all(sentence, "[^a-zA-Z\\s]"), "\\s+")) %>% # 展开为每行一个词汇的长格式 unnest(word) %>% # 仅保留目标动物词汇 filter(word %in% animals) %>% # 按句子分组,标记每个词汇的出现顺序 group_by(id) %>% mutate(rank = row_number()) %>% ungroup() %>% # 转换为宽格式,未出现的词汇填充NA pivot_wider( id_cols = id, names_from = word, values_from = rank, values_fill = NA_integer_ ) %>% # 合并原句子内容 left_join(sentences %>% mutate(id = row_number()), by = "id") %>% # 调整列顺序 select(id, sentence, all_of(animals)) # 查看结果 print(result)
3. 最终输出
运行代码后会得到符合需求的结果:
# A tibble: 2 × 5 id sentence cat dog cow <int> <chr> <int> <int> <int> 1 1 The cat chased the dog 1 2 NA 2 2 The dog barked at the cat, then the cow mooed 2 1 3
关于str_locate_all的局限性
str_locate_all返回的是词汇在句子中的字符位置索引,而非出现顺序位次。要转换为目标结果需要额外排序、去重等操作,不如直接拆分词汇后标记顺序的方法直接高效。
内容的提问来源于stack exchange,提问作者agbarnett
相关产品推荐
相关产品推荐

