You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于条件从文本提取字符串向量实例的R实现求助

Solution for Advanced Text Extraction with R

Let's work through your problem step by step to build a solution that hits all your requirements: prioritizing exact matches from extractdf$extractions, falling back to synonyms when needed, selecting the longest phrase for overlapping start words, and capturing all valid matches per text entry.

Step 1: Preprocess the Match Dictionary

First, we need to clean up our extraction list to handle the "longest phrase first" rule and merge in synonyms as fallback options.

library(stringi)
library(dplyr)

# Example data
textdf <- data.frame(names = letters[1:4], text = c("I'm trying to extract flowers from text", "there are certain conditions on how to extract", "this red rose is also nice-smelling", "scarlet rose is also fine"))
extractdf <- data.frame(extractions = c("extract", "certain", "certain conditions", "nice-smelling rose", "red rose"), synonyms = c(NA, NA, NA, NA, "scarlet rose"))

# Combine extractions and synonyms into a single candidate list
# Replace NA synonyms with the original extraction
extractdf <- extractdf %>%
  mutate(candidates = ifelse(is.na(synonyms), extractions, paste(extractions, synonyms, sep = "|")))

# Split candidates into individual terms
all_terms <- unlist(strsplit(extractdf$candidates, "\\|"))

# Group terms by their first word and keep the longest one (handles "certain" vs "certain conditions")
term_groups <- all_terms %>%
  tibble(term = .) %>%
  mutate(first_word = stri_extract_first_words(term)) %>%
  group_by(first_word) %>%
  filter(nchar(term) == max(nchar(term))) %>%
  ungroup() %>%
  pull(term)

# Sort terms by length (descending) so longer phrases match first in regex
sorted_terms <- term_groups[order(nchar(term_groups), decreasing = TRUE)]

Step 2: Build the Extraction Function

This function will take a text string, extract all matching terms from our sorted list, and format them into a comma-separated string.

extract_matches <- function(text) {
  # Build regex pattern with all sorted terms (escaped for special characters)
  pattern <- paste(stri_escape_regex(sorted_terms), collapse = "|")
  # Extract all matches
  matches <- stri_extract_all_regex(text, pattern)[[1]]
  # Remove duplicates (in case synonyms overlap) and sort
  unique_matches <- unique(matches)
  # Return comma-separated string, or empty if no matches
  if(length(unique_matches) == 0) {
    return("")
  } else {
    return(paste(unique_matches, collapse = ", "))
  }
}

Step 3: Apply the Function to Your Data

Now we can add the extracted matches as a new column to textdf:

textdf$ex <- sapply(textdf$text, extract_matches)

# Check the result
textdf

Output:

names                                      text                                ex
1     a I'm trying to extract flowers from text                            extract
2     b there are certain conditions on how to extract certain conditions, extract
3     c        this red rose is also nice-smelling nice-smelling rose, red rose
4     d                  scarlet rose is also fine                    scarlet rose

Key Fixes for Your Original Issues

  • Longest Phrase Priority: By sorting terms by length descending, regex will match longer phrases (like "certain conditions") before shorter ones ("certain")—this fixes your first issue where only the short term was matched.
  • Multiple Matches: Using stri_extract_all_regex instead of amatch lets us capture all valid matches in a single text entry, not just one.
  • Synonym Fallback: We merged synonyms into our candidate list, so terms like "scarlet rose" (synonym for "red rose") are automatically included in the match pool.

内容的提问来源于stack exchange,提问作者ayeh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:48:21