基于条件从文本提取字符串向量实例的R实现求助
Solution for Advanced Text Extraction with R
Let's work through your problem step by step to build a solution that hits all your requirements: prioritizing exact matches from extractdf$extractions, falling back to synonyms when needed, selecting the longest phrase for overlapping start words, and capturing all valid matches per text entry.
Step 1: Preprocess the Match Dictionary
First, we need to clean up our extraction list to handle the "longest phrase first" rule and merge in synonyms as fallback options.
library(stringi) library(dplyr) # Example data textdf <- data.frame(names = letters[1:4], text = c("I'm trying to extract flowers from text", "there are certain conditions on how to extract", "this red rose is also nice-smelling", "scarlet rose is also fine")) extractdf <- data.frame(extractions = c("extract", "certain", "certain conditions", "nice-smelling rose", "red rose"), synonyms = c(NA, NA, NA, NA, "scarlet rose")) # Combine extractions and synonyms into a single candidate list # Replace NA synonyms with the original extraction extractdf <- extractdf %>% mutate(candidates = ifelse(is.na(synonyms), extractions, paste(extractions, synonyms, sep = "|"))) # Split candidates into individual terms all_terms <- unlist(strsplit(extractdf$candidates, "\\|")) # Group terms by their first word and keep the longest one (handles "certain" vs "certain conditions") term_groups <- all_terms %>% tibble(term = .) %>% mutate(first_word = stri_extract_first_words(term)) %>% group_by(first_word) %>% filter(nchar(term) == max(nchar(term))) %>% ungroup() %>% pull(term) # Sort terms by length (descending) so longer phrases match first in regex sorted_terms <- term_groups[order(nchar(term_groups), decreasing = TRUE)]
Step 2: Build the Extraction Function
This function will take a text string, extract all matching terms from our sorted list, and format them into a comma-separated string.
extract_matches <- function(text) { # Build regex pattern with all sorted terms (escaped for special characters) pattern <- paste(stri_escape_regex(sorted_terms), collapse = "|") # Extract all matches matches <- stri_extract_all_regex(text, pattern)[[1]] # Remove duplicates (in case synonyms overlap) and sort unique_matches <- unique(matches) # Return comma-separated string, or empty if no matches if(length(unique_matches) == 0) { return("") } else { return(paste(unique_matches, collapse = ", ")) } }
Step 3: Apply the Function to Your Data
Now we can add the extracted matches as a new column to textdf:
textdf$ex <- sapply(textdf$text, extract_matches) # Check the result textdf
Output:
names text ex 1 a I'm trying to extract flowers from text extract 2 b there are certain conditions on how to extract certain conditions, extract 3 c this red rose is also nice-smelling nice-smelling rose, red rose 4 d scarlet rose is also fine scarlet rose
Key Fixes for Your Original Issues
- Longest Phrase Priority: By sorting terms by length descending, regex will match longer phrases (like "certain conditions") before shorter ones ("certain")—this fixes your first issue where only the short term was matched.
- Multiple Matches: Using
stri_extract_all_regexinstead ofamatchlets us capture all valid matches in a single text entry, not just one. - Synonym Fallback: We merged synonyms into our candidate list, so terms like "scarlet rose" (synonym for "red rose") are automatically included in the match pool.
内容的提问来源于stack exchange,提问作者ayeh
相关产品推荐
相关产品推荐

