如何在gsub或str_replace中实现文本随机替换
Randomly Replace "It ." with "and" in a Data Frame
Great question! When you need to randomly replace matches instead of replacing all of them, standard gsub or str_replace won't cut it alone—you need to pair them with randomization logic. Here are two practical approaches tailored to different needs:
Approach 1: Replace Matches with a Fixed Probability (Per Match)
If you want each occurrence of "It ." to have a set chance (e.g., 50%) of being replaced with "and", use stringr::str_replace_all with a custom function. This lets you decide replacement on a per-match basis:
library(dplyr) library(stringr) # Example data frame df <- tibble(text_col = c("It . is a test.", "Another line with It . here.", "No match here.")) # Replace each "It ." with 50% probability df <- df %>% mutate(text_col = str_replace_all(text_col, "It \\.", function(match) { # Randomly choose to replace or keep the original match ifelse(sample(c(TRUE, FALSE), size = 1), "and", match) }))
How this works:
- The regex
"It \\."ensures we match the literal "It ." (the\\escapes the special regex character.). str_replace_allpasses every matched "It ." to the custom function, which usessample()to randomly return either "and" or the original match.- Adjust the probability by modifying the
samplecall—e.g.,sample(c(TRUE, FALSE), size = 1, prob = c(0.3, 0.7))for a 30% replacement chance.
Approach 2: Replace a Specific Number/Percentage of Total Matches
If you want to replace an exact number or proportion of all "It ." occurrences across the entire data frame (instead of per-match probability), follow these steps:
library(stringr) # Get all text entries and their match locations all_text <- df$text_col match_locations <- str_locate_all(all_text, "It \\.") # Calculate total number of matches total_matches <- sum(sapply(match_locations, nrow)) # Define how many to replace (e.g., 30% of total matches) num_to_replace <- floor(total_matches * 0.3) replace_indices <- sample(seq(total_matches), num_to_replace) # Iterate through each row and replace selected matches current_match_count <- 0 df$text_col <- mapply(function(text, locs) { if (nrow(locs) == 0) return(text) # Track which matches in this row are marked for replacement row_match_indices <- current_match_count + seq(nrow(locs)) matches_to_replace <- which(row_match_indices %in% replace_indices) # Replace each selected match for (i in matches_to_replace) { text <- str_sub_replace(text, locs[i, 1], locs[i, 2], "and") } current_match_count <<- current_match_count + nrow(locs) return(text) }, all_text, match_locations)
How this works:
str_locate_allfinds the start/end positions of every "It ." match in each row.- We randomly select a subset of all matches (using
sample()) to replace. mapplyiterates through each row, checks which of its matches are in the selected subset, and replaces only those.
Key Notes
- Regex Escaping: Always use
"It \\."instead of"It ."—the unescaped.matches any character, which would lead to unintended replacements. - Flexibility: Use Approach 1 for per-match randomness, Approach 2 for precise control over total replacements.
- Base R Alternative: If you prefer base R, you can vectorize a custom function with
Vectorize(), but it’s less flexible for multiple matches per row:replace_random <- function(x) { if (grepl("It \\.", x)) { sample(c(sub("It \\.", "and", x), x), 1) } else { x } } df$text_col <- Vectorize(replace_random)(df$text_col)
内容的提问来源于stack exchange,提问作者Sebastian Zeki
相关产品推荐
相关产品推荐

