You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在gsub或str_replace中实现文本随机替换

Randomly Replace "It ." with "and" in a Data Frame

Great question! When you need to randomly replace matches instead of replacing all of them, standard gsub or str_replace won't cut it alone—you need to pair them with randomization logic. Here are two practical approaches tailored to different needs:

Approach 1: Replace Matches with a Fixed Probability (Per Match)

If you want each occurrence of "It ." to have a set chance (e.g., 50%) of being replaced with "and", use stringr::str_replace_all with a custom function. This lets you decide replacement on a per-match basis:

library(dplyr)
library(stringr)

# Example data frame
df <- tibble(text_col = c("It . is a test.", "Another line with It . here.", "No match here."))

# Replace each "It ." with 50% probability
df <- df %>%
  mutate(text_col = str_replace_all(text_col, "It \\.", function(match) {
    # Randomly choose to replace or keep the original match
    ifelse(sample(c(TRUE, FALSE), size = 1), "and", match)
  }))

How this works:

  • The regex "It \\." ensures we match the literal "It ." (the \\ escapes the special regex character .).
  • str_replace_all passes every matched "It ." to the custom function, which uses sample() to randomly return either "and" or the original match.
  • Adjust the probability by modifying the sample call—e.g., sample(c(TRUE, FALSE), size = 1, prob = c(0.3, 0.7)) for a 30% replacement chance.

Approach 2: Replace a Specific Number/Percentage of Total Matches

If you want to replace an exact number or proportion of all "It ." occurrences across the entire data frame (instead of per-match probability), follow these steps:

library(stringr)

# Get all text entries and their match locations
all_text <- df$text_col
match_locations <- str_locate_all(all_text, "It \\.")

# Calculate total number of matches
total_matches <- sum(sapply(match_locations, nrow))

# Define how many to replace (e.g., 30% of total matches)
num_to_replace <- floor(total_matches * 0.3)
replace_indices <- sample(seq(total_matches), num_to_replace)

# Iterate through each row and replace selected matches
current_match_count <- 0
df$text_col <- mapply(function(text, locs) {
  if (nrow(locs) == 0) return(text)
  
  # Track which matches in this row are marked for replacement
  row_match_indices <- current_match_count + seq(nrow(locs))
  matches_to_replace <- which(row_match_indices %in% replace_indices)
  
  # Replace each selected match
  for (i in matches_to_replace) {
    text <- str_sub_replace(text, locs[i, 1], locs[i, 2], "and")
  }
  
  current_match_count <<- current_match_count + nrow(locs)
  return(text)
}, all_text, match_locations)

How this works:

  • str_locate_all finds the start/end positions of every "It ." match in each row.
  • We randomly select a subset of all matches (using sample()) to replace.
  • mapply iterates through each row, checks which of its matches are in the selected subset, and replaces only those.

Key Notes

  • Regex Escaping: Always use "It \\." instead of "It ."—the unescaped . matches any character, which would lead to unintended replacements.
  • Flexibility: Use Approach 1 for per-match randomness, Approach 2 for precise control over total replacements.
  • Base R Alternative: If you prefer base R, you can vectorize a custom function with Vectorize(), but it’s less flexible for multiple matches per row:
    replace_random <- function(x) {
      if (grepl("It \\.", x)) {
        sample(c(sub("It \\.", "and", x), x), 1)
      } else {
        x
      }
    }
    df$text_col <- Vectorize(replace_random)(df$text_col)
    

内容的提问来源于stack exchange,提问作者Sebastian Zeki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:00:34