You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中使用gsub提取指定术语前后的n个词?

Extract N Words Before/After a Specified Term in R

I’ve put together a practical function that pulls the exact number of words before and after your target term from each text entry. It handles edge cases like the term being at the start/end of a sentence, ignores messy punctuation when splitting words, and can optionally match terms case-insensitively.

Step-by-Step Solution Code

First, let’s define the function (we’ll use the stringr package for clean string handling):

library(stringr)

extract_context <- function(text, term, n, case_sensitive = FALSE) {
  # Store results for each text entry
  results <- list()
  
  # Process each text one by one
  for (i in seq_along(text)) {
    current_text <- text[i]
    
    # Split text into words (strip punctuation, remove empty strings)
    words <- str_split(current_text, "\\W+")[[1]]
    words <- words[words != ""]
    
    # Find where the term appears in the word list
    if (case_sensitive) {
      term_positions <- which(words == term)
    } else {
      term_positions <- which(tolower(words) == tolower(term))
    }
    
    # Handle cases where the term isn't found
    if (length(term_positions) == 0) {
      results[[i]] <- paste("Term '", term, "' not found in text ", i, sep = "")
      next
    }
    
    # Pull context for each occurrence of the term
    context_list <- list()
    for (pos in term_positions) {
      # Calculate safe start/end indices (don't go out of bounds)
      start_idx <- max(1, pos - n)
      end_idx <- min(length(words), pos + n)
      
      # Extract and format the context
      context_words <- words[start_idx:end_idx]
      context_str <- paste(context_words, collapse = " ")
      
      # Label each occurrence for clarity
      context_list[[paste0("Occurrence ", which(term_positions == pos))]] <- context_str
    }
    
    results[[i]] <- context_list
  }
  
  # Name results to match original text entries
  names(results) <- paste("Text", seq_along(text))
  return(results)
}

How to Use the Function

Let’s test it with your sample data, targeting the term "game" and pulling 2 words before and after:

# Your sample text data
a <- c("The day was nice and dry, when she came for our game we were ready and then she left.", 
       "The day was nice and dry, when she came for our game, but we were not ready. She left after she waited 5 minutes.", 
       "The day was nice and dry, when she came, we were not here. Our game was not completed timely, but it was completed after one hour.")

# Run the extraction
extract_context(a, term = "game", n = 2)

Sample Output

$`Text 1`
$`Text 1`$`Occurrence 1`
[1] "for our game we were"

$`Text 2`
$`Text 2`$`Occurrence 1`
[1] "for our game but we"

$`Text 3`
$`Text 3`$`Occurrence 1`
[1] "here Our game was not"

Key Features

  • Case Insensitivity: By default, it matches terms regardless of case (e.g., "Game" and "game" are treated the same). Toggle case_sensitive = TRUE if you need exact case matching.
  • Edge Case Safety: If the term is the first word in a text, it only pulls words after it. If it’s the last word, it only pulls words before.
  • Multiple Occurrences: If a text has multiple instances of the term, it returns context for each one separately.
  • Punctuation Handling: Splits text on non-word characters, so punctuation doesn’t mess up word counting or matching.

Quick Notes

  • If you don’t have stringr installed, run install.packages("stringr") first.
  • If you want to keep punctuation attached to words (e.g., "game," instead of "game"), change the split pattern to str_split(current_text, "\\s+") (split on whitespace), but you’ll need to adjust your term to include punctuation if needed.

内容的提问来源于stack exchange,提问作者user3357059

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:20:58