如何调整quanteda KWIC函数的关键词前后词数?需前1后5
Great question! The default kwic() function in quanteda uses a symmetric window (same number of words before and after your target term), but there are a couple of straightforward ways to get that asymmetric 1-word before, 5-words after view you need. Let's walk through two reliable approaches:
Method 1: Post-Process a Symmetric KWIC Output
This is the quickest method if you don't want to dive into custom code. We'll first pull a symmetric window large enough to cover both sides, then trim the "pre" column to only the last 1 word.
library(quanteda) # Use quanteda's sample inaugural address corpus for demonstration corpus <- data_corpus_inaugural toks <- tokens(corpus) # First, fetch a symmetric window that covers the larger of the two sides (here, 5 words) kwic_symmetric <- kwic(toks, pattern = "freedom", window = 5) # Trim the "pre" column to keep only the last 1 word; leave "post" as full 5 words kwic_asymmetric <- kwic_symmetric kwic_asymmetric$pre <- stringi::stri_extract_last_words(kwic_asymmetric$pre, n = 1) # View the result – you'll see only 1 word before "freedom" and 5 after kwic_asymmetric
Notes for this method:
- If your target term appears at the very start of a document, the
precolumn will just be empty (since there are no words to extract) – this handles edge cases automatically. - We use
stringi::stri_extract_last_words()because it's efficient and handles whitespace cleanly.
Method 2: Build a Custom Asymmetric KWIC from Tokens
If you need full control over window sizes (or want to avoid relying on post-processing), you can build the KWIC directly from your tokenized text. This is more flexible for complex window requirements.
library(quanteda) corpus <- data_corpus_inaugural toks <- tokens(corpus) # Define your parameters target_term <- "freedom" pre_window_size <- 1 post_window_size <- 5 # Find all positions of the target term in each document term_positions <- tokens_select(toks, pattern = target_term, positions = TRUE) # Initialize an empty data frame to hold our custom KWIC custom_kwic <- data.frame( docname = character(), from = integer(), to = integer(), pre = character(), keyword = character(), post = character(), stringsAsFactors = FALSE ) # Loop through each document and each occurrence of the target term for (doc_id in names(term_positions)) { positions <- term_positions[[doc_id]] for (pos in positions) { # Calculate safe start/end positions (avoid going out of bounds) pre_start <- max(1, pos - pre_window_size) pre_end <- pos - 1 post_start <- pos + 1 post_end <- min(length(toks[[doc_id]]), pos + post_window_size) # Extract tokens for each section pre_text <- if (pre_start <= pre_end) paste(toks[[doc_id]][pre_start:pre_end], collapse = " ") else "" keyword_text <- toks[[doc_id]][pos] post_text <- if (post_start <= post_end) paste(toks[[doc_id]][post_start:post_end], collapse = " ") else "" # Add the row to our custom KWIC custom_kwic <- rbind(custom_kwic, data.frame( docname = doc_id, from = pre_start, to = post_end, pre = pre_text, keyword = keyword_text, post = post_text, stringsAsFactors = FALSE )) } } # Optional: Add quanteda's KWIC class to match the default output format class(custom_kwic) <- c("kwic", class(custom_kwic)) # View your asymmetric KWIC custom_kwic
Notes for this method:
- This approach gives you full control over window sizes, even if you need different values for different terms later.
- The
max()andmin()functions ensure we never try to extract tokens outside the bounds of a document (no errors for terms at the start/end of text).
内容的提问来源于stack exchange,提问作者Ali

