You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整quanteda KWIC函数的关键词前后词数?需前1后5

Asymmetric Window for quanteda's KWIC Function (1 Word Before, 5 Words After)

Great question! The default kwic() function in quanteda uses a symmetric window (same number of words before and after your target term), but there are a couple of straightforward ways to get that asymmetric 1-word before, 5-words after view you need. Let's walk through two reliable approaches:

Method 1: Post-Process a Symmetric KWIC Output

This is the quickest method if you don't want to dive into custom code. We'll first pull a symmetric window large enough to cover both sides, then trim the "pre" column to only the last 1 word.

library(quanteda)

# Use quanteda's sample inaugural address corpus for demonstration
corpus <- data_corpus_inaugural
toks <- tokens(corpus)

# First, fetch a symmetric window that covers the larger of the two sides (here, 5 words)
kwic_symmetric <- kwic(toks, pattern = "freedom", window = 5)

# Trim the "pre" column to keep only the last 1 word; leave "post" as full 5 words
kwic_asymmetric <- kwic_symmetric
kwic_asymmetric$pre <- stringi::stri_extract_last_words(kwic_asymmetric$pre, n = 1)

# View the result – you'll see only 1 word before "freedom" and 5 after
kwic_asymmetric

Notes for this method:

  • If your target term appears at the very start of a document, the pre column will just be empty (since there are no words to extract) – this handles edge cases automatically.
  • We use stringi::stri_extract_last_words() because it's efficient and handles whitespace cleanly.

Method 2: Build a Custom Asymmetric KWIC from Tokens

If you need full control over window sizes (or want to avoid relying on post-processing), you can build the KWIC directly from your tokenized text. This is more flexible for complex window requirements.

library(quanteda)

corpus <- data_corpus_inaugural
toks <- tokens(corpus)

# Define your parameters
target_term <- "freedom"
pre_window_size <- 1
post_window_size <- 5

# Find all positions of the target term in each document
term_positions <- tokens_select(toks, pattern = target_term, positions = TRUE)

# Initialize an empty data frame to hold our custom KWIC
custom_kwic <- data.frame(
  docname = character(),
  from = integer(),
  to = integer(),
  pre = character(),
  keyword = character(),
  post = character(),
  stringsAsFactors = FALSE
)

# Loop through each document and each occurrence of the target term
for (doc_id in names(term_positions)) {
  positions <- term_positions[[doc_id]]
  for (pos in positions) {
    # Calculate safe start/end positions (avoid going out of bounds)
    pre_start <- max(1, pos - pre_window_size)
    pre_end <- pos - 1
    post_start <- pos + 1
    post_end <- min(length(toks[[doc_id]]), pos + post_window_size)
    
    # Extract tokens for each section
    pre_text <- if (pre_start <= pre_end) paste(toks[[doc_id]][pre_start:pre_end], collapse = " ") else ""
    keyword_text <- toks[[doc_id]][pos]
    post_text <- if (post_start <= post_end) paste(toks[[doc_id]][post_start:post_end], collapse = " ") else ""
    
    # Add the row to our custom KWIC
    custom_kwic <- rbind(custom_kwic, data.frame(
      docname = doc_id,
      from = pre_start,
      to = post_end,
      pre = pre_text,
      keyword = keyword_text,
      post = post_text,
      stringsAsFactors = FALSE
    ))
  }
}

# Optional: Add quanteda's KWIC class to match the default output format
class(custom_kwic) <- c("kwic", class(custom_kwic))

# View your asymmetric KWIC
custom_kwic

Notes for this method:

  • This approach gives you full control over window sizes, even if you need different values for different terms later.
  • The max() and min() functions ensure we never try to extract tokens outside the bounds of a document (no errors for terms at the start/end of text).

内容的提问来源于stack exchange,提问作者Ali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:17:21