You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

dfm_replace词元替换后原词映射与KWIC检索方法(附R代码场景)

Hey there, let's work through your quanteda questions step by step—since you need to map lemmas back to original words for KWIC searches after using dfm_replace() (or better, tokens_replace()), here's how to do it properly, including fixing your code snippet.

Key Background

The main thing to remember: KWIC relies on word position context, which gets lost when you jump straight to a dfm. So we need to keep our original tokenized text alongside the lemma-tokenized version to maintain alignment between lemmas and their original words.


1. How to Map Lemmas Back to Original Words for KWIC

Here's the core workflow:

  • Keep a copy of your cleaned, original tokens (before lemma replacement)
  • Create a clear lemma-to-original-word mapping (best done with a named vector)
  • Use the mapping to reverse-lookup original words when you want to run KWIC for a target lemma

2. Fixing Your R Code & Implementing the Mapping

Let's adjust your code to follow the correct workflow, with comments explaining each step:

# Load quanteda (make sure it's up to date!)
library(quanteda)

# Your original dataset
df <- data.frame(
  text = c("Ow now brown cow","Unique New York", "The sassy salesmans agonized about a bigger sale"),
  person = c("Jim", "John", "Jim"),
  year = c(1994, 1995, 1996),
  stringsAsFactors = FALSE
)
x <- corpus(df)

# Step 1: Create cleaned, original tokens (keep padding to maintain position alignment)
toks_original <- tokens(x) %>%
  tokens_remove(stopwords("english"), padding = TRUE) %>%
  tokens_remove(numbers = TRUE, punct = TRUE, symbols = TRUE) # Match dfm cleaning params

# Step 2: Prepare your lemma mapping
# Assume your lemmaFile is a 2-column data frame with "original" (raw words) and "lemma" columns
# Here's a simulated version—replace this with your actual lemma data
lemmaFile <- data.frame(
  original = c("salesmans", "agonized", "sassy", "sale"),
  lemma = c("salesman", "agonize", "sass", "sell"),
  stringsAsFactors = FALSE
)
# Convert to a named vector (critical for quanteda replacements: names = original words, values = lemmas)
lemma_map <- setNames(lemmaFile$lemma, lemmaFile$original)

# Step 3: Create lemma-tokenized version (keep alignment with original tokens)
toks_lemmas <- tokens_replace(toks_original, pattern = names(lemma_map), replacement = lemma_map)

# Step 4: Build your lemma-based dfm (with ngrams as you wanted)
xdfmr <- dfm(toks_lemmas, ngrams = 1:3, concatenator = " ")

# Step 5: Map lemmas back to original words for KWIC
# Example: Search for all instances of the lemma "sell"
target_lemma <- "sell"

# Option 1: Reverse-lookup original words from the lemma map
original_target_words <- names(lemma_map)[lemma_map == target_lemma]
# Run KWIC on original tokens to get full context
kwic_result <- kwic(toks_original, pattern = original_target_words)
print(kwic_result)

# Option 2: Use position alignment (if you need precise location matches)
# Find positions of the target lemma in the lemma tokens
lemma_matches <- tokens_select(toks_lemmas, pattern = target_lemma, selection = "keep")
match_positions <- lapply(lemma_matches, function(doc) which(doc == target_lemma))

# Extract context from original tokens using matching positions
for (doc_idx in seq_along(match_positions)) {
  if (length(match_positions[[doc_idx]]) > 0) {
    for (pos in match_positions[[doc_idx]]) {
      # Grab 2 words of context on each side (adjust as needed)
      left_context <- paste(toks_original[[doc_idx]][max(1, pos-2):(pos-1)], collapse = " ")
      right_context <- paste(toks_original[[doc_idx]][(pos+1):min(length(toks_original[[doc_idx]]), pos+2)], collapse = " ")
      original_word <- toks_original[[doc_idx]][pos]
      
      cat(sprintf("Doc %d, Position %d: %s *%s* %s\n", 
                  doc_idx, pos, left_context, original_word, right_context))
    }
  }
}

Critical Notes

  • Why use tokens_replace() instead of dfm_replace() first? A dfm is a document-feature matrix—you lose word position information, which is essential for KWIC. By working with tokens first, you keep the alignment between original words and their lemmas.
  • Padding = TRUE: This ensures that even if we remove stopwords, the position indices of remaining words stay the same between toks_original and toks_lemmas.
  • Named vector mapping: Quanteda’s replacement functions work best with named vectors, where the names are the terms to replace, and the values are the replacements (lemmas, in this case).

内容的提问来源于stack exchange,提问作者Ted Mosby

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:41:13