dfm_replace词元替换后原词映射与KWIC检索方法(附R代码场景)
Hey there, let's work through your quanteda questions step by step—since you need to map lemmas back to original words for KWIC searches after using dfm_replace() (or better, tokens_replace()), here's how to do it properly, including fixing your code snippet.
Key Background
The main thing to remember: KWIC relies on word position context, which gets lost when you jump straight to a dfm. So we need to keep our original tokenized text alongside the lemma-tokenized version to maintain alignment between lemmas and their original words.
1. How to Map Lemmas Back to Original Words for KWIC
Here's the core workflow:
- Keep a copy of your cleaned, original tokens (before lemma replacement)
- Create a clear lemma-to-original-word mapping (best done with a named vector)
- Use the mapping to reverse-lookup original words when you want to run KWIC for a target lemma
2. Fixing Your R Code & Implementing the Mapping
Let's adjust your code to follow the correct workflow, with comments explaining each step:
# Load quanteda (make sure it's up to date!) library(quanteda) # Your original dataset df <- data.frame( text = c("Ow now brown cow","Unique New York", "The sassy salesmans agonized about a bigger sale"), person = c("Jim", "John", "Jim"), year = c(1994, 1995, 1996), stringsAsFactors = FALSE ) x <- corpus(df) # Step 1: Create cleaned, original tokens (keep padding to maintain position alignment) toks_original <- tokens(x) %>% tokens_remove(stopwords("english"), padding = TRUE) %>% tokens_remove(numbers = TRUE, punct = TRUE, symbols = TRUE) # Match dfm cleaning params # Step 2: Prepare your lemma mapping # Assume your lemmaFile is a 2-column data frame with "original" (raw words) and "lemma" columns # Here's a simulated version—replace this with your actual lemma data lemmaFile <- data.frame( original = c("salesmans", "agonized", "sassy", "sale"), lemma = c("salesman", "agonize", "sass", "sell"), stringsAsFactors = FALSE ) # Convert to a named vector (critical for quanteda replacements: names = original words, values = lemmas) lemma_map <- setNames(lemmaFile$lemma, lemmaFile$original) # Step 3: Create lemma-tokenized version (keep alignment with original tokens) toks_lemmas <- tokens_replace(toks_original, pattern = names(lemma_map), replacement = lemma_map) # Step 4: Build your lemma-based dfm (with ngrams as you wanted) xdfmr <- dfm(toks_lemmas, ngrams = 1:3, concatenator = " ") # Step 5: Map lemmas back to original words for KWIC # Example: Search for all instances of the lemma "sell" target_lemma <- "sell" # Option 1: Reverse-lookup original words from the lemma map original_target_words <- names(lemma_map)[lemma_map == target_lemma] # Run KWIC on original tokens to get full context kwic_result <- kwic(toks_original, pattern = original_target_words) print(kwic_result) # Option 2: Use position alignment (if you need precise location matches) # Find positions of the target lemma in the lemma tokens lemma_matches <- tokens_select(toks_lemmas, pattern = target_lemma, selection = "keep") match_positions <- lapply(lemma_matches, function(doc) which(doc == target_lemma)) # Extract context from original tokens using matching positions for (doc_idx in seq_along(match_positions)) { if (length(match_positions[[doc_idx]]) > 0) { for (pos in match_positions[[doc_idx]]) { # Grab 2 words of context on each side (adjust as needed) left_context <- paste(toks_original[[doc_idx]][max(1, pos-2):(pos-1)], collapse = " ") right_context <- paste(toks_original[[doc_idx]][(pos+1):min(length(toks_original[[doc_idx]]), pos+2)], collapse = " ") original_word <- toks_original[[doc_idx]][pos] cat(sprintf("Doc %d, Position %d: %s *%s* %s\n", doc_idx, pos, left_context, original_word, right_context)) } } }
Critical Notes
- Why use
tokens_replace()instead ofdfm_replace()first? A dfm is a document-feature matrix—you lose word position information, which is essential for KWIC. By working with tokens first, you keep the alignment between original words and their lemmas. - Padding = TRUE: This ensures that even if we remove stopwords, the position indices of remaining words stay the same between
toks_originalandtoks_lemmas. - Named vector mapping: Quanteda’s replacement functions work best with named vectors, where the names are the terms to replace, and the values are the replacements (lemmas, in this case).
内容的提问来源于stack exchange,提问作者Ted Mosby
相关产品推荐
相关产品推荐

