能否结合正则表达式与quanteda包的kwic函数检索邻近非连续词汇?
Hey there! Let's tackle your two technical questions step by step:
1. Can the kwic() function find adjacent words?
Absolutely! The kwic() function from the quanteda package is perfect for this task. By adjusting the window parameter, you can control how close the target words need to be. For example, if you want strictly adjacent words (no gaps between them), set window = 1 (to capture the immediate next/previous word) and filter for your second term. Here's a quick example:
library(quanteda) # Create a sample corpus my_corpus <- corpus(c( "The quick brown fox jumps over the lazy dog", "A quick red fox", "The brown quick fox" )) # Look for "quick" with a 1-word window, then filter for entries containing "brown" adjacent_matches <- kwic(my_corpus, pattern = "quick", window = 1) %>% subset(grepl("brown", target)) print(adjacent_matches)
This will return only the cases where "quick" and "brown" are right next to each other.
2. Retrieving documents with two target words separated by some words
Yes, you can combine regular expressions with kwic() to do this—exactly what you're asking for. The key is to build a regex pattern that matches your two target words with a specified number of words in between, then pass that pattern to kwic() with valuetype = "regex".
Let's say you want to find instances where "quick" and "fox" are separated by 1 to 5 words (and either order works). Here's how to implement it:
# Define your target words word_a <- "quick" word_b <- "fox" # Build regex pattern: matches word_a followed by 1-5 words then word_b, or vice versa gap_regex <- paste0( "(", word_a, "(\\W+\\w+){1,5}", word_b, ")|(", word_b, "(\\W+\\w+){1,5}", word_a, ")" ) # Use kwic with regex mode gap_matches <- kwic(my_corpus, pattern = gap_regex, valuetype = "regex") # Get unique document names that match matching_docs <- unique(docnames(gap_matches)) cat("Documents with the spaced word pairs:", matching_docs, "\n")
A quick breakdown of the regex:
(\\W+\\w+)matches a single word (non-word characters like spaces plus a word){1,5}specifies we want 1 to 5 of these word units (adjust the numbers to change the gap size)- The
|lets us match both word orders (word_a before word_b, or word_b before word_a)
If you don't care about the gap size (just that there's at least one word in between), replace {1,5} with {1,}. And if you need case-insensitive matching, add case_insensitive = TRUE to the kwic() call.
内容的提问来源于stack exchange,提问作者MysteryGuy

