基于词元的R语料库词汇查找:quanteda包kwic函数单复数匹配问题
Great question—handling French morphological variations like singular/plural in text mining is super common, and quanteda has some flexible ways to tackle this. Let's break down two reliable approaches for your use case:
Approach 1: Targeted Pattern Matching (Quick & Simple)
If you only need to cover the singular (trou oblong) and plural (trous oblongs) forms of your phrase, you can directly pass both phrases to kwic using phrase(). This works well for fixed, predictable variations:
# Match both singular and plural explicitly result <- kwic(corp2, pattern = phrase(c("trou oblong", "trous oblongs")))
For even more flexibility (in case you run into other minor variations), you can use a regular expression to match optional plural "s" endings:
# Regex pattern to match both singular and plural result <- kwic(corp2, pattern = regex("trou(s)? oblong(s)?"), valuetype = "regex")
Note: This regex assumes the only variation is the plural "s"—it's fast but might catch unintended matches if other words share these roots. Use it if your phrase is contextually unique.
Approach 2: Lemmatization (Robust for Morphological Variations)
For a more scalable solution that handles all morphological forms (not just singular/plural), use lemmatization to reduce words to their base form (lemma). This is ideal if you might encounter other variations later. We'll use udpipe for accurate French lemmatization:
Step 1: Set up dependencies
# Install and load required packages install.packages(c("quanteda", "udpipe")) library(quanteda) library(udpipe) # Download and load the French Universal Dependencies model ud_model <- udpipe_download_model(language = "french-gsd") ud_model <- udpipe_load_model(ud_model$file_model)
Step 2: Lemmatize your corpus and match lemmas
# Convert your corpus to a data frame for annotation corp_df <- convert(corp2, to = "data.frame") # Annotate the text to extract lemmas annotated_text <- udpipe_annotate(ud_model, x = corp_df$text) annotated_df <- as.data.frame(annotated_text) # Create a lemma-based tokens object lemma_tokens <- tokens(annotated_df$lemma, what = "word") %>% tokens_group(groups = annotated_df$doc_id) # Group tokens back to original documents # Now use kwic to match the lemma phrase ("trou" + "oblong") result <- kwic(lemma_tokens, phrase(c("trou oblong")))
This method maps both trous → trou and oblongs → oblong, so your kwic query will catch all instances regardless of number. It's more accurate for French's rich morphology and scales better if you add more phrases later.
Which to Choose?
- Use Approach 1 if you only need to cover singular/plural and want a quick, no-fuss solution.
- Use Approach 2 if you need robustness across all morphological variations or plan to expand your queries to other French phrases.
内容的提问来源于stack exchange,提问作者MysteryGuy

