在R语言中对俄英双语语料库使用双词干提取的方法咨询
Hey there! Great question about handling bilingual stemming for Russian and English in the tm package. Let's break this down clearly:
First: Why Your Initial Approach Won't Work
The stemDocument function in tm only accepts a single language string for the language parameter. Passing c("russian","english") will throw an error or default to just one language (usually the first in the vector), so it won't correctly stem both languages in your corpus.
The Correct Workflow for tm
To handle bilingual stemming, you need to split your corpus by language, stem each subset separately, then merge them back. Here's how to do it step-by-step:
1. Tag Each Document with Its Language
First, you need to identify which language each document is in. You can use the cld2 package for reliable language detection:
# Install and load the language detection package if you haven't already install.packages("cld2") library(cld2) library(tm) # Add a language metadata tag to each document in your corpus tw.corpus <- tm_map(tw.corpus, function(doc) { doc_content <- content(doc) detected_lang <- detect_language(doc_content) # Map detected language codes to tm's expected names (e.g., "ru" → "russian") lang_name <- ifelse(detected_lang == "ru", "russian", "english") meta(doc, "language") <- lang_name return(doc) })
2. Split the Corpus by Language
Separate your Russian and English documents into two subsets:
# Filter documents by their language metadata russian_corpus <- tw.corpus[meta(tw.corpus, "language") == "russian"] english_corpus <- tw.corpus[meta(tw.corpus, "language") == "english"]
3. Stem Each Subset with the Correct Language
Apply stemDocument to each subset using the appropriate language parameter:
# Stem Russian documents russian_stemmed <- tm_map(russian_corpus, stemDocument, language = "russian") # Stem English documents english_stemmed <- tm_map(english_corpus, stemDocument, language = "english")
4. Merge the Stemmed Corpora
Combine the two stemmed subsets back into a single corpus:
tw.corpus_stemmed <- c(russian_stemmed, english_stemmed)
A Simpler Alternative: Quanteda Package
If you're open to using a different text mining package, quanteda makes bilingual stemming much easier—it supports per-document language parameters. Here's a quick example:
install.packages("quanteda") library(quanteda) # Convert your tm corpus to a quanteda corpus quanteda_corpus <- corpus(tw.corpus) # Add language tags (same as before) docvars(quanteda_corpus, "language") <- sapply(texts(quanteda_corpus), function(x) { lang_code <- detect_language(x) ifelse(lang_code == "ru", "russian", "english") }) # Tokenize and stem each document with its detected language stemmed_tokens <- tokens_wordstem( tokens(quanteda_corpus), language = docvars(quanteda_corpus, "language") )
Note on Mixed-Language Documents
If your corpus has single documents containing both Russian and English text (like your example: "Hello, how are you Привет, как дела"), the above methods won't work perfectly. For this edge case, you'd need to split the text into individual words, detect the language of each word, then stem them accordingly. This requires a custom function, but it's doable with packages like cld2 and string manipulation.
内容的提问来源于stack exchange,提问作者psysky

