You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中对俄英双语语料库使用双词干提取的方法咨询

Bilingual Stemming for Russian & English in the tm Package

Hey there! Great question about handling bilingual stemming for Russian and English in the tm package. Let's break this down clearly:

First: Why Your Initial Approach Won't Work

The stemDocument function in tm only accepts a single language string for the language parameter. Passing c("russian","english") will throw an error or default to just one language (usually the first in the vector), so it won't correctly stem both languages in your corpus.

The Correct Workflow for tm

To handle bilingual stemming, you need to split your corpus by language, stem each subset separately, then merge them back. Here's how to do it step-by-step:

1. Tag Each Document with Its Language

First, you need to identify which language each document is in. You can use the cld2 package for reliable language detection:

# Install and load the language detection package if you haven't already
install.packages("cld2")
library(cld2)
library(tm)

# Add a language metadata tag to each document in your corpus
tw.corpus <- tm_map(tw.corpus, function(doc) {
  doc_content <- content(doc)
  detected_lang <- detect_language(doc_content)
  # Map detected language codes to tm's expected names (e.g., "ru" → "russian")
  lang_name <- ifelse(detected_lang == "ru", "russian", "english")
  meta(doc, "language") <- lang_name
  return(doc)
})

2. Split the Corpus by Language

Separate your Russian and English documents into two subsets:

# Filter documents by their language metadata
russian_corpus <- tw.corpus[meta(tw.corpus, "language") == "russian"]
english_corpus <- tw.corpus[meta(tw.corpus, "language") == "english"]

3. Stem Each Subset with the Correct Language

Apply stemDocument to each subset using the appropriate language parameter:

# Stem Russian documents
russian_stemmed <- tm_map(russian_corpus, stemDocument, language = "russian")
# Stem English documents
english_stemmed <- tm_map(english_corpus, stemDocument, language = "english")

4. Merge the Stemmed Corpora

Combine the two stemmed subsets back into a single corpus:

tw.corpus_stemmed <- c(russian_stemmed, english_stemmed)

A Simpler Alternative: Quanteda Package

If you're open to using a different text mining package, quanteda makes bilingual stemming much easier—it supports per-document language parameters. Here's a quick example:

install.packages("quanteda")
library(quanteda)

# Convert your tm corpus to a quanteda corpus
quanteda_corpus <- corpus(tw.corpus)

# Add language tags (same as before)
docvars(quanteda_corpus, "language") <- sapply(texts(quanteda_corpus), function(x) {
  lang_code <- detect_language(x)
  ifelse(lang_code == "ru", "russian", "english")
})

# Tokenize and stem each document with its detected language
stemmed_tokens <- tokens_wordstem(
  tokens(quanteda_corpus),
  language = docvars(quanteda_corpus, "language")
)

Note on Mixed-Language Documents

If your corpus has single documents containing both Russian and English text (like your example: "Hello, how are you Привет, как дела"), the above methods won't work perfectly. For this edge case, you'd need to split the text into individual words, detect the language of each word, then stem them accordingly. This requires a custom function, but it's doable with packages like cld2 and string manipulation.

内容的提问来源于stack exchange,提问作者psysky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:24:03