如何用R将语料库中复数文本转为单数?tm包无适配函数
Hey there! Let's get your plural-to-singular text handling working with the tm package in R. I see you wrote a custom function but are stuck applying it to your corpus—let's break down what's going wrong and how to fix it.
Option 1: Convert Plurals to Singular During Text Preprocessing
The tm package uses tm_map() to apply transformations to corpus documents, but your current function is built to handle word frequency vectors—not raw text documents. Let's adjust things to process the text directly first:
- First, write a function that takes a single text string, converts plurals to singulars, and returns the modified text. We'll handle common plural rules plus some irregular plurals as an example:
plural_to_singular <- function(text) { # Split text into individual words words <- strsplit(text, "\\s+")[[1]] # Handle regular plurals: remove trailing "s" or "es" words <- sub("es$", "", words) words <- sub("s$", "", words) # Handle irregular plurals (add more mappings as needed!) irregular_plurals <- c( "mice" = "mouse", "children" = "child", "feet" = "foot", "geese" = "goose" ) words <- ifelse(words %in% names(irregular_plurals), irregular_plurals[words], words) # Put the words back into a single string paste(words, collapse = " ") }
- Wrap this function with
content_transformer()sotmrecognizes it as a valid text transformation, then apply it to your corpus:
library(tm) # Example corpus (replace with your own) my_corpus <- VCorpus(VectorSource(c( "The cats are chasing mice across the fields", "Cats and dogs love eating treats" ))) # Standard preprocessing steps first my_corpus <- tm_map(my_corpus, content_transformer(tolower)) my_corpus <- tm_map(my_corpus, removePunctuation) my_corpus <- tm_map(my_corpus, removeWords, stopwords("english")) # Apply the plural-to-singular transformation my_corpus <- tm_map(my_corpus, content_transformer(plural_to_singular))
- Now when you generate a
TermDocumentMatrixorDocumentTermMatrix, all terms will already be in singular form:
tdm <- TermDocumentMatrix(my_corpus) inspect(tdm)
Option 2: Merge Plural/Singular Frequencies in an Existing Term Matrix
If you already have a term frequency matrix and want to merge plural counts into their singular counterparts, let's fix your original function first (note the case-sensitive sep instead of Sep—that was a critical bug!) and apply it correctly:
- Fix the custom aggregation function:
aggregate.plurals <- function(term_vector) { aggro_fen <- function(v, singular, plural) { if (!is.na(v[plural])) { v[singular] <- v[singular] + v[plural] v <- v[-which(names(v) == plural)] } return(v) } # Work with a copy to avoid modifying the original vector mid-loop processed_v <- term_vector original_terms <- names(processed_v) for (term in original_terms) { # Handle "s" plurals plural_s <- paste(term, "s", sep = "") if (plural_s %in% names(processed_v)) { processed_v <- aggro_fen(processed_v, term, plural_s) } # Handle "es" plurals plural_es <- paste(term, "es", sep = "") if (plural_es %in% names(processed_v)) { processed_v <- aggro_fen(processed_v, term, plural_es) } } return(processed_v) }
- Apply this function to your term matrix. First convert it to a regular matrix, process each document's word frequencies, then convert back:
# Convert your TermDocumentMatrix to a regular matrix tdm_matrix <- as.matrix(tdm) # Apply the aggregation to each document (column) aggregated_matrix <- apply(tdm_matrix, 2, aggregate.plurals) # Convert back to a TermDocumentMatrix aggregated_tdm <- as.TermDocumentMatrix(aggregated_matrix, weighting = weightTf) inspect(aggregated_tdm)
Key Notes:
- R is case-sensitive! Your original
Sep='s'was causingpaste()to add spaces instead of concatenating correctly—always use lowercasesep. - Irregular plurals (like
mice→mouse) need explicit mappings, since there's no universal rule for them.
内容的提问来源于stack exchange,提问作者Ayush

