如何仅在removeWords函数内部调用tolower,保留原字符向量单词大小写?
Got it, let's fix this! The issue with your current approach is that converting the entire words vector to lowercase first strips away all original capitalization—hence why "Intelligent" gets turned into "intelligent". We need to match words case-insensitively for removal, but keep the original casing of the words we retain. Here are a couple of straightforward solutions:
Solution 1: Custom Function (Leveraging tm Package Logic)
Instead of modifying the entire input vector, we'll only use tolower() during the matching step to check which words should be removed, while leaving the original words untouched.
First, set up your variables as before:
library(tm) words <- c("the", "The", "Intelligent", "this", "This") words_to_remove <- c("the", "This")
Then define a custom case-insensitive removal function:
removeWordsCaseInsensitive <- function(text, target_words) { # Convert target words to lowercase for case-insensitive matching target_lower <- tolower(unique(target_words)) # Keep only words whose lowercase version isn't in the target list text[!tolower(text) %in% target_lower] }
Call the function to get your desired result:
cleaned_words <- removeWordsCaseInsensitive(words, words_to_remove) print(cleaned_words) # Output: "Intelligent" (preserves original capitalization!)
Solution 2: Regex-Based Filtering
If you prefer a more concise approach without writing a custom function, you can use regular expressions to match entire words case-insensitively, then filter the original vector:
# Create a regex pattern for exact, case-insensitive matches removal_pattern <- paste0("^", tolower(words_to_remove), "$", collapse = "|") # Filter words that don't match the pattern (keep original casing) cleaned_words <- words[!grepl(removal_pattern, tolower(words))]
Why Your Original Approach Didn't Work
Your initial code removeWords(tolower(words), tolower(words_to_remove)) converts the entire input vector to lowercase before removal, so all output words are forced to lowercase. The key here is separating the matching logic (where we ignore case) from the output preservation (where we keep the original text).
内容的提问来源于stack exchange,提问作者WinterMensch

