使用tidytext处理邮件文本遇错误:Can't convert a function to a quosure
Hey there! Let's get that word frequency count working for your email text file. The error you're hitting comes from using an outdated function in tidytext—here's how to fix it and get things running smoothly:
Why the Error Happens
You're using unnest_tokens_() (with the underscore), which is the old standard evaluation version of the function. This function has been deprecated in recent tidytext releases, and the modern tidyverse ecosystem now favors non-standard evaluation functions like unnest_tokens() (no underscore). Trying to use the old function with current tidytext versions leads to the quosure conversion error you're seeing.
Corrected Code
Here's the updated script that will work with modern tidytext:
library(tidytext) library(tibble) # Read in your email data data <- read_lines("emails.txt") text_tibble <- tibble(line = seq_along(data), text = data) # Tokenize the text (the corrected line!) text_words <- text_tibble %>% unnest_tokens(word, text) # Calculate basic word frequencies word_frequencies <- text_words %>% count(word, sort = TRUE)
Bonus: Clean Up Invalid/Stop Words
Since you mentioned your data has lots of invalid content, you'll probably want to filter out common stop words (like "the", "and") to get meaningful results. Add this step after tokenization:
# Load built-in stop words data(stop_words) # Filter out stop words and recalculate frequencies clean_word_frequencies <- text_words %>% anti_join(stop_words) %>% count(word, sort = TRUE)
Quick Note on Package Version
If you're still having issues, make sure you're running the latest version of tidytext:
update.packages("tidytext")
内容的提问来源于stack exchange,提问作者user9365328

