如何高效统计DataFrame中出现频率最高的单词?
Hey there! Great question—let's break this down so you can get the full word frequency stats without worrying about efficiency, even with 20k+ rows.
First off: yes, you do need to combine all the word vectors from the Message column into a single collection of words to count frequencies across the entire dataset. But don't stress about the 20k rows—R handles this kind of workload really well, even with basic methods.
Method 1: Base R (simple and fast)
The easiest way in base R is to use unlist() to flatten all the vectors in df$Message into one long vector of words, then run your table() and sort() combo on that:
# Flatten all word vectors into one single vector all_words <- unlist(df$Message) # Calculate and sort frequencies sorted_freq <- sort(table(all_words), decreasing = TRUE)
unlist() is super efficient, and even with 20k rows (say, average 10 words per row = 200k total words), this will run in a fraction of a second.
Method 2: Tidyverse (cleaner output for further analysis)
If you prefer working with tidy data structures (which makes downstream tasks easier), use unnest_longer() to expand the Message column into one row per word, then count() to get frequencies:
library(tidyverse) word_freq <- df %>% unnest_longer(Message) %>% count(Message, sort = TRUE)
This gives you a DataFrame with two columns (Message for the word, n for the count) that's ready for plotting, filtering, or any other analysis you need. Again, this method is more than fast enough for 20k rows.
Why efficiency isn't a problem here
20k rows is a tiny dataset by R's standards. Even if every row had 50 words, that's only 1 million total words—operations like unlist() or unnest_longer() are optimized to handle this without breaking a sweat. You won't notice any lag at all.
Just a quick side note: if you're generating this DataFrame from a raw text column using strsplit(), you could even skip storing the list-column and directly generate the flattened word list or tidy DataFrame, but either way, the methods above will work perfectly with your existing df.
内容的提问来源于stack exchange,提问作者Arnaud Stephan

