You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效统计DataFrame中出现频率最高的单词?

高效统计DataFrame中所有单词的频率

Hey there! Great question—let's break this down so you can get the full word frequency stats without worrying about efficiency, even with 20k+ rows.

First off: yes, you do need to combine all the word vectors from the Message column into a single collection of words to count frequencies across the entire dataset. But don't stress about the 20k rows—R handles this kind of workload really well, even with basic methods.

Method 1: Base R (simple and fast)

The easiest way in base R is to use unlist() to flatten all the vectors in df$Message into one long vector of words, then run your table() and sort() combo on that:

# Flatten all word vectors into one single vector
all_words <- unlist(df$Message)

# Calculate and sort frequencies
sorted_freq <- sort(table(all_words), decreasing = TRUE)

unlist() is super efficient, and even with 20k rows (say, average 10 words per row = 200k total words), this will run in a fraction of a second.

Method 2: Tidyverse (cleaner output for further analysis)

If you prefer working with tidy data structures (which makes downstream tasks easier), use unnest_longer() to expand the Message column into one row per word, then count() to get frequencies:

library(tidyverse)

word_freq <- df %>%
  unnest_longer(Message) %>%
  count(Message, sort = TRUE)

This gives you a DataFrame with two columns (Message for the word, n for the count) that's ready for plotting, filtering, or any other analysis you need. Again, this method is more than fast enough for 20k rows.

Why efficiency isn't a problem here

20k rows is a tiny dataset by R's standards. Even if every row had 50 words, that's only 1 million total words—operations like unlist() or unnest_longer() are optimized to handle this without breaking a sweat. You won't notice any lag at all.

Just a quick side note: if you're generating this DataFrame from a raw text column using strsplit(), you could even skip storing the list-column and directly generate the flattened word list or tidy DataFrame, but either way, the methods above will work perfectly with your existing df.

内容的提问来源于stack exchange,提问作者Arnaud Stephan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:18:35