使用R语言按年份分组统计字符串中的热门词汇需求
Hey Ian, great question! When dealing with millions of records to extract yearly trending terms from company names, efficiency is key—let’s walk through a practical, scalable solution using R’s text processing and data manipulation tools.
1. 准备工作:加载必要的包
First, we’ll use packages optimized for big data and text handling. data.table blazes through large datasets, tidytext simplifies text tokenization, and lubridate makes date processing a breeze:
install.packages(c("data.table", "tidytext", "lubridate", "dplyr")) library(data.table) library(tidytext) library(lubridate) library(dplyr)
2. 读取与预处理数据
Assuming your data is in a CSV (or similar flat file), use fread() for fast reading—way more efficient than base R’s read.csv() for large datasets:
# 读取数据,替换成你的文件路径 company_data <- fread("your_company_data.csv", col.names = c("ID", "IncorporationDate", "CompanyName")) # 提取年份:从日期列中解析出年份 company_data[, year := year(ymd(IncorporationDate))] # 清理公司名称:转小写、移除常见后缀/标点、过滤空字符串 # 自定义停用词:比如移除"LIMITED", "THE", "LTD"这类无意义的通用词 custom_stop_words <- tibble(word = c("limited", "the", "ltd", "inc", "group", "services")) company_data[, cleaned_name := tolower(CompanyName)] company_data[, cleaned_name := gsub("[^a-z\\s]", "", cleaned_name)] # 移除非字母和空格的字符
3. 拆分词汇并按年份统计词频
We’ll use unnest_tokens() to split each company name into individual words, then group by year to count occurrences:
# 拆分词汇并过滤停用词 word_counts <- company_data %>% unnest_tokens(word, cleaned_name) %>% anti_join(custom_stop_words) %>% group_by(year, word) %>% summarise(count = n(), .groups = "drop") %>% arrange(year, desc(count))
Note: For extremely large datasets (10M+ rows), using data.table syntax might be even faster. Here’s the equivalent:
word_counts_dt <- company_data[, .(word = unlist(strsplit(cleaned_name, "\\s+"))), by = .(year, ID)] %>% .[!word %in% custom_stop_words$word] %>% .[, .(count = .N), by = .(year, word)] %>% .[order(year, -count)]
4. 提取年度热门词汇
Now, pull the top N terms per year (e.g., top 10):
# 用dplyr提取每年Top10热门词汇 top_terms <- word_counts %>% group_by(year) %>% slice_max(n = 10, order_by = count) %>% ungroup() # 查看结果 print(top_terms)
可选:可视化热门词汇
If you want to visualize the trends, use ggplot2:
install.packages("ggplot2") library(ggplot2) ggplot(top_terms, aes(x = reorder(word, count), y = count, fill = factor(year))) + geom_col(show.legend = FALSE) + facet_wrap(~year, scales = "free_y") + coord_flip() + labs(x = "Term", y = "Frequency", title = "Top Company Name Terms by Year") + theme_minimal()
关键注意事项
- Memory management: For datasets with millions of rows, avoid loading the entire dataset into memory if possible.
data.tableuses memory efficiently, but you can also process in chunks if needed. - Custom stop words: Adjust the
custom_stop_wordslist based on your data—you might find other overused terms like "international", "consultants" that you want to exclude. - Tokenization tweaks: If you need to handle hyphenated terms or special cases, modify the
unnest_tokens()orstrsplit()parameters (e.g.,token = "regex", pattern = "\\s|-").
内容的提问来源于stack exchange,提问作者Ian

