You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言按年份分组统计字符串中的热门词汇需求

R语言实现年度公司名称热门词汇统计

Hey Ian, great question! When dealing with millions of records to extract yearly trending terms from company names, efficiency is key—let’s walk through a practical, scalable solution using R’s text processing and data manipulation tools.

1. 准备工作:加载必要的包

First, we’ll use packages optimized for big data and text handling. data.table blazes through large datasets, tidytext simplifies text tokenization, and lubridate makes date processing a breeze:

install.packages(c("data.table", "tidytext", "lubridate", "dplyr"))
library(data.table)
library(tidytext)
library(lubridate)
library(dplyr)

2. 读取与预处理数据

Assuming your data is in a CSV (or similar flat file), use fread() for fast reading—way more efficient than base R’s read.csv() for large datasets:

# 读取数据,替换成你的文件路径
company_data <- fread("your_company_data.csv", col.names = c("ID", "IncorporationDate", "CompanyName"))

# 提取年份:从日期列中解析出年份
company_data[, year := year(ymd(IncorporationDate))]

# 清理公司名称:转小写、移除常见后缀/标点、过滤空字符串
# 自定义停用词:比如移除"LIMITED", "THE", "LTD"这类无意义的通用词
custom_stop_words <- tibble(word = c("limited", "the", "ltd", "inc", "group", "services"))

company_data[, cleaned_name := tolower(CompanyName)]
company_data[, cleaned_name := gsub("[^a-z\\s]", "", cleaned_name)]  # 移除非字母和空格的字符

3. 拆分词汇并按年份统计词频

We’ll use unnest_tokens() to split each company name into individual words, then group by year to count occurrences:

# 拆分词汇并过滤停用词
word_counts <- company_data %>%
  unnest_tokens(word, cleaned_name) %>%
  anti_join(custom_stop_words) %>%
  group_by(year, word) %>%
  summarise(count = n(), .groups = "drop") %>%
  arrange(year, desc(count))

Note: For extremely large datasets (10M+ rows), using data.table syntax might be even faster. Here’s the equivalent:

word_counts_dt <- company_data[, .(word = unlist(strsplit(cleaned_name, "\\s+"))), by = .(year, ID)] %>%
  .[!word %in% custom_stop_words$word] %>%
  .[, .(count = .N), by = .(year, word)] %>%
  .[order(year, -count)]

4. 提取年度热门词汇

Now, pull the top N terms per year (e.g., top 10):

# 用dplyr提取每年Top10热门词汇
top_terms <- word_counts %>%
  group_by(year) %>%
  slice_max(n = 10, order_by = count) %>%
  ungroup()

# 查看结果
print(top_terms)

可选:可视化热门词汇

If you want to visualize the trends, use ggplot2:

install.packages("ggplot2")
library(ggplot2)

ggplot(top_terms, aes(x = reorder(word, count), y = count, fill = factor(year))) +
  geom_col(show.legend = FALSE) +
  facet_wrap(~year, scales = "free_y") +
  coord_flip() +
  labs(x = "Term", y = "Frequency", title = "Top Company Name Terms by Year") +
  theme_minimal()

关键注意事项

  • Memory management: For datasets with millions of rows, avoid loading the entire dataset into memory if possible. data.table uses memory efficiently, but you can also process in chunks if needed.
  • Custom stop words: Adjust the custom_stop_words list based on your data—you might find other overused terms like "international", "consultants" that you want to exclude.
  • Tokenization tweaks: If you need to handle hyphenated terms or special cases, modify the unnest_tokens() or strsplit() parameters (e.g., token = "regex", pattern = "\\s|-").

内容的提问来源于stack exchange,提问作者Ian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:13:00