基于R从长数据中按条件提取指定字符串并计算占比及可视化
解决方案
针对你的需求,我们可以用tidyverse工具链完成数据处理与可视化,步骤如下:
1. 加载依赖包
先加载处理数据和绘图所需的包:
library(tidyverse)
2. 数据处理:提取关键词并计算占比
假设你的长格式数据框为corpus_data,包含tag(语料库类别)和text(单条文本内容)两列。
场景1:按文本条目统计占比(包含目标词的文本数/该语料库总文本数)
先筛选出包含"apple"或"melon"的文本,标记对应的关键词,再结合各语料库的总文本数计算占比:
# 筛选并标记关键词 keyword_records <- corpus_data %>% mutate(keyword = case_when( str_detect(text, fixed("apple")) ~ "apple", str_detect(text, fixed("melon")) ~ "melon", TRUE ~ NA_character_ )) %>% drop_na(keyword) # 计算占比 ratio_df <- keyword_records %>% count(tag, keyword, name = "match_count") %>% left_join( corpus_data %>% count(tag, name = "total_records"), by = "tag" ) %>% mutate(ratio = match_count / total_records)
场景2:按词频统计占比(目标词出现总次数/该语料库总词数)
如果需要统计词级别的占比,先拆分文本为单个词语,再计算占比:
# 拆分文本并统计词频 word_freq_df <- corpus_data %>% unnest_tokens(word, text) %>% mutate(keyword = case_when( word == "apple" ~ "apple", word == "melon" ~ "melon", TRUE ~ NA_character_ )) %>% drop_na(keyword) # 计算词频占比 word_ratio_df <- word_freq_df %>% count(tag, keyword, name = "keyword_freq") %>% left_join( corpus_data %>% unnest_tokens(word, text) %>% count(tag, name = "total_words"), by = "tag" ) %>% mutate(ratio = keyword_freq / total_words)
3. 绘制占比对比条形图
用ggplot2绘制分组条形图,直观对比两个关键词在各语料库的占比:
# 用场景1的ratio_df,场景2替换为word_ratio_df即可 ggplot(ratio_df, aes(x = tag, y = ratio, fill = keyword)) + geom_col(position = position_dodge(width = 0.8), width = 0.7) + # 添加百分比标签 geom_text(aes(label = scales::percent(ratio, accuracy = 0.1)), position = position_dodge(width = 0.8), vjust = -0.3, size = 3.5) + # y轴设为百分比格式 scale_y_continuous(labels = scales::percent, limits = c(0, max(ratio_df$ratio)*1.1)) + labs(title = "Apple与Melon在各语料库中的占比对比", x = "语料库类别", y = "占比", fill = "关键词") + theme_minimal()
这段代码会生成分组条形图,每个语料库下对应两个条形,分别展示"apple"和"melon"的占比,同时在条形上方标注具体百分比,清晰完成对比需求。
内容的提问来源于stack exchange,提问作者pindakazen
相关产品推荐
相关产品推荐

