在R中使用ggplot绘制专利数据嵌套环形图的技术求助
用ggplot2实现专利数据的嵌套环形图(双层环)
以下是针对你提出的两种嵌套环形图需求的完整实现方案,同时解决布局控制和标签添加问题,先修正原数据处理中的统计错误:
library(tidyverse) # 加载并预处理数据 load(url("https://github.com/aterhorst/data/blob/master/patents.Rdata?raw=true")) class_breakdown <- lens_reformatted %>% select(class = value, country_code = `Country Code`, country = `Country Name`) %>% mutate( country_code = ifelse(is.na(country), "XX", country_code), country = ifelse(country_code == "XX", "Unknown", country) ) %>% # 统计每个类别-国家组合的专利数,直接用n()统计分组内的条目数 group_by(class, country) %>% summarise(patents = n(), .groups = "drop")
方案一:内层展示专利类别占比,外层按国家细分对应类别
实现逻辑
内层环展示每个专利类别的总数量占比,外层环将每个类别拆分为对应国家的专利数,通过计算每个分段的起始/中点位置实现标签定位:
# 1. 预处理内层(类别)数据,计算标签中点位置 inner_class_data <- class_breakdown %>% group_by(class) %>% summarise(total = sum(patents), .groups = "drop") %>% mutate( y_start = cumsum(lag(total, default = 0)), y_end = cumsum(total), y_mid = (y_start + y_end) / 2 # 类别标签的中心位置 ) # 2. 预处理外层(国家)数据,计算每个国家在对应类别中的位置 outer_class_data <- class_breakdown %>% left_join(inner_class_data %>% select(class, y_start), by = "class") %>% group_by(class) %>% mutate( country_start = y_start + cumsum(lag(patents, default = 0)), country_end = y_start + cumsum(patents), country_mid = (country_start + country_end) / 2 # 国家标签的中心位置 ) %>% ungroup() # 3. 绘制嵌套环形图 ggplot() + # 内层类别环:x=1控制环的半径,width控制环的宽度 geom_col( data = inner_class_data, aes(x = 1, y = total, fill = class), width = 0.8 ) + # 外层国家环:x=2控制外层半径,白色边框区分不同国家 geom_col( data = outer_class_data, aes(x = 2, y = patents, fill = class), width = 0.8, color = "white", size = 0.1 ) + # 添加内层类别标签:x=0.6控制标签在内环内侧 geom_text( data = inner_class_data, aes(x = 0.6, y = y_mid, label = class), size = 3, fontface = "bold" ) + # 添加外层国家标签:过滤专利数≥10的国家避免标签重叠,x=2.4控制标签在外环外侧 geom_text( data = outer_class_data %>% filter(patents >= 10), aes(x = 2.4, y = country_mid, label = country), size = 2.5, hjust = 0 ) + scale_fill_aaas() + coord_polar(theta = "y", start = 0) + # 控制x轴范围,避免标签超出画布 scale_x_continuous(limits = c(0, 3)) + theme_void() + theme(legend.position = "none")
方案二:内层展示国家专利占比,外层细分专利类别
实现逻辑
内层环展示每个国家的总专利数占比,外层环将每个国家拆分为对应类别的专利数,同样通过计算分段中点定位标签:
# 1. 预处理内层(国家)数据,过滤少量专利的国家避免标签拥挤 inner_country_data <- class_breakdown %>% group_by(country) %>% summarise(total = sum(patents), .groups = "drop") %>% mutate( y_start = cumsum(lag(total, default = 0)), y_end = cumsum(total), y_mid = (y_start + y_end) / 2 ) %>% filter(total >= 20) # 仅保留专利数≥20的国家 # 2. 预处理外层(类别)数据,计算每个类别在对应国家中的位置 outer_country_data <- class_breakdown %>% inner_join(inner_country_data %>% select(country, y_start), by = "country") %>% group_by(country) %>% mutate( class_start = y_start + cumsum(lag(patents, default = 0)), class_end = y_start + cumsum(patents), class_mid = (class_start + class_end) / 2 ) %>% ungroup() # 3. 绘制嵌套环形图 ggplot() + # 内层国家环 geom_col( data = inner_country_data, aes(x = 1, y = total, fill = country), width = 0.8 ) + # 外层类别环 geom_col( data = outer_country_data, aes(x = 2, y = patents, fill = class), width = 0.8, color = "white", size = 0.1 ) + # 添加内层国家标签 geom_text( data = inner_country_data, aes(x = 0.6, y = y_mid, label = country), size = 3, fontface = "bold" ) + # 添加外层类别标签:过滤专利数≥5的条目 geom_text( data = outer_country_data %>% filter(patents >= 5), aes(x = 2.4, y = class_mid, label = class), size = 2.5, hjust = 0 ) + scale_fill_viridis_d(option = "plasma") + coord_polar(theta = "y", start = 0) + scale_x_continuous(limits = c(0, 3)) + theme_void() + theme(legend.position = "bottom")
关键说明
- 环的半径由
geom_col的x参数控制(如x=1、x=2),环的宽度由width参数调整 - 标签定位的核心是通过累加前序分段的数值,计算每个分段的中点位置,确保标签居中显示
- 由于国家数量多达119个,必须通过
filter过滤少量专利的条目,否则标签会严重重叠,可根据需求调整阈值 - 原代码中
sum(n())是错误用法:group_by后直接用n()即可统计每组的专利数量
内容的提问来源于stack exchange,提问作者aterhorst
相关产品推荐
相关产品推荐

