R语言爬取维基联合国人口表返回空列表,求合并病例数据方案
解决维基百科人口数据抓取及与total_cases合并问题
问题根源分析
- 选择器错误:原代码用
html_nodes('.jquery-tablesorter td')选中表格单元格,而html_table()需要作用在<table>节点上,导致返回空列表。 - 表头处理不当:目标维基百科表格的第二行是说明性副标题,直接
slice(-1)会导致列名混乱。 - 数据格式未清洗:人口数值带千分位逗号,未转换为数值类型;年份列处理逻辑有误。
- 列名不匹配:合并时
total_cases用小写country/year,而抓取的数据集用大写Country/Year,会导致匹配失败。 - 国家名称不一致:WHO数据与维基百科的国家名称存在差异(如"US" vs "United States of America"),需要统一。
修正后的Part D代码
library(rvest) library(tidyverse) # URL of the Wikipedia page containing the population data url <- "https://en.wikipedia.org/wiki/List_of_countries_by_population_(United_Nations)" # 抓取表格:选择table节点而非td节点,提取第一个符合条件的表格 population_table <- url %>% read_html() %>% html_nodes('.jquery-tablesorter') %>% html_table(header = TRUE, fill = TRUE) %>% pluck(1) # 页面可能有多个同class表格,取第一个目标表格 # 清洗人口数据:跳过副标题行,选择需要的列,处理年份和人口数值 population_data <- population_table %>% slice(-1) %>% # 跳过第二行的副标题 select(`Country or area`, contains("Est.")) %>% # 选择国家列和各年份人口估算列 rename(Country = `Country or area`) %>% pivot_longer(-Country, names_to = "Year", values_to = "Population") %>% # 提取年份数字,清洗人口数值(去掉逗号和空格) mutate( Year = as.numeric(str_extract(Year, "\\d{4}")), Population = as.numeric(str_remove_all(Population, "[\\s,]")) ) %>% drop_na(Year, Population) # 移除无效数据 # 统一列名大小写,匹配total_cases的列名 population_data <- population_data %>% rename(country = Country, year = Year) # 处理国家名称匹配(示例:部分常见差异,可根据需要扩展) population_data <- population_data %>% mutate( country = case_when( country == "United States of America" ~ "United States", country == "United Kingdom of Great Britain and Northern Ireland" ~ "United Kingdom", country == "Republic of Korea" ~ "South Korea", country == "Iran (Islamic Republic of)" ~ "Iran", TRUE ~ country ) ) # 合并数据集 merged_data <- total_cases %>% left_join(population_data, by = c("country", "year")) # 查看结果 merged_data
关键修改说明
- 表格选择:改用
html_nodes('.jquery-tablesorter')选中表格节点,并用pluck(1)确保拿到目标表格。 - 表头与列选择:使用
header = TRUE正确识别表头,选择Country or area和含"Est."的年份列,过滤无关数据。 - 数据清洗:用
str_extract提取4位年份,str_remove_all清理人口数值中的逗号和空格,转换为数值类型。 - 列名统一:将
Country/Year重命名为country/year,确保合并时列名匹配。 - 国家名称对齐:通过
case_when修正常见的国家名称差异,提升合并匹配率。
内容的提问来源于stack exchange,提问作者khali-blinks
相关产品推荐
相关产品推荐

