You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言爬取维基联合国人口表返回空列表,求合并病例数据方案

解决维基百科人口数据抓取及与total_cases合并问题

问题根源分析

  • 选择器错误:原代码用html_nodes('.jquery-tablesorter td')选中表格单元格,而html_table()需要作用在<table>节点上,导致返回空列表。
  • 表头处理不当:目标维基百科表格的第二行是说明性副标题,直接slice(-1)会导致列名混乱。
  • 数据格式未清洗:人口数值带千分位逗号,未转换为数值类型;年份列处理逻辑有误。
  • 列名不匹配:合并时total_cases用小写country/year,而抓取的数据集用大写Country/Year,会导致匹配失败。
  • 国家名称不一致:WHO数据与维基百科的国家名称存在差异(如"US" vs "United States of America"),需要统一。

修正后的Part D代码

library(rvest)
library(tidyverse)

# URL of the Wikipedia page containing the population data
url <- "https://en.wikipedia.org/wiki/List_of_countries_by_population_(United_Nations)"

# 抓取表格:选择table节点而非td节点,提取第一个符合条件的表格
population_table <- url %>%
  read_html() %>%
  html_nodes('.jquery-tablesorter') %>%
  html_table(header = TRUE, fill = TRUE) %>%
  pluck(1) # 页面可能有多个同class表格,取第一个目标表格

# 清洗人口数据:跳过副标题行,选择需要的列,处理年份和人口数值
population_data <- population_table %>%
  slice(-1) %>% # 跳过第二行的副标题
  select(`Country or area`, contains("Est.")) %>% # 选择国家列和各年份人口估算列
  rename(Country = `Country or area`) %>%
  pivot_longer(-Country, names_to = "Year", values_to = "Population") %>%
  # 提取年份数字,清洗人口数值(去掉逗号和空格)
  mutate(
    Year = as.numeric(str_extract(Year, "\\d{4}")),
    Population = as.numeric(str_remove_all(Population, "[\\s,]"))
  ) %>%
  drop_na(Year, Population) # 移除无效数据

# 统一列名大小写,匹配total_cases的列名
population_data <- population_data %>%
  rename(country = Country, year = Year)

# 处理国家名称匹配(示例:部分常见差异,可根据需要扩展)
population_data <- population_data %>%
  mutate(
    country = case_when(
      country == "United States of America" ~ "United States",
      country == "United Kingdom of Great Britain and Northern Ireland" ~ "United Kingdom",
      country == "Republic of Korea" ~ "South Korea",
      country == "Iran (Islamic Republic of)" ~ "Iran",
      TRUE ~ country
    )
  )

# 合并数据集
merged_data <- total_cases %>%
  left_join(population_data, by = c("country", "year"))

# 查看结果
merged_data

关键修改说明

  1. 表格选择:改用html_nodes('.jquery-tablesorter')选中表格节点,并用pluck(1)确保拿到目标表格。
  2. 表头与列选择:使用header = TRUE正确识别表头,选择Country or area和含"Est."的年份列,过滤无关数据。
  3. 数据清洗:用str_extract提取4位年份,str_remove_all清理人口数值中的逗号和空格,转换为数值类型。
  4. 列名统一:将Country/Year重命名为country/year,确保合并时列名匹配。
  5. 国家名称对齐:通过case_when修正常见的国家名称差异,提升合并匹配率。

内容的提问来源于stack exchange,提问作者khali-blinks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 16:27:55