You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R读取HTML页面时如何拆分单行大文本提取表格生成dataframe

解决方法

你之前的代码没有得到预期结果,是因为grep是对字符串向量的每个元素做匹配,你拿到的data是长度为1的字符串向量,所以只要包含<td>就只会返回1。你可以按以下两种方式处理:

方式1:基于你现有代码继续处理

首先把超长单行文本按表格行的结束标签拆分,再逐行提取单元格内容:

# 1. 按表格行结束标签拆分单行文本,得到每一行表格的原始内容
row_list <- strsplit(thepage[11], "</tr>")[[1]]
# 过滤掉不包含<td>的无效行
valid_rows <- row_list[grep("<td>", row_list)]

# 2. 定义单元格内容清洗函数
clean_cell <- function(cell_str) {
  # 移除HTML标签、千位分隔符、百分号,去除前后空格
  res <- gsub("<.*?>|,|%", "", cell_str)
  res <- trimws(res)
  return(res)
}

# 3. 逐行提取清洗单元格内容
row_content <- lapply(valid_rows, function(x) {
  cells <- strsplit(x, "<td>")[[1]]
  cells <- cells[cells != ""]
  sapply(cells, clean_cell, USE.NAMES = FALSE)
})

# 4. 提取表头并生成最终数据框
header_raw <- strsplit(thepage[grep("<th>", thepage)[1]], "</th>")[[1]]
headers <- sapply(header_raw, clean_cell, USE.NAMES = FALSE)
headers <- headers[headers != ""]

pop_df <- as.data.frame(do.call(rbind, row_content), stringsAsFactors = FALSE)
colnames(pop_df) <- headers
# 把非国家名称的列转成数值类型
pop_df[, -2] <- lapply(pop_df[, -2], as.numeric)

方式2:用HTML解析包直接提取(更稳定,推荐)

手动拆分字符串容易因为页面结构变动出错,直接用rvest包解析HTML更简单:

# 安装加载包
install.packages("rvest")
library(rvest)

# 读取页面并提取表格
page <- read_html("https://www.worldometers.info/world-population/population-by-country/")
pop_df <- page %>% 
  html_element("table#example2") %>% 
  html_table()

提取完成后你可以根据需求再做列名调整、数值格式清理即可。

内容的提问来源于stack exchange,提问作者user16596066

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 22:45:03