R读取HTML页面时如何拆分单行大文本提取表格生成dataframe
解决方法
你之前的代码没有得到预期结果,是因为grep是对字符串向量的每个元素做匹配,你拿到的data是长度为1的字符串向量,所以只要包含<td>就只会返回1。你可以按以下两种方式处理:
方式1:基于你现有代码继续处理
首先把超长单行文本按表格行的结束标签拆分,再逐行提取单元格内容:
# 1. 按表格行结束标签拆分单行文本,得到每一行表格的原始内容 row_list <- strsplit(thepage[11], "</tr>")[[1]] # 过滤掉不包含<td>的无效行 valid_rows <- row_list[grep("<td>", row_list)] # 2. 定义单元格内容清洗函数 clean_cell <- function(cell_str) { # 移除HTML标签、千位分隔符、百分号,去除前后空格 res <- gsub("<.*?>|,|%", "", cell_str) res <- trimws(res) return(res) } # 3. 逐行提取清洗单元格内容 row_content <- lapply(valid_rows, function(x) { cells <- strsplit(x, "<td>")[[1]] cells <- cells[cells != ""] sapply(cells, clean_cell, USE.NAMES = FALSE) }) # 4. 提取表头并生成最终数据框 header_raw <- strsplit(thepage[grep("<th>", thepage)[1]], "</th>")[[1]] headers <- sapply(header_raw, clean_cell, USE.NAMES = FALSE) headers <- headers[headers != ""] pop_df <- as.data.frame(do.call(rbind, row_content), stringsAsFactors = FALSE) colnames(pop_df) <- headers # 把非国家名称的列转成数值类型 pop_df[, -2] <- lapply(pop_df[, -2], as.numeric)
方式2:用HTML解析包直接提取(更稳定,推荐)
手动拆分字符串容易因为页面结构变动出错,直接用rvest包解析HTML更简单:
# 安装加载包 install.packages("rvest") library(rvest) # 读取页面并提取表格 page <- read_html("https://www.worldometers.info/world-population/population-by-country/") pop_df <- page %>% html_element("table#example2") %>% html_table()
提取完成后你可以根据需求再做列名调整、数值格式清理即可。
内容的提问来源于stack exchange,提问作者user16596066
相关产品推荐
相关产品推荐

