如何用R抓取含交互元素网页中同CSS类的两个表格数据
解决思路与代码实现
因为两个切换表格的CSS类完全相同,html_element()只会返回匹配到的第一个表格,改用html_elements()获取所有符合条件的表格,再按索引提取目标表格即可解决问题。
完整代码实现
library(rvest) library(tidyverse) # 读取目标网页 mammoth <- read_html('https://www.mammothmountain.com/on-the-mountain/historical-snowfall') # 获取所有匹配指定CSS类的表格(共2个) all_tables <- mammoth %>% html_elements('table.css-86hwhl') %>% map(~html_table(.x, header = TRUE, convert = TRUE)) # 提取第一个表格:扩展降雪历史数据 extended_snowfall <- all_tables[[1]] %>% mutate_if(is.character, as.factor) %>% mutate_if(is.integer, as.double) %>% select(-Total) # 提取第二个表格:2022-2023冬季季数据 winter_2022_2023 <- all_tables[[2]] %>% mutate_if(is.character, as.factor) %>% mutate_if(is.integer, as.double) # ---------------------- # 这里添加对冬季数据的处理计算逻辑(示例) # 假设需要将冬季数据汇总为一行,添加到扩展历史表格中 winter_summary <- winter_2022_2023 %>% summarise( Season = "2022-2023", across(where(is.numeric), sum, na.rm = TRUE) ) %>% mutate_if(is.character, as.factor) # 合并两个表格 combined_snowfall <- bind_rows(extended_snowfall, winter_summary)
关键说明
html_elements()会返回页面中所有匹配指定CSS选择器的元素,此处得到包含两个目标表格的列表。- 通过
[[1]]和[[2]]分别提取两个表格,对应页面上的“extended snowfall history”和“2022-2023 winter season”。 - 可根据实际需求调整冬季数据的处理逻辑,最终用
bind_rows()将处理后的冬季数据添加到扩展降雪历史表格中。
内容的提问来源于stack exchange,提问作者agf1997
相关产品推荐
相关产品推荐

