使用rvest抓取选区列表:正确CSS选择器与数据整理求助
用rvest提取选区列表的CSS选择器及代码实现
目标页面的选区数据嵌套在div.col下的表格结构里,你可以通过以下CSS选择器精准定位所需元素:
- 选区名称:
div.col table tr td:first-child(匹配表格每行的第一列文本) - 截止日期:
div.col table tr td:last-child(匹配表格每行的最后一列文本)
完整代码示例
library(rvest) library(tibble) # 读取目标页面 url <- "https://labour.org.uk/activist-hub/governance-and-legal-hub/selections/parliamentary-candidate-application-form/" page <- read_html(url) # 提取并清洗数据(跳过表头行) rows <- page %>% html_elements("div.col table tr:not(:first-child)") constituencies <- rows %>% html_element("td:first-child") %>% html_text2() # 自动去除多余空格、换行 deadlines <- rows %>% html_element("td:last-child") %>% html_text2() # 整理为干净的tibble result_tibble <- tibble( 选区名称 = constituencies, 截止日期 = deadlines ) print(result_tibble)
补充说明
html_text2()比html_text()更适合处理网页文本,能自动清理冗余格式- 如果页面表格结构有变动,可通过浏览器开发者工具(F12)重新确认元素路径,调整CSS选择器
内容的提问来源于stack exchange,提问作者C.Robin
相关产品推荐
相关产品推荐

