使用R语言rvest库爬取Trustpilot时无法获取最后页码的问题
解决Trustpilot评论页最后页码爬取结果为0的问题
问题原因
- 直接使用
read_html()未添加请求头,可能被网站识别为非浏览器请求,返回不完整的静态页面内容。 - 原教程中的分页元素类名
.pagination-page已失效,当前Trustpilot的分页组件使用了新的类名结构。
解决方案
使用httr模拟浏览器请求,并通过定位"最后一页"按钮的链接来提取页码,这种方式比依赖分页按钮列表更可靠:
# 加载所需包 library(rvest) library(httr) library(stringr) # 目标URL url <- "https://www.trustpilot.com/review/www.ikea.com" # 模拟浏览器发送请求,添加User-Agent避免被拦截 response <- GET( url, user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") ) first_page <- content(response, as = "parsed") # 定义获取最后页码的函数 get_last_page <- function(html) { # 定位"最后一页"按钮的链接 last_page_href <- html %>% html_nodes('a[aria-label="Go to last page"]') %>% html_attr("href") # 从链接中提取page参数的数值 str_extract(last_page_href, "page=(\\d+)") %>% str_remove("page=") %>% as.numeric() } # 执行并打印结果 latest_page_number <- get_last_page(first_page) print(latest_page_number)
说明
- 添加
User-Agent头可以让请求更接近真实浏览器行为,避免网站返回不完整内容。 - 通过
aria-label="Go to last page"定位最后一页按钮,无需依赖易变的CSS类名,稳定性更高。
内容的提问来源于stack exchange,提问作者Millie Nguyen
相关产品推荐
相关产品推荐

