使用R语言rvest包提取PGA官网球员图片时返回空白问题求助
问题原因与修复方案
核心问题说明
- PGA官网
/players.html为客户端动态渲染页面,rvest::read_html仅能获取初始静态HTML,此时球员相关数据还未通过接口加载、渲染到页面上,你定位的player-card类节点在静态源码中不存在,所以返回空值。 - 你使用的
player-card、player-image-wrapper都是旧版PGA官网的节点类名,当前官网迭代后页面结构已经更新,类名不匹配。 - 即便能定位到图片节点,当前PGA官网的图片都做了懒加载优化,初始状态下
src属性为占位空值,真实图片地址存储在data-src等自定义属性中。
可行解决方法
方法1:调用官方公开数据接口(推荐,稳定性更高)
PGA Tour公开了结构化的球员数据接口,无需解析前端页面即可直接拿到球员头像地址,示例代码如下:
if(!require(pacman))install.packages("pacman") pacman::p_load('httr', 'jsonlite') # 配置请求头避免被反爬拦截 req_headers <- c( "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" ) # 请求公开球员列表接口 res <- GET("https://statdata.pgatour.com/players/v1/players.json", add_headers(req_headers)) players_data <- fromJSON(rawToChar(res$content)) # 拼接球员头像地址,PGA头像地址为固定前缀+球员ID的格式 avatar_prefix <- "https://pga-tour-res.cloudinary.com/image/upload/c_fill,g_face,h_240,q_auto,w_180/v1/headshots/" plyers_photo <- paste0(avatar_prefix, players_data$players$id, ".png")
方法2:动态渲染页面后解析
如果需要直接解析前端页面,可以使用chromote等工具调用浏览器渲染完整页面后再提取节点:
if(!require(pacman))install.packages("pacman") pacman::p_load('chromote', 'rvest') # 启动Chrome无头实例 b <- ChromoteSession$new(headless = TRUE) # 访问页面等待渲染完成 b$Page$navigate("https://www.pgatour.com/players.html") b$Page$loadEventFired(timeout = 10000) # 获取渲染后的完整HTML源码 full_html <- read_html(b$Runtime$evaluate("document.documentElement.outerHTML")$result$value) # 关闭浏览器实例 b$close() # 按最新页面结构提取懒加载图片地址 plyers_photo <- full_html %>% html_nodes("img[data-testid='player-headshot']") %>% html_attr("data-src")
内容的提问来源于stack exchange,提问作者user86907
相关产品推荐
相关产品推荐

