Yahoo Finance API移除后,如何用R爬取关键统计页面指定数据?
解决Yahoo Finance关键统计页面爬取问题(R语言)
问题原因
Yahoo Finance的关键统计页面大量内容通过JavaScript动态渲染,直接用rvest::read_html()只能获取静态HTML结构,无法加载JS生成的实际数据,因此返回空结果。
方案1:提取页面内嵌JSON数据(轻量高效)
Yahoo会将页面数据以JSON格式内嵌在<script type="application/json">标签中,无需模拟浏览器即可提取:
代码实现
library(httr) library(rvest) library(jsonlite) library(dplyr) # 目标股票的关键统计页面URL url <- "https://uk.finance.yahoo.com/quote/IPX.L/key-statistics?p=IPX.L" # 模拟浏览器请求,避免被拦截 page <- GET( url, user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") ) html_content <- content(page, "text", encoding = "UTF-8") # 提取页面中的内嵌JSON数据块 script_tags <- html_content %>% read_html() %>% html_elements("script[type='application/json']") %>% html_text() # 筛选包含关键统计数据的JSON片段 target_json <- script_tags[sapply(script_tags, function(x) grepl("trailingPE|bookValue", x))][1] data_list <- fromJSON(target_json) # 提取指定数据 ## Trailing P/E trailing_pe <- data_list$context$dispatcher$stores$QuoteSummaryStore$summaryDetail$trailingPE$raw ## Book value per share (mrq) book_value_per_share <- data_list$context$dispatcher$stores$QuoteSummaryStore$summaryDetail$bookValue$raw ## 最新股价及时间 latest_price <- data_list$context$dispatcher$stores$QuoteSummaryStore$price$regularMarketPrice$raw price_time <- data_list$context$dispatcher$stores$QuoteSummaryStore$price$regularMarketTime %>% as.POSIXct(origin = "1970-01-01", tz = "Europe/London") %>% format("%H:%M %p BST") # 输出结果 cat("Trailing P/E:", trailing_pe, "\n") cat("Book value per share (mrq):", book_value_per_share, "\n") cat("Latest price and time:", latest_price, "At close :", price_time, "\n")
通用方案:将所有关键统计导出为DataFrame
# 提取所有关键统计数据 key_stats <- data_list$context$dispatcher$stores$QuoteSummaryStore$defaultKeyStatistics # 转换为结构化DataFrame,处理空值 stats_df <- key_stats %>% lapply(function(x) ifelse(is.null(x$raw), NA, x$raw)) %>% as.data.frame() %>% t() %>% as.data.frame() %>% rownames_to_column(var = "Statistic") %>% rename(Value = V1) # 优化统计名称格式(可选) stats_df$Statistic <- gsub("_", " ", stats_df$Statistic) stats_df$Statistic <- stringr::str_to_title(stats_df$Statistic) # 查看结果 head(stats_df)
方案2:用RSelenium模拟浏览器加载(兼容动态渲染)
如果内嵌JSON结构发生变化,可通过模拟浏览器完全加载页面:
代码实现
library(RSelenium) library(rvest) library(dplyr) library(stringr) # 启动Chrome浏览器驱动(需提前安装Chrome和对应版本的ChromeDriver) driver <- rsDriver(browser = "chrome", port = 4567L, chromever = "latest") remDr <- driver$client # 访问目标页面 remDr$navigate("https://uk.finance.yahoo.com/quote/IPX.L/key-statistics?p=IPX.L") Sys.sleep(5) # 等待页面完全加载 # 获取页面HTML page_html <- remDr$getPageSource()[[1]] %>% read_html() # 提取指定数据 trailing_pe <- page_html %>% html_elements(xpath = "//td[text()='Trailing P/E']/following-sibling::td") %>% html_text() %>% as.numeric() book_value <- page_html %>% html_elements(xpath = "//td[text()='Book value per share (mrq)']/following-sibling::td") %>% html_text() %>% as.numeric() latest_price <- page_html %>% html_elements("fin-streamer[data-field='regularMarketPrice']") %>% html_text() %>% as.numeric() price_time <- page_html %>% html_elements("div[data-testid='qsp-price-time']") %>% html_text() %>% str_extract("At close : .*") # 关闭浏览器及驱动 remDr$close() driver$server$stop() # 输出结果 cat("Trailing P/E:", trailing_pe, "\n") cat("Book value per share (mrq):", book_value, "\n") cat("Latest price and time:", latest_price, price_time, "\n")
注意事项
- 必须设置
user-agent模拟浏览器请求,否则Yahoo会拦截爬虫请求。 - Yahoo的页面结构(包括内嵌JSON路径)可能随时调整,需定期验证代码有效性。
- 使用RSelenium时,需确保本地安装对应版本的浏览器和驱动程序。
内容的提问来源于stack exchange,提问作者user22322433
相关产品推荐
相关产品推荐

