You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Yahoo Finance API移除后,如何用R爬取关键统计页面指定数据?

解决Yahoo Finance关键统计页面爬取问题(R语言)

问题原因

Yahoo Finance的关键统计页面大量内容通过JavaScript动态渲染,直接用rvest::read_html()只能获取静态HTML结构,无法加载JS生成的实际数据,因此返回空结果。

方案1:提取页面内嵌JSON数据(轻量高效)

Yahoo会将页面数据以JSON格式内嵌在<script type="application/json">标签中,无需模拟浏览器即可提取:

代码实现

library(httr)
library(rvest)
library(jsonlite)
library(dplyr)

# 目标股票的关键统计页面URL
url <- "https://uk.finance.yahoo.com/quote/IPX.L/key-statistics?p=IPX.L"

# 模拟浏览器请求,避免被拦截
page <- GET(
  url,
  user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
)
html_content <- content(page, "text", encoding = "UTF-8")

# 提取页面中的内嵌JSON数据块
script_tags <- html_content %>%
  read_html() %>%
  html_elements("script[type='application/json']") %>%
  html_text()

# 筛选包含关键统计数据的JSON片段
target_json <- script_tags[sapply(script_tags, function(x) grepl("trailingPE|bookValue", x))][1]
data_list <- fromJSON(target_json)

# 提取指定数据
## Trailing P/E
trailing_pe <- data_list$context$dispatcher$stores$QuoteSummaryStore$summaryDetail$trailingPE$raw
## Book value per share (mrq)
book_value_per_share <- data_list$context$dispatcher$stores$QuoteSummaryStore$summaryDetail$bookValue$raw
## 最新股价及时间
latest_price <- data_list$context$dispatcher$stores$QuoteSummaryStore$price$regularMarketPrice$raw
price_time <- data_list$context$dispatcher$stores$QuoteSummaryStore$price$regularMarketTime %>%
  as.POSIXct(origin = "1970-01-01", tz = "Europe/London") %>%
  format("%H:%M %p BST")

# 输出结果
cat("Trailing P/E:", trailing_pe, "\n")
cat("Book value per share (mrq):", book_value_per_share, "\n")
cat("Latest price and time:", latest_price, "At close :", price_time, "\n")

通用方案:将所有关键统计导出为DataFrame

# 提取所有关键统计数据
key_stats <- data_list$context$dispatcher$stores$QuoteSummaryStore$defaultKeyStatistics

# 转换为结构化DataFrame,处理空值
stats_df <- key_stats %>%
  lapply(function(x) ifelse(is.null(x$raw), NA, x$raw)) %>%
  as.data.frame() %>%
  t() %>%
  as.data.frame() %>%
  rownames_to_column(var = "Statistic") %>%
  rename(Value = V1)

# 优化统计名称格式(可选)
stats_df$Statistic <- gsub("_", " ", stats_df$Statistic)
stats_df$Statistic <- stringr::str_to_title(stats_df$Statistic)

# 查看结果
head(stats_df)

方案2:用RSelenium模拟浏览器加载(兼容动态渲染)

如果内嵌JSON结构发生变化,可通过模拟浏览器完全加载页面:

代码实现

library(RSelenium)
library(rvest)
library(dplyr)
library(stringr)

# 启动Chrome浏览器驱动(需提前安装Chrome和对应版本的ChromeDriver)
driver <- rsDriver(browser = "chrome", port = 4567L, chromever = "latest")
remDr <- driver$client

# 访问目标页面
remDr$navigate("https://uk.finance.yahoo.com/quote/IPX.L/key-statistics?p=IPX.L")
Sys.sleep(5) # 等待页面完全加载

# 获取页面HTML
page_html <- remDr$getPageSource()[[1]] %>% read_html()

# 提取指定数据
trailing_pe <- page_html %>%
  html_elements(xpath = "//td[text()='Trailing P/E']/following-sibling::td") %>%
  html_text() %>%
  as.numeric()

book_value <- page_html %>%
  html_elements(xpath = "//td[text()='Book value per share (mrq)']/following-sibling::td") %>%
  html_text() %>%
  as.numeric()

latest_price <- page_html %>%
  html_elements("fin-streamer[data-field='regularMarketPrice']") %>%
  html_text() %>%
  as.numeric()

price_time <- page_html %>%
  html_elements("div[data-testid='qsp-price-time']") %>%
  html_text() %>%
  str_extract("At close : .*")

# 关闭浏览器及驱动
remDr$close()
driver$server$stop()

# 输出结果
cat("Trailing P/E:", trailing_pe, "\n")
cat("Book value per share (mrq):", book_value, "\n")
cat("Latest price and time:", latest_price, price_time, "\n")

注意事项

  • 必须设置user-agent模拟浏览器请求,否则Yahoo会拦截爬虫请求。
  • Yahoo的页面结构(包括内嵌JSON路径)可能随时调整,需定期验证代码有效性。
  • 使用RSelenium时,需确保本地安装对应版本的浏览器和驱动程序。

内容的提问来源于stack exchange,提问作者user22322433

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 14:15:54