You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用rvest爬取雅虎财经的完整比特币历史数据集

雅虎财经比特币历史数据全量爬取解决方案

你遇到的问题根源是雅虎财经的历史价格页面采用了无限滚动加载的交互逻辑,静态请求read_html()只能获取页面首次渲染加载的前101条数据,剩余数据需要触发页面滚动后才会异步加载,因此普通的静态爬虫无法直接拿到全量数据。

方案1:调用公开数据接口获取(最推荐)

雅虎财经本身有公开的历史数据查询接口,不需要爬取前端页面,直接构造请求即可拿到指定时间区间的全量数据,示例代码如下:

library(htTR)
library(tidyverse)

# 配置请求参数
params <- list(
  period1 = 1480464000,
  period2 = 1638230400,
  interval = "1d",
  events = "history",
  includeAdjustedClose = "true"
)

# 发送请求并解析返回结果
response <- GET("https://query1.finance.yahoo.com/v8/finance/chart/BTC-USD", query = params)
raw_data <- content(response, as = "parsed")

# 转换为结构化表格
timestamp_list <- raw_data$chart$result[[1]]$timestamp
quote_data <- raw_data$chart$result[[1]]$indicators$quote[[1]]

btc_full_data <- tibble(
  date = as.POSIXct(unlist(timestamp_list), origin = "1970-01-01", tz = "UTC"),
  open = unlist(quote_data$open),
  high = unlist(quote_data$high),
  low = unlist(quote_data$low),
  close = unlist(quote_data$close),
  volume = unlist(quote_data$volume),
  adjusted_close = unlist(raw_data$chart$result[[1]]$indicators$adjclose[[1]]$adjclose)
)

该方案无需处理前端渲染逻辑,运行效率高,返回数据完整,是最优选择。

方案2:动态爬虫模拟滚动加载

如果必须从前端页面提取数据,可以使用RSelenium模拟浏览器操作,反复滚动页面触发数据加载,全部加载完成后再提取表格,示例逻辑如下:

library(RSelenium)
library(rvest)

# 启动Chrome浏览器驱动,需提前安装对应版本的chromedriver
driver <- rsDriver(browser = "chrome", chromever = "你的Chrome浏览器对应版本号")
remote_driver <- driver[["client"]]

# 打开目标页面
remote_driver$navigate("https://finance.yahoo.com/quote/BTC-USD/history?period1=1480464000&period2=1638230400&interval=1d&filter=history&frequency=1d&includeAdjustedClose=true")

# 循环滚动直到页面不再加载新内容
last_page_height <- 0
while(TRUE) {
  # 滚动到页面底部
  remote_driver$executeScript("window.scrollTo(0, document.body.scrollHeight);")
  Sys.sleep(2) # 等待数据加载完成
  current_page_height <- remote_driver$executeScript("return document.body.scrollHeight;")[[1]]
  if (current_page_height == last_page_height) {
    break
  }
  last_page_height <- current_page_height
}

# 提取完整表格数据
page_source <- remote_driver$getPageSource()[[1]]
cryp_table <- read_html(page_source) %>% 
  html_nodes(xpath = '//*[@id="Col1-1-HistoricalDataTable-Proxy"]/section/div[2]/table') %>% 
  html_table(fill = T)
cryp_final <- cryp_table[[1]]

# 关闭驱动释放资源
remote_driver$close()
driver$server$stop()

该方案运行效率较低,需要提前配置浏览器驱动环境,仅作为备选方案。

内容的提问来源于stack exchange,提问作者sebastian fanchi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 03:36:04