You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R的rvest包爬取Yahoo! Finance表格时未获取全部行的问题

Hey there! Let's break down why you're only getting 101 rows of AAPL data instead of the full 2012-2018 range you requested, and how to fix it.

The Root Cause

Yahoo Finance's historical data page uses dynamic content loading: when you first load the page, it only renders the most recent ~100 days of data. Older data is loaded via AJAX requests only when you scroll down the page. Since rvest grabs the initial HTML source code (before any scrolling), it can't access the data that hasn't been loaded yet.

Solution 1: Use a Purpose-Built Yahoo Finance R Package (Best & Most Reliable)

Instead of scraping HTML, use packages that directly interact with Yahoo Finance's API—this avoids dealing with dynamic page loading entirely, and gives you clean, structured data out of the box.

Option A: quantmod (Classic Choice)

library(quantmod)
# Pull AAPL data for your desired date range
getSymbols("AAPL", from = "2012-01-01", to = "2018-05-14")
# Convert the xts object to a data frame for easy manipulation
aapl_df <- data.frame(Date = index(AAPL), coredata(AAPL))
# Check the number of rows (should match the full date range)
nrow(aapl_df)

Option B: yfinance (Modern, Up-to-Date Alternative)

library(yfinance)
# Fetch the data directly via the API
aapl_data <- yf_get("AAPL", start = "2012-01-01", end = "2018-05-14")
# Verify the row count
nrow(aapl_data)

Solution 2: Simulate Browser Scrolling (Advanced & Less Stable)

If you absolutely need to use rvest, you can pair it with RSelenium to simulate a browser scrolling down the page, triggering the dynamic data loads. Note that this method is fragile—Yahoo could change their page structure at any time, breaking the code.

Here's a rough example:

library(RSelenium)
library(rvest)

# Start a Chrome browser session (you'll need ChromeDriver installed)
driver <- rsDriver(browser = "chrome")
remDr <- driver[["client"]]

# Navigate to the Yahoo Finance page
url <- 'https://finance.yahoo.com/quote/AAPL/history?period1=1325350800&period2=1526230800&interval=1d&filter=history&frequency=1d'
remDr$navigate(url)

# Scroll to the bottom repeatedly to load all data
# Adjust the loop count based on how much data you need
for(i in 1:20){
  remDr$executeScript("window.scrollTo(0, document.body.scrollHeight);")
  Sys.sleep(2) # Wait for the data to load
}

# Grab the fully loaded page source and parse the table
page_html <- read_html(remDr$getPageSource()[[1]])
df <- page_html %>% html_nodes("table") %>% html_table() %>% .[[1]]

# Check the row count now
nrow(df)

# Clean up: close the browser and stop the driver
remDr$close()
driver$server$stop()

Final Note

Stick with the API-based packages if you can—they're designed for exactly this use case, require less maintenance, and give you more reliable results than web scraping for dynamic sites.

内容的提问来源于stack exchange,提问作者lukmanedwindra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:54:06