使用R的rvest包爬取Yahoo! Finance表格时未获取全部行的问题
Hey there! Let's break down why you're only getting 101 rows of AAPL data instead of the full 2012-2018 range you requested, and how to fix it.
The Root Cause
Yahoo Finance's historical data page uses dynamic content loading: when you first load the page, it only renders the most recent ~100 days of data. Older data is loaded via AJAX requests only when you scroll down the page. Since rvest grabs the initial HTML source code (before any scrolling), it can't access the data that hasn't been loaded yet.
Solution 1: Use a Purpose-Built Yahoo Finance R Package (Best & Most Reliable)
Instead of scraping HTML, use packages that directly interact with Yahoo Finance's API—this avoids dealing with dynamic page loading entirely, and gives you clean, structured data out of the box.
Option A: quantmod (Classic Choice)
library(quantmod) # Pull AAPL data for your desired date range getSymbols("AAPL", from = "2012-01-01", to = "2018-05-14") # Convert the xts object to a data frame for easy manipulation aapl_df <- data.frame(Date = index(AAPL), coredata(AAPL)) # Check the number of rows (should match the full date range) nrow(aapl_df)
Option B: yfinance (Modern, Up-to-Date Alternative)
library(yfinance) # Fetch the data directly via the API aapl_data <- yf_get("AAPL", start = "2012-01-01", end = "2018-05-14") # Verify the row count nrow(aapl_data)
Solution 2: Simulate Browser Scrolling (Advanced & Less Stable)
If you absolutely need to use rvest, you can pair it with RSelenium to simulate a browser scrolling down the page, triggering the dynamic data loads. Note that this method is fragile—Yahoo could change their page structure at any time, breaking the code.
Here's a rough example:
library(RSelenium) library(rvest) # Start a Chrome browser session (you'll need ChromeDriver installed) driver <- rsDriver(browser = "chrome") remDr <- driver[["client"]] # Navigate to the Yahoo Finance page url <- 'https://finance.yahoo.com/quote/AAPL/history?period1=1325350800&period2=1526230800&interval=1d&filter=history&frequency=1d' remDr$navigate(url) # Scroll to the bottom repeatedly to load all data # Adjust the loop count based on how much data you need for(i in 1:20){ remDr$executeScript("window.scrollTo(0, document.body.scrollHeight);") Sys.sleep(2) # Wait for the data to load } # Grab the fully loaded page source and parse the table page_html <- read_html(remDr$getPageSource()[[1]]) df <- page_html %>% html_nodes("table") %>% html_table() %>% .[[1]] # Check the row count now nrow(df) # Clean up: close the browser and stop the driver remDr$close() driver$server$stop()
Final Note
Stick with the API-based packages if you can—they're designed for exactly this use case, require less maintenance, and give you more reliable results than web scraping for dynamic sites.
内容的提问来源于stack exchange,提问作者lukmanedwindra

