求助:Python爬取Yahoo Finance股票历史数据仅获前100条
Hey there! I totally get where you're coming from—scrolling-based lazy loading is one of the most common hurdles new web scrapers hit with sites like Yahoo Finance. Let’s walk through two solid solutions to get all that historical stock data you need, no more stuck at 100 rows!
If you want to mimic how a human would interact with the page, Selenium is perfect for this. It controls a real browser, so you can replicate the "scroll to bottom" action that triggers more data to load. Here's how to set it up:
First, install Selenium and grab a webdriver (like ChromeDriver for Chrome):
pip install seleniumMake sure ChromeDriver is downloaded and added to your system PATH, or you can specify its file path directly in the code.
Now, write the code to scroll until all data is loaded:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd import time # Launch the browser and navigate to the stock's history page driver = webdriver.Chrome() stock_url = "https://finance.yahoo.com/quote/AAPL/history?p=AAPL" # Swap AAPL for your ticker driver.get(stock_url) # Wait for the initial table to load before doing anything WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.XPATH, "//table[@data-test='historical-prices']"))) # Scroll repeatedly until no new data loads last_scroll_height = driver.execute_script("return document.body.scrollHeight") while True: # Scroll to the very bottom of the page driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Give the page time to load new rows (adjust the sleep time if needed) time.sleep(2) # Check if we've reached the bottom new_scroll_height = driver.execute_script("return document.body.scrollHeight") if new_scroll_height == last_scroll_height: break last_scroll_height = new_scroll_height # Extract the full table data and convert to a DataFrame table_element = driver.find_element(By.XPATH, "//table[@data-test='historical-prices']") stock_df = pd.read_html(table_element.get_attribute('outerHTML'))[0] # Clean up messy rows (like dividend announcements that break the table structure) stock_df = stock_df[stock_df['Open'].notna()] # Export to CSV stock_df.to_csv('full_stock_history.csv', index=False) # Close the browser when done driver.quit()Quick notes: If you see a cookie consent pop-up, add a line to click the "Accept" button using Selenium before scrolling. Also, tweak the
time.sleep(2)if your internet is slower—you want to give the page enough time to load new data.
Here's a pro tip: Yahoo Finance has a hidden API that serves historical data directly, so you don't need to mess with browser automation at all. The yfinance library wraps this API, making it super simple to pull all available data in one shot.
Install the library first:
pip install yfinanceThen, write a few lines of code to get and export your data:
import yfinance as yf # Define your stock ticker stock_ticker = yf.Ticker("AAPL") # Replace with your desired ticker # Fetch ALL historical data (use period="max" for everything, or specify dates) full_history = stock_ticker.history(period="max") # Export to CSV full_history.to_csv('complete_stock_data.csv')This method is way more reliable and efficient. You can also filter by date range if you don't need every single row—just use
startandendparameters, likestock_ticker.history(start="2015-01-01", end="2024-01-01").
- For Selenium users: If the table's HTML structure changes, use your browser's dev tools (F12) to inspect the table and update the XPath/CSS selector in your code.
- Always check Yahoo Finance's terms of service before scraping to make sure you're complying with their rules.
- The
yfinancemethod is hands-down the best choice for most cases—it handles all the lazy loading and API complexity behind the scenes.
内容的提问来源于stack exchange,提问作者Dhanush M

