web.DataReader调用Yahoo的最大股票代码数及400个批量请求方案咨询
web.DataReader 1. Maximum Number of Tickers Supported by web.DataReader for Yahoo Finance
Yahoo Finance doesn’t publish an official limit on the number of tickers you can request in a single web.DataReader() call. But from real-world usage, most users run into failures when trying to fetch more than 200-250 tickers at once. This usually happens because:
- Anti-scraping safeguards: Yahoo’s servers flag large, rapid requests as suspicious traffic.
- Payload size limits: A single request with 400 tickers may exceed the server’s threshold for response size, leading to timeouts or truncated data.
- Implicit rate limits: Even without public rules, Yahoo enforces unstated limits to prevent server overload.
2. How to Fetch Data for 400 Tickers (Plus Adding Timers/Delays)
Here are two reliable approaches to pull data for 400 tickers successfully:
Option 1: Batch Requests with Delays (Using web.DataReader)
Split your 400-ticker list into smaller batches (e.g., 200 tickers per batch), fetch each batch individually, and add a delay between requests to avoid triggering blocks.
import pandas as pd import pandas_datareader.data as web import time # Your full list of 400 tickers ticker_list = ["AAPL", "MSFT", ...] # Insert all 400 tickers here batch_size = 200 all_historical_data = [] # Iterate over tickers in batches for start_idx in range(0, len(ticker_list), batch_size): end_idx = start_idx + batch_size current_batch = ticker_list[start_idx:end_idx] try: # Fetch data for the current batch batch_data = web.DataReader(current_batch, 'yahoo', '2013-01-01', '2018-01-20') all_historical_data.append(batch_data) print(f"Successfully fetched batch {start_idx//batch_size + 1}/{len(ticker_list)//batch_size + 1}") except Exception as e: print(f"Failed to fetch batch {start_idx//batch_size + 1}: {str(e)}") # Add a 3-5 second delay between batches to avoid overwhelming the server time.sleep(3) # Combine all batches into a single DataFrame final_data = pd.concat(all_historical_data, axis=1)
Option 2: Use yfinance (A More Reliable Yahoo Finance Library)
The pandas_datareader Yahoo Finance interface isn’t actively maintained—yfinance is now the de facto standard for Yahoo data. It handles batch requests more gracefully, supports multi-threading, and has built-in safeguards against rate limits.
import yfinance as yf # Your 400-ticker list ticker_list = ["AAPL", "MSFT", ...] # Fetch data with multi-threading and automatic rate limiting historical_data = yf.download( ticker_list, start='2013-01-01', end='2018-01-21', # Note: end date is exclusive in yfinance threads=4, group_by='ticker', auto_adjust=True # Automatically adjust prices for splits/dividends )
Key Notes on Timers/Delays
- For
web.DataReader, usetime.sleep()between batches to slow your request rate (3-5 seconds is a safe starting point). - For
yfinance, the library handles basic rate limiting internally, but you can add custom delays if you still hit issues (e.g., split into smaller batches and addtime.sleep()). - Consider adding retry logic (with libraries like
tenacity) for failed requests to handle temporary server errors.
内容的提问来源于stack exchange,提问作者Ross Demtschyna

