获取标普500全量历史数据时遇KeyError: 'Date'问题求助
Fixing KeyError: 'Date' & Incomplete S&P 500 Data from Yahoo Finance
Hey there, let's break down why you're hitting that KeyError: 'Date' and only getting 64 stocks loaded successfully:
Root Causes
- Yahoo Finance API Deprecation: The old
pandas_datareaderYahoo data source is no longer maintained and can't interact with Yahoo's updated API structure. That's why you're seeing theKeyError—the data returned no longer has the "Date" column the legacy code expects. - Unreliable Ticker Cleaning: Your current code uses
ticker[:-1]to trim trailing characters from Wikipedia's ticker entries, but this is risky. Some valid tickers (likeBRK.B) contain dots, and slicing might accidentally mangle them if the trailing character isn't just a newline. - Lack of Error Handling: If a single ticker fails to load, your entire script stops dead, leaving you with incomplete data instead of continuing to fetch the rest.
Solution: Switch to yfinance & Refactor Code
yfinance is the community-maintained library built specifically for Yahoo Finance data, and it works seamlessly with pandas. Here's the revised code that fixes all your issues:
import bs4 as bs import datetime as dt import os import yfinance as yf import pickle import requests def save_sp500_tickers(): resp = requests.get('https://en.wikipedia.org/wiki/List_of_S%26P_500_companies') soup = bs.BeautifulSoup(resp.text, 'lxml') table = soup.find('table', {'class': 'wikitable sortable'}) tickers = [] for row in table.findAll('tr')[1:]: # Use strip() to safely remove whitespace/newlines instead of risky slicing ticker = row.findAll('td')[0].text.strip() tickers.append(ticker) with open("sp500tickers.pickle", "wb") as f: pickle.dump(tickers, f) return tickers def get_data_from_yahoo(reload_sp500=False): if reload_sp500: tickers = save_sp500_tickers() else: with open("sp500tickers.pickle", "rb") as f: tickers = pickle.load(f) if not os.path.exists('stock_dfs'): os.makedirs('stock_dfs') start = dt.datetime(2000, 1, 1) end = dt.datetime(2016,12,31) for ticker in tickers: csv_path = f'stock_dfs/{ticker}.csv' if not os.path.exists(csv_path): try: # Use yfinance to fetch data instead of outdated pandas_datareader df = yf.download(ticker, start=start, end=end) if not df.empty: # Only save if we actually retrieved data df.to_csv(csv_path) print(f'Saved {ticker}') else: print(f'No historical data found for {ticker}') except Exception as e: print(f'Failed to fetch {ticker}: {str(e)}') else: print(f'Already have {ticker}') get_data_from_yahoo()
Key Improvements
- Switched to
yfinance: This library handles Yahoo's current API endpoints correctly, eliminating theKeyErrorentirely. - Safer Ticker Cleaning:
strip()removes any leading/trailing whitespace (like newlines) without risking valid ticker characters. - Error Handling: The
try-exceptblock catches issues with individual tickers, so your script keeps running even if some stocks fail to load. - Empty Data Check: We only save the CSV if we actually got data for the ticker, avoiding empty or broken files.
Quick Setup Step
- Install
yfinanceif you haven't already:pip install yfinance
内容的提问来源于stack exchange,提问作者dkar
相关产品推荐
相关产品推荐

