You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas将雅虎财经爬取数据存入可索引DataFrame

问题

从雅虎财经爬取数据时,爬取过程正常,但将追加的列表存入可索引DataFrame(financial_dir)时返回空DataFrame,存入普通DataFrame(data5)则能正常存储。打印temp可以看到数据,将其转为DataFrame也成功,但执行financial_dir[ticker]=temp.append(...)时无法生成可通过股票代码(如INDUSINDBK.NS)调用对应数据的可索引结构。

用户原代码如下:

import requests
from bs4 import BeautifulSoup
import pandas as pd

tickers = ['KOTAKBANK.NS','WIPRO.NS','HINDALCO.NS','RELIANCE.NS',
           'INDUSINDBK.NS','HDFCLIFE.NS','TATACONSUM.NS','TITAN.NS',
           'ULTRACEMCO.NS']

financial_dir = pd.DataFrame()
temp = []
for ticker in tickers:
    url = 'https://finance.yahoo.com/quote/'+ticker+'/financials?p='+ticker
    page = requests.get(url, headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/105.0.0.0 Safari/537.36'})
    soup = BeautifulSoup(page.text, 'html.parser')
    a = list(range(0,2000,1))
    try:
        for i in a:
            financial_dir[ticker]=temp.append(soup.find('div', {'class' : "D(tbrg)"}).find_all('div')[i].get_text(separator='|').split('|'))
    except:
        pass

temp
data5 = pd.DataFrame(temp)
financial_dir

解决方案

问题根源

  1. list.append()返回值问题:temp.append(...)的返回值是None,而非追加后的列表,直接赋值给financial_dir[ticker]会导致该列全为None,最终DataFrame为空。
  2. 数据结构设计错误:想要通过股票代码调用对应DataFrame,应使用字典存储每个股票的DataFrame,而非将所有数据塞进一个大DataFrame的列中——不同股票的财务数据行列数可能不一致,强行按列存储会导致数据错位或丢失。

修改后的代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

tickers = ['KOTAKBANK.NS','WIPRO.NS','HINDALCO.NS','RELIANCE.NS',
           'INDUSINDBK.NS','HDFCLIFE.NS','TATACONSUM.NS','TITAN.NS',
           'ULTRACEMCO.NS']

# 用字典存储每个股票的DataFrame,键为股票代码,满足按代码调用的需求
financial_dir = {}

for ticker in tickers:
    url = f'https://finance.yahoo.com/quote/{ticker}/financials?p={ticker}'
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/105.0.0.0 Safari/537.36'}
    page = requests.get(url, headers=headers)
    soup = BeautifulSoup(page.text, 'html.parser')
    
    temp = []
    # 定位财务数据容器,避免无效的大范围循环
    table_container = soup.find('div', {'class': "D(tbrg)"})
    if not table_container:
        print(f"{ticker} 未找到财务数据容器")
        continue
    
    rows = table_container.find_all('div')
    for row in rows:
        try:
            # 提取并分割每行数据
            row_data = row.get_text(separator='|').split('|')
            temp.append(row_data)
        except Exception as e:
            print(f"{ticker} 处理单行数据出错: {e}")
            continue
    
    # 将当前股票的有效数据转为DataFrame并存入字典
    if temp:
        financial_dir[ticker] = pd.DataFrame(temp)
    else:
        print(f"{ticker} 无有效财务数据")

# 示例:调用INDUSINDBK.NS对应的财务数据
print(financial_dir['INDUSINDBK.NS'])

关键改进点

  • 改用字典financial_dir存储单股票DataFrame,直接通过股票代码键调用,符合需求。
  • 移除无效的range(0,2000,1)循环,直接遍历页面中实际存在的行元素,提升效率并减少异常触发。
  • 每次处理新股票时重置temp列表,避免不同股票的数据混淆。
  • 增加异常捕获和日志输出,方便排查单股票的数据爬取问题。

内容的提问来源于stack exchange,提问作者Sonu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 23:10:55