You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将爬取的雅虎财经<tr>标签内容转换为DataFrame?

爬取Yahoo财经GOOG历史数据并转换为DataFrame的优化方案

问题描述

我正在爬取Yahoo财经的GOOG历史数据页面,希望把爬取结果整理成和页面展示格式一致的DataFrame。目前已经能获取<tr>标签的内容,但用以下代码把所有内容存入列表后,不知道怎么转换成DataFrame。我觉得除了给页面每列单独创建列表外,应该有更优的方法,求思路。

url = 'https://www.finance.yahoo.com/quote/GOOG/history?p=GOOG' 
headers = {'User-Agent': 
           'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36'}
r = requests.get(url, headers = headers)
soup = BeautifulSoup(r.content, 'html.parser')
data = []
for t in soup.select('table'):
    for tr in t.select('tr:has(td)'):
        print(type(tr))
        for span in tr: 
            data.append([d.get_text(strip = True) for d in span])
print(data)

优化思路与解决方案

1. 修正数据提取逻辑

你当前的代码在遍历<tr>下的元素时逻辑有误,导致data列表结构混乱。正确的做法是每一行<tr>对应DataFrame的一行,直接提取该行所有<td>的文本值,组成一个列表作为行数据:

import pandas as pd
import requests
from bs4 import BeautifulSoup

url = 'https://www.finance.yahoo.com/quote/GOOG/history?p=GOOG' 
headers = {'User-Agent': 
           'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36'}
r = requests.get(url, headers=headers)
soup = BeautifulSoup(r.content, 'html.parser')

# 提取表格头部(列名)
columns = [th.get_text(strip=True) for th in soup.select('table th')]
# 提取表格行数据
rows = []
for tr in soup.select('table tr:has(td)'):
    row_data = [td.get_text(strip=True) for td in tr.select('td')]
    # 跳过空行(页面可能存在分隔行)
    if row_data:
        rows.append(row_data)

# 直接转换为DataFrame
df = pd.DataFrame(rows, columns=columns)
print(df.head())

2. 核心优化点

  • 按行提取数据:不再拆分<tr>内的<span>,直接以行为单位收集数据,保证结构和页面表格完全匹配。
  • 自动获取列名:从页面<th>标签提取原生列名,直接作为DataFrame的列,无需手动定义。
  • 过滤无效行:页面可能存在无数据的分隔行,加入判断跳过,避免DataFrame出现空行或异常数据。

3. 额外注意事项

  • 若页面存在分页,需要解析下一页URL并循环请求,实现全量数据爬取。
  • Yahoo财经有反爬机制,频繁请求需加入time.sleep()延迟,避免IP被封禁。
  • 部分数据包含逗号、美元符号等非数值字符,后续可通过df.replace()或正则表达式清洗数据,方便后续数值计算。

内容的提问来源于stack exchange,提问作者MKN17

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 09:50:15