You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Yahoo Finance营收数据:BeautifulSoup拆分难题

Fixing Total Revenue Extraction from Yahoo Finance

Hey there! I see exactly what's going on here—you're grabbing the entire <td> tag object instead of just the text inside it. Let's get that clean revenue number for you.

The Problem

Right now, stock_total_revenue is a BeautifulSoup Tag object, so printing it dumps the full HTML of that element. The actual revenue value is nested inside a <span> tag within the <td>, but we don't need to target the span directly—BeautifulSoup gives us easy ways to pull out the text content.

The Solution

Use either the .text attribute or the .get_text() method on the Tag object to extract just the text. Even better, add a check to handle cases where the element isn't found (to avoid AttributeErrors), and use strip=True to clean up any extra whitespace.

Here's your updated code:

from bs4 import BeautifulSoup
import requests

def get_fundamentals(ticker):
    # Grab fundamental stock data from Yahoo Finance
    url = f"https://finance.yahoo.com/quote/{ticker.upper()}/financials?p={ticker.upper()}"
    html = requests.get(url)
    soup = BeautifulSoup(html.text, 'html.parser')
    
    stock_total_revenue = soup.find('td', {'class':'Fz(s) Ta(end) Pstart(10px)'})
    
    # Handle cases where the element might not exist
    if stock_total_revenue:
        # Extract clean text, stripping extra whitespace
        revenue_value = stock_total_revenue.get_text(strip=True)
        print(f"Total Revenue: {revenue_value}")
    else:
        print("Oops, couldn't find total revenue data for that ticker.")

ticker = input("Please enter a stock ticker: ")
get_fundamentals(ticker)

A More Reliable Alternative

Yahoo Finance's page structure can change over time, so targeting the "Total Revenue" row directly might be more robust than relying on a single class. Here's how to do that:

from bs4 import BeautifulSoup
import requests

def get_fundamentals(ticker):
    url = f"https://finance.yahoo.com/quote/{ticker.upper()}/financials?p={ticker.upper()}"
    html = requests.get(url)
    soup = BeautifulSoup(html.text, 'html.parser')
    
    # First find the element containing "Total Revenue"
    revenue_label = soup.find(string='Total Revenue')
    if revenue_label:
        # Traverse to the adjacent td with the actual value
        revenue_value = revenue_label.find_next('td', class_='Fz(s) Ta(end) Pstart(10px)').get_text(strip=True)
        print(f"Total Revenue: {revenue_value}")
    else:
        print("Couldn't locate the Total Revenue row.")

ticker = input("Please enter a stock ticker: ")
get_fundamentals(ticker)

This way, you're first identifying the row that corresponds to Total Revenue, then grabbing the value from the next cell—this reduces the chance of pulling the wrong data if other cells share the same class.

内容的提问来源于stack exchange,提问作者Simeon Ikudabo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:00:57