Python爬取Yahoo Finance营收数据:BeautifulSoup拆分难题
Hey there! I see exactly what's going on here—you're grabbing the entire <td> tag object instead of just the text inside it. Let's get that clean revenue number for you.
The Problem
Right now, stock_total_revenue is a BeautifulSoup Tag object, so printing it dumps the full HTML of that element. The actual revenue value is nested inside a <span> tag within the <td>, but we don't need to target the span directly—BeautifulSoup gives us easy ways to pull out the text content.
The Solution
Use either the .text attribute or the .get_text() method on the Tag object to extract just the text. Even better, add a check to handle cases where the element isn't found (to avoid AttributeErrors), and use strip=True to clean up any extra whitespace.
Here's your updated code:
from bs4 import BeautifulSoup import requests def get_fundamentals(ticker): # Grab fundamental stock data from Yahoo Finance url = f"https://finance.yahoo.com/quote/{ticker.upper()}/financials?p={ticker.upper()}" html = requests.get(url) soup = BeautifulSoup(html.text, 'html.parser') stock_total_revenue = soup.find('td', {'class':'Fz(s) Ta(end) Pstart(10px)'}) # Handle cases where the element might not exist if stock_total_revenue: # Extract clean text, stripping extra whitespace revenue_value = stock_total_revenue.get_text(strip=True) print(f"Total Revenue: {revenue_value}") else: print("Oops, couldn't find total revenue data for that ticker.") ticker = input("Please enter a stock ticker: ") get_fundamentals(ticker)
A More Reliable Alternative
Yahoo Finance's page structure can change over time, so targeting the "Total Revenue" row directly might be more robust than relying on a single class. Here's how to do that:
from bs4 import BeautifulSoup import requests def get_fundamentals(ticker): url = f"https://finance.yahoo.com/quote/{ticker.upper()}/financials?p={ticker.upper()}" html = requests.get(url) soup = BeautifulSoup(html.text, 'html.parser') # First find the element containing "Total Revenue" revenue_label = soup.find(string='Total Revenue') if revenue_label: # Traverse to the adjacent td with the actual value revenue_value = revenue_label.find_next('td', class_='Fz(s) Ta(end) Pstart(10px)').get_text(strip=True) print(f"Total Revenue: {revenue_value}") else: print("Couldn't locate the Total Revenue row.") ticker = input("Please enter a stock ticker: ") get_fundamentals(ticker)
This way, you're first identifying the row that corresponds to Total Revenue, then grabbing the value from the next cell—this reduces the chance of pulling the wrong data if other cells share the same class.
内容的提问来源于stack exchange,提问作者Simeon Ikudabo

