如何用Beautiful Soup从Yahoo财经HTML提取Reported EPS Basic数据
Hey there! Let's walk through how to pull those Reported EPS Basic values (0.23 and 0.20) from the 1928.HK Yahoo Finance page. I'll start by refining your existing code and adding the extraction logic.
Step 1: Fix the Request (Avoid Anti-Crawl Blocks)
First, Yahoo Finance often blocks requests that don't include a browser-like user agent. Let's update your request to include this, which will prevent 403 errors:
import requests from bs4 import BeautifulSoup url = "https://finance.yahoo.com/quote/1928.HK/financials?p=1928.HK" # Add a user agent to mimic a browser request headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } result = requests.get(url, headers=headers) result.raise_for_status() soup = BeautifulSoup(result.text, 'lxml') # Use result.text instead of content for easier handling
Step 2: Locate the "Reported EPS Basic" Row
Yahoo Finance structures its financial tables using div elements with specific classes. The rows are marked with D(tbr), and each cell uses D(tbc). We'll find the row where the first cell contains "Reported EPS Basic":
# Find the row containing "Reported EPS Basic" target_row = None for row in soup.find_all('div', class_='D(tbr)'): first_cell = row.find('div', class_='D(tbc)') if first_cell and 'Reported EPS Basic' in first_cell.get_text(strip=True): target_row = row break
Step 3: Extract the Numeric Values
Once we have the target row, we'll pull the values from the subsequent cells (skipping the first cell which holds the label):
if target_row: # Extract value cells (skip the first cell which is the label) value_cells = target_row.find_all('div', class_='D(tbc)', recursive=False)[1:] # Clean and collect the values reported_eps = [cell.get_text(strip=True) for cell in value_cells if cell.get_text(strip=True)] print("Reported EPS Basic Values:", reported_eps) # If you want to convert them to floats: # reported_eps_floats = [float(val) for val in reported_eps] else: print("Could not locate the 'Reported EPS Basic' row. Page structure might have changed.")
Notes for Future Maintenance
Yahoo Finance occasionally updates its page structure, so if this stops working:
- Right-click the "Reported EPS Basic" text on the page and select "Inspect" to check the latest class names for rows/cells.
- Adjust the
class_parameters in thefind_allcalls to match the updated structure.
内容的提问来源于stack exchange,提问作者Arthur Law

