多变量Python股票数据爬取求助:代码仅获取数组最后一项数据
Let's break down what's going wrong and build a complete solution that meets all your requirements.
Why Your Current Code Only Gets the Last Item
First, the most likely issue is that your selector soup.find('td', attrs={'class': 'col2'}) isn't reliably targeting the "Dividend Yield Range, Past 5 Years - Average" value across all YCharts pages. The site's HTML structure might have shifted, or that selector matches irrelevant elements for some stocks, leaving you with valid data only for the last ticker.
Second, your CSV writing logic appends single values with timestamps, which doesn't align with your goal of having rows for each stock and columns for every data point.
Step 1: Fix YCharts Data Extraction
We'll update the selector to target the exact metric row using text matching (far more reliable than fragile class names). We'll also switch to requests (a modern, easier-to-use alternative to urllib2) and add error handling to avoid crashes if pages fail to load or data is missing.
Step 2: Add Finviz Data Extraction
For Finviz, we'll parse the "Book/sh" and "LT Debt/Eq" values from the valuation table on each quote page, using row-by-row label matching to find the right metrics.
Step 3: Structure the CSV Correctly
We'll build a CSV where each row represents one stock ticker, and columns correspond to all your requested data points:
- Ticker
- 5-Year Avg Dividend Yield
- 5-Year Avg PE Ratio
- Book/Share
- LT Debt/Equity
Full Working Code
import requests from bs4 import BeautifulSoup import csv from datetime import datetime # List of stock tickers to process tickers = ['AAPL', 'T', 'MMM'] # Initialize storage for all stock data stock_data = [] # Mimic a browser request to avoid being blocked by target sites headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } def get_ycharts_5year_avg(ticker, metric): """Fetch 5-year average metric from YCharts""" url = f'https://ycharts.com/companies/{ticker}/{metric}' try: response = requests.get(url, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # Find the row with the exact "Past 5 Years - Average" label avg_row = soup.find('tr', string='Past 5 Years - Average') if avg_row: # Grab the value from the adjacent table cell return avg_row.find_next_sibling('td').text.strip() return 'N/A' except Exception as e: print(f"Error fetching {metric} for {ticker}: {str(e)}") return 'N/A' def get_finviz_metrics(ticker): """Fetch Book/sh and LT Debt/Eq from Finviz""" url = f'https://www.finviz.com/quote.ashx?t={ticker}' try: response = requests.get(url, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') val_table = soup.find('table', class_='snapshot-table2') book_sh, lt_debt_eq = 'N/A', 'N/A' if val_table: # Iterate through table rows to find target metrics for row in val_table.find_all('tr'): cells = row.find_all('td') if len(cells) == 2: label = cells[0].text.strip() value = cells[1].text.strip() if label == 'Book/sh': book_sh = value elif label == 'LT Debt/Eq': lt_debt_eq = value return book_sh, lt_debt_eq except Exception as e: print(f"Error fetching Finviz data for {ticker}: {str(e)}") return 'N/A', 'N/A' # Process each ticker for ticker in tickers: print(f"Processing {ticker}...") # Pull YCharts data div_yield_avg = get_ycharts_5year_avg(ticker, 'dividend_yield') pe_ratio_avg = get_ycharts_5year_avg(ticker, 'pe_ratio') # Pull Finviz data book_sh, lt_debt_eq = get_finviz_metrics(ticker) # Store all data for the ticker stock_data.append([ ticker, div_yield_avg, pe_ratio_avg, book_sh, lt_debt_eq ]) # Write data to CSV csv_filename = f'stock_metrics_{datetime.now().strftime("%Y%m%d")}.csv' with open(csv_filename, 'w', newline='', encoding='utf-8') as csv_file: writer = csv.writer(csv_file) # Write header row writer.writerow([ 'Ticker', '5-Year Avg Dividend Yield', '5-Year Avg PE Ratio', 'Book/Share', 'LT Debt/Equity' ]) # Write all stock data rows writer.writerows(stock_data) print(f"Success! Data saved to {csv_filename}")
Key Improvements
- Reliable Selectors: Uses text matching to target exact metric rows, avoiding issues from site layout changes.
- Error Handling: Catches network errors and missing data, returning
N/Ainstead of crashing. - Browser Headers: Mimics real traffic to reduce the chance of being blocked by YCharts or Finviz.
- Proper CSV Structure: Aligns with your requested format (rows = tickers, columns = data points).
- Modular Code: Breaks extraction into reusable functions for easy updates or additions of new metrics.
Quick Setup Note
Before running the code, install the required packages with this command:
pip install requests beautifulsoup4
内容的提问来源于stack exchange,提问作者p3nd0l0

