如何清洗Bloomberg股票指数爬虫数据并保存为CSV文件?
Clean & Save Bloomberg Stock Index Data to CSV
Hey there! You're already halfway there with your code—let's refine it to extract clean, usable data and save it directly to a CSV file. Here's how to do it step by step:
Step 1: Add Required Modules
First, we'll need the csv module to handle writing the CSV file. We'll also use built-in string cleaning methods to strip out extra whitespace and HTML cruft.
Step 2: Extract & Clean Data
Instead of printing raw HTML siblings, we'll target each table cell, extract its text content, and clean it up. We'll also grab the table headers to make our CSV readable and structured.
Full Working Code
from urllib.request import urlopen from bs4 import BeautifulSoup import csv # Fetch and parse the Bloomberg page html = urlopen('https://www.bloomberg.com/markets/stocks') bs = BeautifulSoup(html, 'html.parser') # Extract table headers (to use as CSV column names) header_cells = bs.find('thead').find_all('th') clean_headers = [header.get_text(strip=True) for header in header_cells] # Extract and clean table rows clean_rows = [] tbody = bs.find('tbody', {'class': 'data-table-body'}) for row in tbody.find_all('tr'): # Pull text from each cell, stripping extra spaces/newlines cell_contents = [cell.get_text(strip=True) for cell in row.find_all('td')] # Skip any empty rows that might come from hidden HTML elements if cell_contents: clean_rows.append(cell_contents) # Save the cleaned data to a CSV file with open('bloomberg_stock_indices.csv', 'w', newline='', encoding='utf-8') as csv_file: csv_writer = csv.writer(csv_file) # Write headers first csv_writer.writerow(clean_headers) # Write all the cleaned rows csv_writer.writerows(clean_rows) print("Data saved successfully to bloomberg_stock_indices.csv!")
Key Improvements Explained
- Header Extraction: We pull the table headers to give your CSV meaningful column names (like "Index", "Price", "Change %", etc.) instead of generic columns.
- Text Cleaning:
get_text(strip=True)removes extra spaces, line breaks, and hidden HTML formatting from each cell, leaving only the raw, usable data. - CSV Handling: The
csv.writerhandles proper formatting (like escaping commas within values, correct encoding) so your file opens seamlessly in Excel, Google Sheets, or any data tool. - Empty Row Filter: We skip any empty rows that might be present in the HTML to avoid cluttering your CSV.
Quick Troubleshooting Tip
If you run into blocked requests (Bloomberg sometimes uses anti-scraping measures), add a User-Agent header to mimic a browser request:
from urllib.request import Request, urlopen req = Request('https://www.bloomberg.com/markets/stocks', headers={'User-Agent': 'Mozilla/5.0'}) html = urlopen(req)
内容的提问来源于stack exchange,提问作者Martin598
相关产品推荐
相关产品推荐

