You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清洗Bloomberg股票指数爬虫数据并保存为CSV文件?

Clean & Save Bloomberg Stock Index Data to CSV

Hey there! You're already halfway there with your code—let's refine it to extract clean, usable data and save it directly to a CSV file. Here's how to do it step by step:

Step 1: Add Required Modules

First, we'll need the csv module to handle writing the CSV file. We'll also use built-in string cleaning methods to strip out extra whitespace and HTML cruft.

Step 2: Extract & Clean Data

Instead of printing raw HTML siblings, we'll target each table cell, extract its text content, and clean it up. We'll also grab the table headers to make our CSV readable and structured.

Full Working Code

from urllib.request import urlopen
from bs4 import BeautifulSoup
import csv

# Fetch and parse the Bloomberg page
html = urlopen('https://www.bloomberg.com/markets/stocks')
bs = BeautifulSoup(html, 'html.parser')

# Extract table headers (to use as CSV column names)
header_cells = bs.find('thead').find_all('th')
clean_headers = [header.get_text(strip=True) for header in header_cells]

# Extract and clean table rows
clean_rows = []
tbody = bs.find('tbody', {'class': 'data-table-body'})
for row in tbody.find_all('tr'):
    # Pull text from each cell, stripping extra spaces/newlines
    cell_contents = [cell.get_text(strip=True) for cell in row.find_all('td')]
    # Skip any empty rows that might come from hidden HTML elements
    if cell_contents:
        clean_rows.append(cell_contents)

# Save the cleaned data to a CSV file
with open('bloomberg_stock_indices.csv', 'w', newline='', encoding='utf-8') as csv_file:
    csv_writer = csv.writer(csv_file)
    # Write headers first
    csv_writer.writerow(clean_headers)
    # Write all the cleaned rows
    csv_writer.writerows(clean_rows)

print("Data saved successfully to bloomberg_stock_indices.csv!")

Key Improvements Explained

  • Header Extraction: We pull the table headers to give your CSV meaningful column names (like "Index", "Price", "Change %", etc.) instead of generic columns.
  • Text Cleaning: get_text(strip=True) removes extra spaces, line breaks, and hidden HTML formatting from each cell, leaving only the raw, usable data.
  • CSV Handling: The csv.writer handles proper formatting (like escaping commas within values, correct encoding) so your file opens seamlessly in Excel, Google Sheets, or any data tool.
  • Empty Row Filter: We skip any empty rows that might be present in the HTML to avoid cluttering your CSV.

Quick Troubleshooting Tip

If you run into blocked requests (Bloomberg sometimes uses anti-scraping measures), add a User-Agent header to mimic a browser request:

from urllib.request import Request, urlopen
req = Request('https://www.bloomberg.com/markets/stocks', headers={'User-Agent': 'Mozilla/5.0'})
html = urlopen(req)

内容的提问来源于stack exchange,提问作者Martin598

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:39:10