You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用BeautifulSoup正确抓取表格并规整数据为CSV文件?

解决DSE收盘价表格抓取并生成规范CSV的问题

问题分析

直接提取表格原始文本写入CSV会忽略表格的行/单元格结构,导致输出内容杂乱。要生成规范CSV,需按表格的<tr>(行)和<td>/<th>(单元格)结构解析内容,再用CSV标准格式输出。

修正后的代码

from bs4 import BeautifulSoup
import requests
import csv

def get_stats():
    url = 'https://dsebd.org/dse_close_price_archive.php?startDate=2021-08-28&endDate=2023-08-28&inst=REPUBLIC&archive=data'
    # 模拟浏览器请求头,避免被反爬拦截
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/117.0.0.0 Safari/537.36'
    }
    response = requests.get(url, headers=headers)
    response.encoding = 'utf-8'  # 确保文本编码正确,避免乱码
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # 定位目标表格
    table = soup.find("table", class_="table table-bordered background-white")
    if not table:
        print("未找到目标表格")
        return
    
    # 写入CSV文件
    with open("stats.csv", 'w', newline='', encoding='utf-8') as file:
        writer = csv.writer(file)
        # 遍历表格每一行
        for row in table.find_all('tr'):
            # 提取每个单元格的文本,清理多余空格和换行
            cells = [cell.get_text(strip=True) for cell in row.find_all(['td', 'th'])]
            # 将该行数据写入CSV
            writer.writerow(cells)
    
    print("规范CSV文件已生成")
    return "success"

get_stats()

关键改进点

  • 使用csv.writer模块:自动处理CSV格式规则(逗号分隔、换行、特殊字符转义),避免手动拼接的混乱。
  • 结构化解析表格:按行遍历<tr>元素,再提取每行中的<td>(数据单元格)和<th>(表头单元格)内容,保证输出结构与原表格一致。
  • 文本清理:通过get_text(strip=True)去除单元格文本前后的空格、换行符,避免多余空白干扰CSV格式。
  • 防拦截优化:添加浏览器请求头,设置正确编码,提升请求成功率并避免乱码。

内容的提问来源于stack exchange,提问作者agag gagaga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 00:37:20