You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网站多表格爬取难题求助:无法生成可用CSV文件

问题描述

近两个月一直无法正确爬取目标页面的表格数据并生成可用CSV文件。试过WebScraper、Octoparse等无代码工具,也用了Python的Pandas、BeautifulSoup库,都没成功。原代码如下:

link = "https://sitem.herts.ac.uk/aeru/ppdb/en/atoz.htm"
    f = urllib.request.urlopen(link)
    html_doc= f.read()



   soup = BeautifulSoup(html_doc)
    #print(soup)
    pages = soup.find_all("a")
    filtred_pages = []
    for p in pages :
     if(p.has_attr('href') and p['href'].startswith("Report")):
       filtred_pages.append(p)

    #pages = list(filter(lambda el: el['href'].startswith("Report"), pages))
    print(filtred_pages)

    result = []
    for page in filtred_pages[:100]:
    f = urllib.request.urlopen('http://sitem.herts.ac.uk/aeru/ppdb/en/'+page['href'])
    html_doc= f.read()
    soup = BeautifulSoup(html_doc)
    #print(soup)
    titreNode = soup.find_all("td", attrs={"class" : "title"})[0].text
      trs = soup.select('table.report_data tr')
      rowDict = {}
      for tr in trs:
    if len(tr.select('td.row_header')) == 0:
      continue
    if len(tr.select('td.data1')) == 0:
      continue
    tdTitre = tr.select('td.row_header')[0].text
    tdValue = tr.select('td.data1')[0].text.replace('&nbsp','').rstrip('\n').strip()
    rowDict[tdTitre] = tdValue
  result.append(rowDict)
       #print(result)
       df = pd.DataFrame(result)
       df.to_csv('file.csv')
解决方案

原代码核心问题

  • 缩进混乱:Python对缩进要求严格,多处错误缩进会直接导致运行报错
  • 解析器缺失:初始化BeautifulSoup未指定解析器,可能引发解析异常
  • 协议不一致:主页面用HTTPS,子页面请求用HTTP,易触发访问限制或跳转
  • 字符处理不全:&nbsp是 的转义,直接替换无法清除实际的非换行空格
  • 无异常捕获:网络请求、节点查找出错时会直接中断爬取

修改后的可运行代码

import urllib.request
from bs4 import BeautifulSoup
import pandas as pd

# 主页面链接
link = "https://sitem.herts.ac.uk/aeru/ppdb/en/atoz.htm"
try:
    f = urllib.request.urlopen(link)
    html_doc = f.read()
except Exception as e:
    print(f"访问主页面失败: {e}")
    exit()

# 指定解析器初始化BeautifulSoup
soup = BeautifulSoup(html_doc, 'html.parser')
pages = soup.find_all("a")
filtred_pages = []

# 筛选目标报告链接
for p in pages:
    if p.has_attr('href') and p['href'].startswith("Report"):
        filtred_pages.append(p)

print(f"共找到 {len(filtred_pages)} 个目标页面")

result = []
# 遍历前100个页面(可按需调整数量)
for idx, page in enumerate(filtred_pages[:100]):
    page_url = f"https://sitem.herts.ac.uk/aeru/ppdb/en/{page['href']}"
    try:
        f = urllib.request.urlopen(page_url)
        html_doc = f.read()
        soup = BeautifulSoup(html_doc, 'html.parser')
        
        # 获取页面标题并加入数据字典
        titre_node = soup.find("td", attrs={"class": "title"})
        if not titre_node:
            print(f"第 {idx+1} 个页面无标题,跳过")
            continue
        row_dict = {"标题": titre_node.text.strip()}
        
        # 提取表格数据
        trs = soup.select('table.report_data tr')
        for tr in trs:
            row_header = tr.select('td.row_header')
            data1 = tr.select('td.data1')
            if not row_header or not data1:
                continue
            td_titre = row_header[0].text.strip()
            # 彻底清理空格和换行符
            td_value = data1[0].text.replace('\xa0', '').replace('\n', '').strip()
            row_dict[td_titre] = td_value
        
        result.append(row_dict)
        print(f"已处理第 {idx+1} 个页面")
    except Exception as e:
        print(f"处理第 {idx+1} 个页面失败: {e}")
        continue

# 生成CSV文件
if result:
    df = pd.DataFrame(result)
    # 指定utf-8-sig编码避免中文乱码
    df.to_csv('ppdb_data.csv', index=False, encoding='utf-8-sig')
    print("CSV文件已成功生成")
else:
    print("未获取到有效数据,无法生成CSV")

优化说明

  • 修复所有缩进问题,保证代码语法正确
  • 统一使用HTTPS协议请求,避免协议冲突
  • 明确指定html.parser解析器,提升解析稳定性
  • 完善字符处理:清除\xa0(非换行空格)和多余换行符
  • 添加异常捕获,单个页面出错不会中断整个爬取流程
  • 将页面标题加入数据,CSV内容更直观
  • 指定utf-8-sig编码,避免导出CSV后中文乱码
  • 增加进度提示,方便跟踪爬取状态

内容的提问来源于stack exchange,提问作者Marie Gainche

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 07:24:22