You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取遇相同页面HTML结构差异问题求助

解决网页爬取中HTML结构不一致导致程序崩溃的问题

核心思路

先判断当前页面的结构类型,针对不同结构执行对应逻辑,同时给表格提取逻辑加上容错处理,避免程序直接终止。

具体实现步骤

  • 结构检测:
    通过检查页面中是否存在目标表格的标志性元素(比如表格的id、class,或特定表头文本),或是异常结构的特征(比如包含错误提示的脚本内容)来区分结构。示例代码:

    import requests
    from bs4 import BeautifulSoup
    
    url = "https://www.ibm.com/docs/en/imdm/12.0?topic=t-accessdateval"
    # 模拟浏览器请求头,降低异常结构出现概率
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36",
        "Accept-Language": "zh-CN,zh;q=0.9"
    }
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, "html.parser")
    
    # 检测异常结构:替换成实际异常脚本中的关键词
    is_exception = False
    for script in soup.find_all("script"):
        if script.string and "错误提示翻译" in script.string:
            is_exception = True
            break
    
    # 检测目标表格:替换成实际表格的class/id
    target_table = soup.find("table", class_="ibm-table")
    
  • 分支处理逻辑:

    • 若为异常结构:直接跳过或记录日志后继续执行后续任务;
    • 若为预期结构:执行表格提取和CSV保存,同时用try-except包裹核心代码,处理提取时的意外错误。
      示例代码:
    import csv
    
    if is_exception:
        print(f"页面 {url} 为异常结构,跳过")
        with open("crawl_error.log", "a", encoding="utf-8") as f:
            f.write(f"{url} - 异常结构\n")
    elif target_table:
        try:
            # 提取表头
            headers = [th.get_text(strip=True) for th in target_table.find_all("th")]
            # 提取表格内容行
            rows = []
            for tr in target_table.find_all("tr")[1:]:
                row_data = [td.get_text(strip=True) for td in tr.find_all("td")]
                rows.append(row_data)
            
            # 追加写入CSV
            with open("table_data.csv", "a", newline="", encoding="utf-8") as csvfile:
                writer = csv.writer(csvfile)
                # 仅当文件为空时写入表头
                if csvfile.tell() == 0:
                    writer.writerow(headers)
                writer.writerows(rows)
            print(f"成功处理页面 {url}")
        except Exception as e:
            print(f"处理页面 {url} 出错: {str(e)}")
            with open("crawl_error.log", "a", encoding="utf-8") as f:
                f.write(f"{url} - 提取错误: {str(e)}\n")
    else:
        print(f"页面 {url} 未找到目标表格,跳过")
        with open("crawl_error.log", "a", encoding="utf-8") as f:
            f.write(f"{url} - 无目标表格\n")
    
  • 额外优化:

    • 加入请求重试机制:使用requests.adapters.HTTPAdapter设置重试次数,应对临时请求异常;
    • 定期清理日志文件,避免日志过大。

内容的提问来源于stack exchange,提问作者Agreeable_Math_314

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 05:42:40