You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网站无响应时Webscraping爬虫程序卡顿,如何加载默认数据?

解决爬虫遇网站卡顿无响应的问题

你的核心问题是网站卡顿导致爬虫卡住无法继续,且超时处理未生效,可以从以下几个方面修复:

1. 修正表单选择器的语法错误

你代码里的表单选择器少了闭合引号:

browser.select_form('form[action="https://myclimate.com]')

应该改为:

browser.select_form('form[action="https://myclimate.com"]')

这个错误可能导致每次循环都无法正确定位表单,间接加重请求异常的概率。

2. 强化超时与异常捕获

仅捕获requests.Timeout不够,网站卡顿可能引发多种请求异常(比如连接超时、读取超时、服务器无响应等),需要扩大捕获范围,同时确保每个请求的超时设置生效。另外,mechanicalsoup的submit_selected可以通过timeout参数单独指定超时时间,覆盖session的全局设置。

3. 增加重试机制(可选)

对于偶尔卡顿的日期,可以尝试重试1-2次,再失败就用默认数据,避免直接跳过。

修改后的完整代码

import re
import pandas as pd
import requests
from bs4 import BeautifulSoup
import mechanicalsoup
import time

# 初始化浏览器,设置全局超时
browser = mechanicalsoup.StatefulBrowser()
browser.session.timeout = (10, 20)  # 连接超时10s,读取超时20s

url = "https://myclimate.com"
browser.open(url)

dateList = pd.date_range(start='2023-01-01', end='2023-02-02').strftime("%Y-%m-%d").tolist()
df = pd.DataFrame()

print("----开始爬取----")

for dte in dateList:
    print(f"正在处理日期: {dte}")
    retry_count = 0
    max_retries = 2
    success = False
    
    while retry_count <= max_retries and not success:
        try:
            # 重新选择表单(确保每次循环都正确定位)
            browser.select_form('form[action="https://myclimate.com"]')
            browser['date_on'] = dte
            
            # 提交表单时单独指定超时,确保生效
            response = browser.submit_selected(timeout=20)
            
            # 你的爬取逻辑,填充df
            # 示例:假设解析页面获取数据
            soup = BeautifulSoup(response.text, 'html.parser')
            # ... 你的数据提取代码 ...
            
            success = True
            print(f"日期 {dte} 爬取成功")
            
        except requests.exceptions.RequestException as e:
            retry_count += 1
            if retry_count > max_retries:
                print(f"日期 {dte} 请求失败(已重试{max_retries}次): {str(e)},使用默认数据")
                # 填充默认数据到df,示例:
                default_row = pd.DataFrame({'date': [dte], 'temp': [None], 'humidity': [None]})  # 根据你的字段调整
                df = pd.concat([df, default_row], ignore_index=True)
            else:
                print(f"日期 {dte} 请求超时,正在重试第{retry_count}次...")
                time.sleep(2)
        except Exception as e:
            print(f"日期 {dte} 处理出错: {str(e)},使用默认数据")
            default_row = pd.DataFrame({'date': [dte], 'temp': [None], 'humidity': [None]})
            df = pd.concat([df, default_row], ignore_index=True)
            success = True
    
    # 爬取间隔,避免频繁请求被封
    time.sleep(1)

df.to_csv('data_year_J2023.csv', index=False)
print("----爬取完成,数据已保存----")

关键改动说明

  • 修复了表单选择器的语法错误,确保每次都能正确定位表单
  • 增加重试机制,对超时的请求尝试2次后再使用默认数据
  • 扩大异常捕获范围,覆盖所有请求类异常和未知异常,避免程序中途崩溃
  • 给submit_selected单独设置超时时间,确保超时触发时能及时中断请求
  • 添加了爬取间隔,减轻目标网站服务器压力,降低被封禁风险
  • 明确了默认数据的填充逻辑,你可以根据实际字段调整default_row的内容

内容的提问来源于stack exchange,提问作者user3462318

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 14:26:14