You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用requests模块爬虫时出现Connection Aborted Error报错

报错原因
  • 代理配置不匹配+代理失效:代码中请求的是https开头的韦氏词典页面,但传入的代理字典仅配置了http键,无HTTPS请求对应的代理规则;同时你使用的公网免费代理存活时间极短,绝大多数不支持HTTPS流量转发,连接建立后代理或目标站点直接断开连接,就会触发远程无响应断开的报错。
  • 重试逻辑未生效:你仅给http://前缀的请求挂载了重试适配器,实际所有请求都是https://开头,预设的3次重连规则完全没有触发,遇到单次连接异常直接抛出错误终止程序。
  • 反爬策略触发:仅配置了单一User-Agent字段,缺少正常浏览器访问时会携带的Accept、Accept-Language、Referer等常规请求头,且无请求间隔的高频访问很容易被站点的WAF识别为爬虫流量,直接掐断连接不返回任何响应内容。
  • 代码存在基础语法疏漏:未导入BeautifulSoup依赖,读取pickle文件的句柄f、分段变量second/third均未提前定义,即使连接正常代码也无法完整运行。
修复方法
  • 补全代码基础逻辑:导入缺失的依赖,提前定义文件读取句柄、分段边界变量。
  • 修正代理配置:代理字典同时配置http和https键值;使用代理前先本地测试连通性,优先使用高匿付费代理,若本地网络可直连目标站点建议暂时移除失效的免费代理。
  • 补全重试规则:同时给HTTP、HTTPS请求挂载重试适配器,把连接断开、429限流、5xx服务端错误都纳入重试范围,配置退避间隔避免频繁重连被封。
  • 优化反爬绕过配置:补全常规浏览器请求头,每次请求后增加1-3秒的随机访问间隔,降低请求频率。
  • 增加异常捕获:单个请求失败时记录失败词条后跳过,不要直接终止整个爬取流程,后续可对失败词条做补爬。

修正后的可运行参考代码:

import requests
import pickle
import time
import random
from requests.adapters import HTTPAdapter
from requests.packages.urllib3.util.retry import Retry
from bs4 import BeautifulSoup

# 初始化session
session = requests.Session()
# 配置重试策略,覆盖连接错误、限流、服务端错误
retry = Retry(
    total=3,
    connect=3,
    backoff_factor=1,
    allowed_methods=frozenset(['GET']),
    status_forcelist=[429, 500, 502, 503, 504]
)
adapter = HTTPAdapter(max_retries=retry)
# 同时给http和https请求挂载适配器
session.mount('http://', adapter)
session.mount('https://', adapter)

# 补全请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:101.0) Gecko/20100101 Firefox/101.0',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'zh-CN,zh;q=0.8,zh-TW;q=0.7,zh-HK;q=0.5,en-US;q=0.3,en;q=0.2',
    'Referer': 'https://www.merriam-webster.com/',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

# 代理配置示例,注意同时填http和https键,无效代理请删除后用直连测试
proxies1 = {"http": "http://23.238.33.186:80", "https": "http://23.238.33.186:80"}
# proxies2、proxies3按相同规则配置

# 提前定义缺失变量
with open('your_word_list.pkl', 'rb') as f: # 替换成实际的词表文件路径
    lstofwords = pickle.load(f)
dic = {}
failed_words = []
word_count = len(lstofwords)
first = int(word_count/3)
second = int(word_count*2/3)
third = word_count

# 第一段爬取
for x in range(0, first):
    word = lstofwords[x]
    try:
        r = session.get(
            f'https://www.merriam-webster.com/dictionary/{word}/',
            headers=headers,
            proxies=proxies1,
            timeout=10
        )
        soup = BeautifulSoup(r.content, 'html.parser')
        descriptions = soup.find_all('span', class_='unText')
        descs = [i.text for i in descriptions]
        dic[word] = descs
        # 随机延时降低被封概率
        time.sleep(random.uniform(1, 3))
    except Exception as e:
        print(f"爬取词条{word}失败: {str(e)}")
        failed_words.append(word)
        continue

# 第二段爬取
for x in range(first, second):
    word = lstofwords[x]
    try:
        r = session.get(
            f'https://www.merriam-webster.com/dictionary/{word}/',
            headers=headers,
            proxies=proxies2,
            timeout=10
        )
        soup = BeautifulSoup(r.content, 'html.parser')
        descriptions = soup.find_all('span', class_='unText')
        descs = [i.text for i in descriptions]
        dic[word] = descs
        time.sleep(random.uniform(1, 3))
    except Exception as e:
        print(f"爬取词条{word}失败: {str(e)}")
        failed_words.append(word)
        continue

# 第三段爬取
for x in range(second, third):
    word = lstofwords[x]
    try:
        r = session.get(
            f'https://www.merriam-webster.com/dictionary/{word}/',
            headers=headers,
            proxies=proxies3,
            timeout=10
        )
        soup = BeautifulSoup(r.content, 'html.parser')
        descriptions = soup.find_all('span', class_='unText')
        descs = [i.text for i in descriptions]
        dic[word] = descs
        time.sleep(random.uniform(1, 3))
    except Exception as e:
        print(f"爬取词条{word}失败: {str(e)}")
        failed_words.append(word)
        continue

# 保存爬取结果
with open('dictionary.pkl', 'wb') as f:
    pickle.dump(dic, f)
# 保存失败词条后续补爬
with open('failed_words.pkl', 'wb') as f:
    pickle.dump(failed_words, f)
print(f"爬取完成,成功{len(dic)}个词条,失败{len(failed_words)}个词条")

注意:如果更换代理后依然出现连接断开,大概率是代理质量差或被站点封禁,可先移除proxies参数用本地网络测试连通性,再调整代理和请求频率。

内容的提问来源于stack exchange,提问作者Ehab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 12:31:01