Python使用requests模块爬虫时出现Connection Aborted Error报错
报错原因
- 代理配置不匹配+代理失效:代码中请求的是
https开头的韦氏词典页面,但传入的代理字典仅配置了http键,无HTTPS请求对应的代理规则;同时你使用的公网免费代理存活时间极短,绝大多数不支持HTTPS流量转发,连接建立后代理或目标站点直接断开连接,就会触发远程无响应断开的报错。 - 重试逻辑未生效:你仅给
http://前缀的请求挂载了重试适配器,实际所有请求都是https://开头,预设的3次重连规则完全没有触发,遇到单次连接异常直接抛出错误终止程序。 - 反爬策略触发:仅配置了单一
User-Agent字段,缺少正常浏览器访问时会携带的Accept、Accept-Language、Referer等常规请求头,且无请求间隔的高频访问很容易被站点的WAF识别为爬虫流量,直接掐断连接不返回任何响应内容。 - 代码存在基础语法疏漏:未导入
BeautifulSoup依赖,读取pickle文件的句柄f、分段变量second/third均未提前定义,即使连接正常代码也无法完整运行。
修复方法
- 补全代码基础逻辑:导入缺失的依赖,提前定义文件读取句柄、分段边界变量。
- 修正代理配置:代理字典同时配置
http和https键值;使用代理前先本地测试连通性,优先使用高匿付费代理,若本地网络可直连目标站点建议暂时移除失效的免费代理。 - 补全重试规则:同时给HTTP、HTTPS请求挂载重试适配器,把连接断开、429限流、5xx服务端错误都纳入重试范围,配置退避间隔避免频繁重连被封。
- 优化反爬绕过配置:补全常规浏览器请求头,每次请求后增加1-3秒的随机访问间隔,降低请求频率。
- 增加异常捕获:单个请求失败时记录失败词条后跳过,不要直接终止整个爬取流程,后续可对失败词条做补爬。
修正后的可运行参考代码:
import requests import pickle import time import random from requests.adapters import HTTPAdapter from requests.packages.urllib3.util.retry import Retry from bs4 import BeautifulSoup # 初始化session session = requests.Session() # 配置重试策略,覆盖连接错误、限流、服务端错误 retry = Retry( total=3, connect=3, backoff_factor=1, allowed_methods=frozenset(['GET']), status_forcelist=[429, 500, 502, 503, 504] ) adapter = HTTPAdapter(max_retries=retry) # 同时给http和https请求挂载适配器 session.mount('http://', adapter) session.mount('https://', adapter) # 补全请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:101.0) Gecko/20100101 Firefox/101.0', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'zh-CN,zh;q=0.8,zh-TW;q=0.7,zh-HK;q=0.5,en-US;q=0.3,en;q=0.2', 'Referer': 'https://www.merriam-webster.com/', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } # 代理配置示例,注意同时填http和https键,无效代理请删除后用直连测试 proxies1 = {"http": "http://23.238.33.186:80", "https": "http://23.238.33.186:80"} # proxies2、proxies3按相同规则配置 # 提前定义缺失变量 with open('your_word_list.pkl', 'rb') as f: # 替换成实际的词表文件路径 lstofwords = pickle.load(f) dic = {} failed_words = [] word_count = len(lstofwords) first = int(word_count/3) second = int(word_count*2/3) third = word_count # 第一段爬取 for x in range(0, first): word = lstofwords[x] try: r = session.get( f'https://www.merriam-webster.com/dictionary/{word}/', headers=headers, proxies=proxies1, timeout=10 ) soup = BeautifulSoup(r.content, 'html.parser') descriptions = soup.find_all('span', class_='unText') descs = [i.text for i in descriptions] dic[word] = descs # 随机延时降低被封概率 time.sleep(random.uniform(1, 3)) except Exception as e: print(f"爬取词条{word}失败: {str(e)}") failed_words.append(word) continue # 第二段爬取 for x in range(first, second): word = lstofwords[x] try: r = session.get( f'https://www.merriam-webster.com/dictionary/{word}/', headers=headers, proxies=proxies2, timeout=10 ) soup = BeautifulSoup(r.content, 'html.parser') descriptions = soup.find_all('span', class_='unText') descs = [i.text for i in descriptions] dic[word] = descs time.sleep(random.uniform(1, 3)) except Exception as e: print(f"爬取词条{word}失败: {str(e)}") failed_words.append(word) continue # 第三段爬取 for x in range(second, third): word = lstofwords[x] try: r = session.get( f'https://www.merriam-webster.com/dictionary/{word}/', headers=headers, proxies=proxies3, timeout=10 ) soup = BeautifulSoup(r.content, 'html.parser') descriptions = soup.find_all('span', class_='unText') descs = [i.text for i in descriptions] dic[word] = descs time.sleep(random.uniform(1, 3)) except Exception as e: print(f"爬取词条{word}失败: {str(e)}") failed_words.append(word) continue # 保存爬取结果 with open('dictionary.pkl', 'wb') as f: pickle.dump(dic, f) # 保存失败词条后续补爬 with open('failed_words.pkl', 'wb') as f: pickle.dump(failed_words, f) print(f"爬取完成,成功{len(dic)}个词条,失败{len(failed_words)}个词条")
注意:如果更换代理后依然出现连接断开,大概率是代理质量差或被站点封禁,可先移除proxies参数用本地网络测试连通性,再调整代理和请求频率。
内容的提问来源于stack exchange,提问作者Ehab
相关产品推荐
相关产品推荐

