使用Requests与BeautifulSoup4爬取剑桥词典遇ConnectionError求助
问题描述
我编写了网页爬虫代码,尝试获取剑桥词典网站的HTML代码,但运行时出现错误。错误发生在html = requests.get(url, headers=headers).text行,错误信息为:
Exception has occurred: ConnectionError
('Connection aborted.', RemoteDisconnected('Remote end closed connection without response'))
以下是我的代码:
import requests from bs4 import BeautifulSoup def checkWord(word): url_top = "https://dictionary.cambridge.org/dictionary/english/" url = url_top + word headers = requests.utils.default_headers() headers.update( { 'User-Agent': 'My User Agent 1.0', } ) html = requests.get(url, headers=headers).text soup = BeautifulSoup(html, 'html.parser') check = soup.find("title") boolean = check.string if boolean == "Cambridge English Dictionary: Meanings & Definitions": return False else: return True word = "App" checkWord(word)
错误原因分析
- 反爬机制拦截:剑桥词典服务器识别到请求来自爬虫,直接断开连接。你设置的
User-Agent过于简单,不符合真实浏览器的请求头格式,极易被识别。 - 缺少必要请求头:真实浏览器请求会携带
Accept-Language、Accept-Encoding等字段,缺少这些会被判定为异常请求。 - 网络波动(概率低):少数情况下,网络不稳定可能导致连接意外断开,但优先考虑反爬拦截问题。
解决方案
1. 完善请求头信息
使用贴近真实浏览器的User-Agent,补充必要请求头字段:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'Accept-Encoding': 'gzip, deflate, br', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8' }
2. 添加重试与异常处理
针对临时拦截或网络波动,添加请求重试逻辑,并处理可能的异常:
import requests from bs4 import BeautifulSoup from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry def checkWord(word): url_top = "https://dictionary.cambridge.org/dictionary/english/" url = url_top + word headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'Accept-Encoding': 'gzip, deflate, br', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8' } # 创建会话并设置重试策略 session = requests.Session() retry = Retry(total=3, backoff_factor=1, status_forcelist=[429, 500, 502, 503, 504]) adapter = HTTPAdapter(max_retries=retry) session.mount('https://', adapter) try: response = session.get(url, headers=headers, timeout=10) response.raise_for_status() # 抛出HTTP错误状态码异常 html = response.text except requests.exceptions.RequestException as e: print(f"请求失败: {e}") return False soup = BeautifulSoup(html, 'html.parser') check = soup.find("title") if not check or not check.string: return False return check.string != "Cambridge English Dictionary: Meanings & Definitions" word = "App" print(checkWord(word))
3. 遵守网站爬取规则
爬取前查看剑桥词典的robots.txt,确认允许爬取的路径,避免访问受限内容,降低被拦截概率。
内容的提问来源于stack exchange,提问作者わんわんワコワコ
相关产品推荐
相关产品推荐

