You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Requests与BeautifulSoup4爬取剑桥词典遇ConnectionError求助

问题描述

我编写了网页爬虫代码,尝试获取剑桥词典网站的HTML代码,但运行时出现错误。错误发生在html = requests.get(url, headers=headers).text行,错误信息为:

Exception has occurred: ConnectionError
('Connection aborted.', RemoteDisconnected('Remote end closed connection without response'))

以下是我的代码:

import requests
from bs4 import BeautifulSoup


def checkWord(word):
    url_top = "https://dictionary.cambridge.org/dictionary/english/"
    url = url_top + word

    headers = requests.utils.default_headers()

    headers.update(
        {
            'User-Agent': 'My User Agent 1.0',
        }       
    )

    html = requests.get(url, headers=headers).text 
    soup = BeautifulSoup(html, 'html.parser') 
    check = soup.find("title")
    boolean = check.string

    
    if boolean == "Cambridge English Dictionary: Meanings & Definitions":
        return False
    else:
        return True

word = "App"
checkWord(word)
错误原因分析
  • 反爬机制拦截:剑桥词典服务器识别到请求来自爬虫,直接断开连接。你设置的User-Agent过于简单,不符合真实浏览器的请求头格式,极易被识别。
  • 缺少必要请求头:真实浏览器请求会携带Accept-Language、Accept-Encoding等字段,缺少这些会被判定为异常请求。
  • 网络波动(概率低):少数情况下,网络不稳定可能导致连接意外断开,但优先考虑反爬拦截问题。
解决方案

1. 完善请求头信息

使用贴近真实浏览器的User-Agent,补充必要请求头字段:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept-Language': 'en-US,en;q=0.9',
    'Accept-Encoding': 'gzip, deflate, br',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8'
}

2. 添加重试与异常处理

针对临时拦截或网络波动,添加请求重试逻辑,并处理可能的异常:

import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

def checkWord(word):
    url_top = "https://dictionary.cambridge.org/dictionary/english/"
    url = url_top + word

    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept-Language': 'en-US,en;q=0.9',
        'Accept-Encoding': 'gzip, deflate, br',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8'
    }

    # 创建会话并设置重试策略
    session = requests.Session()
    retry = Retry(total=3, backoff_factor=1, status_forcelist=[429, 500, 502, 503, 504])
    adapter = HTTPAdapter(max_retries=retry)
    session.mount('https://', adapter)
    
    try:
        response = session.get(url, headers=headers, timeout=10)
        response.raise_for_status()  # 抛出HTTP错误状态码异常
        html = response.text
    except requests.exceptions.RequestException as e:
        print(f"请求失败: {e}")
        return False

    soup = BeautifulSoup(html, 'html.parser') 
    check = soup.find("title")
    if not check or not check.string:
        return False
    
    return check.string != "Cambridge English Dictionary: Meanings & Definitions"

word = "App"
print(checkWord(word))

3. 遵守网站爬取规则

爬取前查看剑桥词典的robots.txt,确认允许爬取的路径,避免访问受限内容,降低被拦截概率。

内容的提问来源于stack exchange,提问作者わんわんワコワコ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 07:00:57