You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取exam-mate网站时遭遇连接错误求助

使用Python爬取exam-mate网站时遭遇连接错误求助

看起来你遇到了网站反爬拦截或者请求模拟不够真实的问题,我来帮你分析下可能的原因和解决办法:

1. 补全请求头信息

你目前只设置了User-Agent,很多网站会校验更多请求头字段来识别爬虫,模拟更完整的浏览器请求才能降低被拦截的概率。可以把请求头改成这样:

headers = {
    "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/87.0.4280.88 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.5",
    "Referer": "https://www.exam-mate.com/",
    "Accept-Encoding": "gzip, deflate, br",
    "Connection": "keep-alive",
    "Upgrade-Insecure-Requests": "1"
}

2. 排查动态加载问题

报错里提到了XMLHttpRequest,这说明目标页面的内容很可能是通过AJAX动态加载的——直接用requests.get()只能拿到静态的HTML骨架,根本获取不到包含图片链接的核心内容。这种情况下,你需要用浏览器自动化工具模拟真实浏览行为,比如Selenium:
先安装Selenium和对应浏览器的驱动(比如ChromeDriver),然后修改爬取页面的代码:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

# 配置无头浏览器(可选,不想弹出浏览器窗口的话启用)
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument(f"user-agent={headers['user-agent']}")

driver = webdriver.Chrome(options=chrome_options)

for topic, url in topic_urls.items():
    print(f"Scraping {topic}...")
    os.makedirs(f"images/{topic}/questions", exist_ok=True)
    os.makedirs(f"images/{topic}/answers", exist_ok=True)

    try:
        driver.get(url)
        # 等待页面加载完成(可根据实际情况调整等待时间,或者用显式等待更精准)
        sleep(randint(3, 5))
        soup = BeautifulSoup(driver.page_source, "html.parser")
    except Exception as e:
        print(f"Error fetching URL {url}: {e}")
        continue

    # 后面的提取图片链接和下载逻辑保持不变...

driver.quit()

3. 调整请求策略

  • 换运行环境试试:你提到在ChatGPT控制台运行代码,这大概率有问题——ChatGPT的服务器IP可能已经被exam-mate封禁了,建议在本地环境运行代码测试。
  • 给页面请求加重试:和图片下载的重试逻辑一样,给页面请求也加上重试机制,避免单次网络波动导致失败:
    def fetch_page(url, headers, retries=3):
        for attempt in range(retries):
            try:
                response = requests.get(url, headers=headers, timeout=15)
                response.raise_for_status()
                return response
            except requests.exceptions.RequestException as e:
                print(f"Page fetch attempt {attempt+1} failed: {e}")
                sleep(randint(2, 4))
        return None
    
    然后在循环里调用这个函数替代原来的requests.get。

4. 其他注意事项

  • 适当延长请求间隔,不要过于频繁发起请求,避免触发网站的反爬阈值。
  • 提前查看目标网站的robots.txt,确认是否允许爬取这类内容,遵守网站的规则。

备注:内容来源于stack exchange,提问作者Anonymous Coder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 17:49:33