使用Python爬取exam-mate网站时遭遇连接错误求助
使用Python爬取exam-mate网站时遭遇连接错误求助
看起来你遇到了网站反爬拦截或者请求模拟不够真实的问题,我来帮你分析下可能的原因和解决办法:
1. 补全请求头信息
你目前只设置了User-Agent,很多网站会校验更多请求头字段来识别爬虫,模拟更完整的浏览器请求才能降低被拦截的概率。可以把请求头改成这样:
headers = { "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/87.0.4280.88 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.5", "Referer": "https://www.exam-mate.com/", "Accept-Encoding": "gzip, deflate, br", "Connection": "keep-alive", "Upgrade-Insecure-Requests": "1" }
2. 排查动态加载问题
报错里提到了XMLHttpRequest,这说明目标页面的内容很可能是通过AJAX动态加载的——直接用requests.get()只能拿到静态的HTML骨架,根本获取不到包含图片链接的核心内容。这种情况下,你需要用浏览器自动化工具模拟真实浏览行为,比如Selenium:
先安装Selenium和对应浏览器的驱动(比如ChromeDriver),然后修改爬取页面的代码:
from selenium import webdriver from selenium.webdriver.chrome.options import Options # 配置无头浏览器(可选,不想弹出浏览器窗口的话启用) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument(f"user-agent={headers['user-agent']}") driver = webdriver.Chrome(options=chrome_options) for topic, url in topic_urls.items(): print(f"Scraping {topic}...") os.makedirs(f"images/{topic}/questions", exist_ok=True) os.makedirs(f"images/{topic}/answers", exist_ok=True) try: driver.get(url) # 等待页面加载完成(可根据实际情况调整等待时间,或者用显式等待更精准) sleep(randint(3, 5)) soup = BeautifulSoup(driver.page_source, "html.parser") except Exception as e: print(f"Error fetching URL {url}: {e}") continue # 后面的提取图片链接和下载逻辑保持不变... driver.quit()
3. 调整请求策略
- 换运行环境试试:你提到在ChatGPT控制台运行代码,这大概率有问题——ChatGPT的服务器IP可能已经被exam-mate封禁了,建议在本地环境运行代码测试。
- 给页面请求加重试:和图片下载的重试逻辑一样,给页面请求也加上重试机制,避免单次网络波动导致失败:
然后在循环里调用这个函数替代原来的def fetch_page(url, headers, retries=3): for attempt in range(retries): try: response = requests.get(url, headers=headers, timeout=15) response.raise_for_status() return response except requests.exceptions.RequestException as e: print(f"Page fetch attempt {attempt+1} failed: {e}") sleep(randint(2, 4)) return Nonerequests.get。
4. 其他注意事项
- 适当延长请求间隔,不要过于频繁发起请求,避免触发网站的反爬阈值。
- 提前查看目标网站的
robots.txt,确认是否允许爬取这类内容,遵守网站的规则。
备注:内容来源于stack exchange,提问作者Anonymous Coder
相关产品推荐
相关产品推荐

