使用BeautifulSoup爬取Quizlet网站遭Cloudflare拦截,求解决方案
问题描述
尝试爬取Quizlet链接https://quizlet.com/751002352/d2k2-die-kleidung-3-gegenteile-1-flash-cards/的标题内容,预期拿到<h1>D2K2 die Kleidung 3 & Gegenteile 1</h1>,但实际返回的是<h1>One more step…</h1>。调用find_all('h1')后得到的结果是:
[<h1>One more step…</h1>, <h1 style="color:#bd2426;">Please turn JavaScript on and reload the page.</h1>]
查看页面源码后发现,该网站受Cloudflare保护,需要完成验证码验证才能访问。当前使用的代码如下:
from bs4 import BeautifulSoup import requests html_text = requests.get(input('Link: ')).text soup = BeautifulSoup(html_text, 'lxml') flashcard_title = soup.find_all('h1') print(flashcard_title)
请问怎么绕过这个拦截?
解决方案
方法1:用cloudscraper库绕过基础Cloudflare验证
cloudscraper是专门用来绕过Cloudflare反爬的工具库,能模拟浏览器行为通过基础验证:
- 先安装这个库:
pip install cloudscraper
- 修改代码如下:
from bs4 import BeautifulSoup import cloudscraper # 创建scraper实例 scraper = cloudscraper.create_scraper() url = input('Link: ') html_text = scraper.get(url).text soup = BeautifulSoup(html_text, 'lxml') # 用find直接获取单个h1即可,不用find_all flashcard_title = soup.find('h1') print(flashcard_title.text)
方法2:用Selenium模拟真实浏览器
如果cloudscraper失效,Selenium可以模拟带JS渲染的真实浏览器,支持手动完成验证码验证(也可以配合验证码识别工具实现自动验证):
- 先安装库和浏览器驱动:
pip install selenium
另外需要下载对应浏览器的驱动(比如Chrome的ChromeDriver),并把驱动路径配置到系统环境变量里。
- 修改代码:
from bs4 import BeautifulSoup from selenium import webdriver import time url = input('Link: ') # 初始化Chrome浏览器 driver = webdriver.Chrome() driver.get(url) # 这里留10秒时间让你手动完成Cloudflare的验证码验证,时间不够可以调整 time.sleep(10) # 获取验证后的页面源码 html_text = driver.page_source soup = BeautifulSoup(html_text, 'lxml') flashcard_title = soup.find('h1') print(flashcard_title.text) # 关闭浏览器 driver.quit()
注意点
- 频繁爬取可能触发Quizlet的反爬限制,建议添加请求间隔,遵守网站的
robots.txt规则。 - Cloudflare的验证机制会持续更新,上述方法可能随时间失效,需要根据实际情况调整。
内容的提问来源于stack exchange,提问作者Shonky
相关产品推荐
相关产品推荐

