You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取Quizlet网站遭Cloudflare拦截,求解决方案

问题描述

尝试爬取Quizlet链接https://quizlet.com/751002352/d2k2-die-kleidung-3-gegenteile-1-flash-cards/的标题内容,预期拿到<h1>D2K2 die Kleidung 3 & Gegenteile 1</h1>,但实际返回的是<h1>One more step…</h1>。调用find_all('h1')后得到的结果是:

[<h1>One more step…</h1>, <h1 style="color:#bd2426;">Please turn JavaScript on and reload the page.</h1>]

查看页面源码后发现,该网站受Cloudflare保护,需要完成验证码验证才能访问。当前使用的代码如下:

from bs4 import BeautifulSoup
import requests

html_text = requests.get(input('Link: ')).text
soup = BeautifulSoup(html_text, 'lxml')
flashcard_title = soup.find_all('h1')
print(flashcard_title)

请问怎么绕过这个拦截?

解决方案

方法1:用cloudscraper库绕过基础Cloudflare验证

cloudscraper是专门用来绕过Cloudflare反爬的工具库,能模拟浏览器行为通过基础验证:

  1. 先安装这个库:
pip install cloudscraper
  1. 修改代码如下:
from bs4 import BeautifulSoup
import cloudscraper

# 创建scraper实例
scraper = cloudscraper.create_scraper()
url = input('Link: ')
html_text = scraper.get(url).text
soup = BeautifulSoup(html_text, 'lxml')
# 用find直接获取单个h1即可,不用find_all
flashcard_title = soup.find('h1')
print(flashcard_title.text)

方法2:用Selenium模拟真实浏览器

如果cloudscraper失效,Selenium可以模拟带JS渲染的真实浏览器,支持手动完成验证码验证(也可以配合验证码识别工具实现自动验证):

  1. 先安装库和浏览器驱动:
pip install selenium

另外需要下载对应浏览器的驱动(比如Chrome的ChromeDriver),并把驱动路径配置到系统环境变量里。

  1. 修改代码:
from bs4 import BeautifulSoup
from selenium import webdriver
import time

url = input('Link: ')
# 初始化Chrome浏览器
driver = webdriver.Chrome()
driver.get(url)

# 这里留10秒时间让你手动完成Cloudflare的验证码验证,时间不够可以调整
time.sleep(10)

# 获取验证后的页面源码
html_text = driver.page_source
soup = BeautifulSoup(html_text, 'lxml')
flashcard_title = soup.find('h1')
print(flashcard_title.text)

# 关闭浏览器
driver.quit()

注意点

  • 频繁爬取可能触发Quizlet的反爬限制,建议添加请求间隔,遵守网站的robots.txt规则。
  • Cloudflare的验证机制会持续更新,上述方法可能随时间失效,需要根据实际情况调整。

内容的提问来源于stack exchange,提问作者Shonky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 00:10:52