You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取网站返回Cloudflare链接,求解决方案

问题

我想用Python的requests和BeautifulSoup爬取不同网站,提取所有带href属性的a标签链接并逐行打印。比如对于<a href="https://www.google.com">some text</a>,应返回https://www.google.com,但爬取https://www.swappa.com时,只返回Cloudflare验证链接。

代码示例

import requests
from bs4 import BeautifulSoup
from selenium import webdriver
from webdriver_manager.chrome import ChromeDriverManager

page = requests.get("https://www.swappa.com")

soup = BeautifulSoup(page.content, "html.parser")

links = soup.find_all("a", href=True)

for link in links:
    print(link['href'])

输出结果

https://www.cloudflare.com/?utm_source=challenge&utm_campaign=j
解决办法

方法1:用Selenium模拟浏览器访问

直接用requests请求会被Cloudflare反爬机制拦截,你已经导入了Selenium相关库,改用浏览器模拟真实访问行为即可绕过验证:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
import time

# 初始化Chrome浏览器
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
driver.get("https://www.swappa.com")

# 等待页面加载完成(可根据实际情况调整等待时长)
time.sleep(3)

# 获取加载后的页面源码
page_source = driver.page_source
soup = BeautifulSoup(page_source, "html.parser")

# 提取并打印所有带href的a标签链接
links = soup.find_all("a", href=True)
for link in links:
    print(link['href'])

# 关闭浏览器
driver.quit()

方法2:用cfscrape库绕过Cloudflare验证

如果不想依赖浏览器,可以用专门针对Cloudflare的cfscrape库自动处理验证:

  1. 先安装库:
pip install cfscrape
  1. 修改代码:
import cfscrape
from bs4 import BeautifulSoup

# 创建scraper对象
scraper = cfscrape.create_scraper()
page = scraper.get("https://www.swappa.com")

soup = BeautifulSoup(page.content, "html.parser")
links = soup.find_all("a", href=True)

for link in links:
    print(link['href'])

注意:cfscrape对部分新版Cloudflare验证可能失效,此时优先选择Selenium方法。

内容的提问来源于stack exchange,提问作者Arsalan Sasan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 00:15:50