使用BeautifulSoup爬取网站返回Cloudflare链接,求解决方案
问题
我想用Python的requests和BeautifulSoup爬取不同网站,提取所有带href属性的a标签链接并逐行打印。比如对于<a href="https://www.google.com">some text</a>,应返回https://www.google.com,但爬取https://www.swappa.com时,只返回Cloudflare验证链接。
代码示例
import requests from bs4 import BeautifulSoup from selenium import webdriver from webdriver_manager.chrome import ChromeDriverManager page = requests.get("https://www.swappa.com") soup = BeautifulSoup(page.content, "html.parser") links = soup.find_all("a", href=True) for link in links: print(link['href'])
输出结果
https://www.cloudflare.com/?utm_source=challenge&utm_campaign=j
解决办法
方法1:用Selenium模拟浏览器访问
直接用requests请求会被Cloudflare反爬机制拦截,你已经导入了Selenium相关库,改用浏览器模拟真实访问行为即可绕过验证:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager import time # 初始化Chrome浏览器 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get("https://www.swappa.com") # 等待页面加载完成(可根据实际情况调整等待时长) time.sleep(3) # 获取加载后的页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") # 提取并打印所有带href的a标签链接 links = soup.find_all("a", href=True) for link in links: print(link['href']) # 关闭浏览器 driver.quit()
方法2:用cfscrape库绕过Cloudflare验证
如果不想依赖浏览器,可以用专门针对Cloudflare的cfscrape库自动处理验证:
- 先安装库:
pip install cfscrape
- 修改代码:
import cfscrape from bs4 import BeautifulSoup # 创建scraper对象 scraper = cfscrape.create_scraper() page = scraper.get("https://www.swappa.com") soup = BeautifulSoup(page.content, "html.parser") links = soup.find_all("a", href=True) for link in links: print(link['href'])
注意:cfscrape对部分新版Cloudflare验证可能失效,此时优先选择Selenium方法。
内容的提问来源于stack exchange,提问作者Arsalan Sasan
相关产品推荐
相关产品推荐

