使用Selenium爬取Sayurbox与Segari站点返回403错误求助
网页抓取403错误排查与解决方案
问题描述
尝试抓取https://www.sayurbox.com/或https://segari.id/products内容时,返回<h1>403 ERROR</h1>,怀疑与Sayurbox网站弹出的验证表单(对应页面:https://www.sayurbox.com/confirmation?redirectTo=%2F)有关,使用的代码如下:
!pip install chromedriver-autoinstaller import sys sys.path.insert(0,'/usr/lib/chromium-browser/chromedriver') import time import pandas as pd from bs4 import BeautifulSoup from selenium import webdriver import chromedriver_autoinstaller # setup chrome options chrome_options = webdriver.ChromeOptions() chrome_options.add_argument('--headless') # ensure GUI is off chrome_options.add_argument('--no-sandbox') chrome_options.add_argument("--disable-gpu") chrome_options.add_argument('--disable-dev-shm-usage') chrome_options.add_argument('ignore-certificate-errors') # set path to chromedriver as per your configuration chromedriver_autoinstaller.install() driver = webdriver.Chrome(options=chrome_options) # Access a URL (replace with your URL) url = "https://segari.id/p/timun-lalap" # Make sure to include the protocol (https://) driver.get(url) # Get the page content or perform other actions as needed page_content = driver.page_source print(page_content) # Close the webdriver driver.quit()
原因分析
- 无头浏览器特征被识别:默认
--headless模式下,Chrome的User-Agent会带有"HeadlessChrome"标识,网站反爬机制会直接判定为爬虫,返回403。 - 未处理验证流程:网站弹出的验证表单是人机校验环节,当前代码没有任何处理验证的逻辑,无法通过网站的拦截。
解决步骤
1. 伪装浏览器特征,绕过反爬检测
修改Chrome配置,使用新版无头模式并模拟真实用户的浏览器信息:
chrome_options = webdriver.ChromeOptions() # 使用新版无头模式,行为更接近真实浏览器 chrome_options.add_argument('--headless=new') chrome_options.add_argument('--no-sandbox') chrome_options.add_argument("--disable-gpu") chrome_options.add_argument('--disable-dev-shm-usage') chrome_options.add_argument('ignore-certificate-errors') # 添加真实的User-Agent(可从自己常用浏览器中复制) chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36') # 禁用自动化特征标记,避免被识别为爬虫 chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False)
2. 处理网站验证表单
添加等待逻辑,识别并完成验证操作(需根据实际页面的验证元素调整定位方式):
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver.get(url) # 等待验证元素加载并执行验证(示例,需根据实际页面修改) try: # 等待验证按钮可点击并点击 verify_button = WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "验证")]')) ) verify_button.click() # 等待验证完成后页面跳转 time.sleep(4) except Exception as e: print("验证处理失败:", str(e)) page_content = driver.page_source
3. 模拟用户浏览节奏
在请求之间添加随机延迟,避免被判定为高频爬虫:
import random # 随机等待2-6秒,模拟用户停留时间 time.sleep(random.uniform(2, 6))
内容的提问来源于stack exchange,提问作者Hal
相关产品推荐
相关产品推荐

