You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取Sayurbox与Segari站点返回403错误求助

网页抓取403错误排查与解决方案

问题描述

尝试抓取https://www.sayurbox.com/或https://segari.id/products内容时,返回<h1>403 ERROR</h1>,怀疑与Sayurbox网站弹出的验证表单(对应页面:https://www.sayurbox.com/confirmation?redirectTo=%2F)有关,使用的代码如下:

!pip install chromedriver-autoinstaller

import sys
sys.path.insert(0,'/usr/lib/chromium-browser/chromedriver')

import time
import pandas as pd
from bs4 import BeautifulSoup
from selenium import webdriver
import chromedriver_autoinstaller

# setup chrome options
chrome_options = webdriver.ChromeOptions()
chrome_options.add_argument('--headless') # ensure GUI is off
chrome_options.add_argument('--no-sandbox')
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument('--disable-dev-shm-usage')
chrome_options.add_argument('ignore-certificate-errors')

# set path to chromedriver as per your configuration
chromedriver_autoinstaller.install()

driver = webdriver.Chrome(options=chrome_options)
# Access a URL (replace with your URL)
url = "https://segari.id/p/timun-lalap"  # Make sure to include the protocol (https://)
driver.get(url)

# Get the page content or perform other actions as needed
page_content = driver.page_source
print(page_content)

# Close the webdriver
driver.quit()

原因分析

  1. 无头浏览器特征被识别:默认--headless模式下,Chrome的User-Agent会带有"HeadlessChrome"标识,网站反爬机制会直接判定为爬虫,返回403。
  2. 未处理验证流程:网站弹出的验证表单是人机校验环节,当前代码没有任何处理验证的逻辑,无法通过网站的拦截。

解决步骤

1. 伪装浏览器特征,绕过反爬检测

修改Chrome配置,使用新版无头模式并模拟真实用户的浏览器信息:

chrome_options = webdriver.ChromeOptions()
# 使用新版无头模式,行为更接近真实浏览器
chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--no-sandbox')
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument('--disable-dev-shm-usage')
chrome_options.add_argument('ignore-certificate-errors')
# 添加真实的User-Agent(可从自己常用浏览器中复制)
chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36')
# 禁用自动化特征标记,避免被识别为爬虫
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)

2. 处理网站验证表单

添加等待逻辑,识别并完成验证操作(需根据实际页面的验证元素调整定位方式):

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver.get(url)

# 等待验证元素加载并执行验证(示例,需根据实际页面修改)
try:
    # 等待验证按钮可点击并点击
    verify_button = WebDriverWait(driver, 15).until(
        EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "验证")]'))
    )
    verify_button.click()
    # 等待验证完成后页面跳转
    time.sleep(4)
except Exception as e:
    print("验证处理失败:", str(e))

page_content = driver.page_source

3. 模拟用户浏览节奏

在请求之间添加随机延迟,避免被判定为高频爬虫:

import random
# 随机等待2-6秒,模拟用户停留时间
time.sleep(random.uniform(2, 6))

内容的提问来源于stack exchange,提问作者Hal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 02:50:20