You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup+Selenium获取全部结果并规避自动化检测页面

解决FairPrice奶粉页面爬取的两个问题:加载全部结果+规避自动化检测

一、核心问题拆解

  • 仅获取24条结果:页面采用滚动加载机制,初始仅渲染第一屏内容,需主动滚动触发后续商品加载
  • 自动化检测:网站会识别Selenium默认的特征(如webdriver标识),需修改浏览器配置规避识别

二、完整解决方案代码

1. 依赖安装(按需执行)

如果需要更稳定的反检测驱动,先安装依赖:

pip install undetected-chromedriver selenium beautifulsoup4 pandas

2. 优化后的爬取代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import time
import pandas as pd
import bs4

# 配置浏览器反检测参数
chrome_options = Options()
# 禁用自动化特征提示
chrome_options.add_argument("--disable-blink-features=AutomationControlled")
# 关闭自动化扩展与提示
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)
# 设置真实用户代理(可替换成自己浏览器的UA)
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")

# 初始化驱动:二选一,undetected版本反爬能力更强
# 方案1:使用增强版反检测驱动
# import undetected_chromedriver as uc
# driver = uc.Chrome(options=chrome_options)

# 方案2:原生ChromeDriver(需配合上面的参数)
driver = webdriver.Chrome(options=chrome_options)
# 移除webdriver标识,伪装成正常浏览器
driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})")

# 访问目标页面
driver.get("https://www.fairprice.com.sg/category/milk-powder")
wait = WebDriverWait(driver, 10)

# 循环滚动加载全部商品
last_product_num = 0
while True:
    # 等待商品容器加载完成
    wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "product-container")))
    # 获取当前已加载的商品数量
    current_num = len(driver.find_elements(By.CLASS_NAME, "product-container"))
    # 数量不再变化说明加载完成
    if current_num == last_product_num:
        break
    last_product_num = current_num
    # 滚动到页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # 等待新内容加载(可根据网络速度调整时长)
    time.sleep(2)

# 解析页面内容
soup = bs4.BeautifulSoup(driver.page_source, 'lxml')
allelem = soup.find_all('div', class_='sc-1plwklf-0 iknXK product-container')

items = []
prices = []
volumes = []

for item in allelem:
    # 提取商品名称,容错处理避免报错
    name_tag = item.find('span', class_='sc-1bsd7ul-1 eJoyLL')
    items.append(name_tag.text.strip() if name_tag else "无名称")
    # 提取商品价格
    price_tag = item.find('span', class_='sc-1bsd7ul-1 sc-1svix5t-1 gJhHzP biBzHY')
    prices.append(price_tag.text.strip() if price_tag else "无价格")
    # 提取商品规格
    volume_tag = item.find('span', class_='sc-1bsd7ul-1 eeyOqy')
    volumes.append(volume_tag.text.strip() if volume_tag else "无规格")

# 生成DataFrame并导出Excel
final_data = [{'Item': i, 'Volume': v, 'Price': p} for i, v, p in zip(items, volumes, prices)]
df = pd.DataFrame(final_data)
print(df)
df.to_excel('ntucv4milk.xlsx', index=False)

# 关闭浏览器
driver.quit()

三、关键细节说明

1. 反检测核心措施

  • 关闭浏览器的自动化特征标识,移除navigator.webdriver属性,让爬虫行为更接近人工访问
  • 设置真实用户代理,避免被网站识别为异常请求
  • 优先推荐undetected-chromedriver:该库针对反爬机制做了专门优化,比原生Selenium更难被检测

2. 滚动加载逻辑

  • 通过对比滚动前后的商品数量,判断是否加载完成
  • 使用WebDriverWait等待元素加载,避免因网络延迟导致的元素查找失败
  • 滚动后设置等待时间,确保新商品有足够时间渲染

3. 容错处理

  • 提取元素时增加存在性判断,避免因个别商品的元素缺失导致代码崩溃,用默认值填充缺失数据

内容的提问来源于stack exchange,提问作者jjbkd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 13:55:37