如何用BeautifulSoup+Selenium获取全部结果并规避自动化检测页面
解决FairPrice奶粉页面爬取的两个问题:加载全部结果+规避自动化检测
一、核心问题拆解
- 仅获取24条结果:页面采用滚动加载机制,初始仅渲染第一屏内容,需主动滚动触发后续商品加载
- 自动化检测:网站会识别Selenium默认的特征(如
webdriver标识),需修改浏览器配置规避识别
二、完整解决方案代码
1. 依赖安装(按需执行)
如果需要更稳定的反检测驱动,先安装依赖:
pip install undetected-chromedriver selenium beautifulsoup4 pandas
2. 优化后的爬取代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By import time import pandas as pd import bs4 # 配置浏览器反检测参数 chrome_options = Options() # 禁用自动化特征提示 chrome_options.add_argument("--disable-blink-features=AutomationControlled") # 关闭自动化扩展与提示 chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) # 设置真实用户代理(可替换成自己浏览器的UA) chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 初始化驱动:二选一,undetected版本反爬能力更强 # 方案1:使用增强版反检测驱动 # import undetected_chromedriver as uc # driver = uc.Chrome(options=chrome_options) # 方案2:原生ChromeDriver(需配合上面的参数) driver = webdriver.Chrome(options=chrome_options) # 移除webdriver标识,伪装成正常浏览器 driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") # 访问目标页面 driver.get("https://www.fairprice.com.sg/category/milk-powder") wait = WebDriverWait(driver, 10) # 循环滚动加载全部商品 last_product_num = 0 while True: # 等待商品容器加载完成 wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "product-container"))) # 获取当前已加载的商品数量 current_num = len(driver.find_elements(By.CLASS_NAME, "product-container")) # 数量不再变化说明加载完成 if current_num == last_product_num: break last_product_num = current_num # 滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待新内容加载(可根据网络速度调整时长) time.sleep(2) # 解析页面内容 soup = bs4.BeautifulSoup(driver.page_source, 'lxml') allelem = soup.find_all('div', class_='sc-1plwklf-0 iknXK product-container') items = [] prices = [] volumes = [] for item in allelem: # 提取商品名称,容错处理避免报错 name_tag = item.find('span', class_='sc-1bsd7ul-1 eJoyLL') items.append(name_tag.text.strip() if name_tag else "无名称") # 提取商品价格 price_tag = item.find('span', class_='sc-1bsd7ul-1 sc-1svix5t-1 gJhHzP biBzHY') prices.append(price_tag.text.strip() if price_tag else "无价格") # 提取商品规格 volume_tag = item.find('span', class_='sc-1bsd7ul-1 eeyOqy') volumes.append(volume_tag.text.strip() if volume_tag else "无规格") # 生成DataFrame并导出Excel final_data = [{'Item': i, 'Volume': v, 'Price': p} for i, v, p in zip(items, volumes, prices)] df = pd.DataFrame(final_data) print(df) df.to_excel('ntucv4milk.xlsx', index=False) # 关闭浏览器 driver.quit()
三、关键细节说明
1. 反检测核心措施
- 关闭浏览器的自动化特征标识,移除
navigator.webdriver属性,让爬虫行为更接近人工访问 - 设置真实用户代理,避免被网站识别为异常请求
- 优先推荐
undetected-chromedriver:该库针对反爬机制做了专门优化,比原生Selenium更难被检测
2. 滚动加载逻辑
- 通过对比滚动前后的商品数量,判断是否加载完成
- 使用
WebDriverWait等待元素加载,避免因网络延迟导致的元素查找失败 - 滚动后设置等待时间,确保新商品有足够时间渲染
3. 容错处理
- 提取元素时增加存在性判断,避免因个别商品的元素缺失导致代码崩溃,用默认值填充缺失数据
内容的提问来源于stack exchange,提问作者jjbkd
相关产品推荐
相关产品推荐

