You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup查找已存在的average-rating元素返回空值问题

问题描述

使用Selenium启动Chrome驱动访问指定flaconi商品页面后,通过BeautifulSoup解析获取到的page_source,调用find方法查找class为average-rating的div元素无返回结果,确认目标页面存在对应元素,复现代码如下:

from bs4 import BeautifulSoup  
import pandas as pd
from selenium import webdriver
driver = webdriver.Chrome()

driver.get('https://www.flaconi.de/haare/maria-nila/head-and-hair-heal/maria-nila-head-and-hair-heal-haarshampoo.html#sku=80021856-100')
soup = BeautifulSoup(driver.page_source,'html.parser')


soup.find('div', class_ = 'average-rating')
故障原因
  • 动态内容未加载完成:driver.get()仅等待页面初始框架加载完成,评分模块属于异步拉取渲染的动态内容,代码在元素渲染到DOM前就抓取了页面源码,自然找不到目标节点。
  • 解析器容错性不足:默认使用的html.parser对部分动态生成的非标准DOM结构解析兼容性差,可能出现节点漏匹配的情况。
  • 反爬机制拦截:默认启动的Chrome驱动带有明显的自动化特征,容易被站点反爬策略识别,返回的页面源码本身就剔除了评分等正常用户可见的内容。
修复方案
  1. 配置Chrome启动选项,添加常规请求头、关闭自动化标识,绕过基础反爬检测。
  2. 引入Selenium显式等待逻辑,等待目标评分元素加载完成后,再获取页面源码。
  3. 将BeautifulSoup的解析器替换为容错性更强的lxml(需提前执行pip install lxml安装依赖)。

修复后可正常运行的代码如下:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.chrome.options import Options

# 配置Chrome启动参数
chrome_options = Options()
chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
chrome_options.add_argument('--disable-blink-features=AutomationControlled')
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)

driver = webdriver.Chrome(options=chrome_options)
# 覆盖浏览器自动化特征属性
driver.execute_cdp_cmd('Page.addScriptToEvaluateOnNewDocument', {
    'source': 'Object.defineProperty(navigator, "webdriver", {get: () => undefined})'
})

driver.get('https://www.flaconi.de/haare/maria-nila/head-and-hair-heal/maria-nila-head-and-hair-heal-haarshampoo.html#sku=80021856-100')
# 显式等待最多10秒,直到目标评分元素加载完成
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, 'average-rating'))
)
# 元素加载完成后再抓取源码,使用lxml解析器提升兼容性
soup = BeautifulSoup(driver.page_source, 'lxml')
rating_node = soup.find('div', class_='average-rating')
print(rating_node)

driver.quit()

内容的提问来源于stack exchange,提问作者Bill Lin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.02 07:06:34