You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析电商嵌套标签获取商品标题价格为空如何解决

问题排查与解决方法

常见原因

  • 页面动态渲染未完成:该电商站点商品列表为动态加载内容,你还没等Selenium把页面完全加载出来,就提前把源码传给BeautifulSoup解析,自然找不到对应元素。
  • 反爬拦截:站点识别到爬虫请求后返回验证码、错误页等非目标内容,源码里根本没有商品相关的HTML结构。
  • 类名匹配规则问题:BeautifulSoup传字符串匹配多类名时,要求类名顺序、空格完全和HTML一致,容易出现匹配失败的情况。
  • 元素内容嵌套问题:用.string提取文本时,如果a标签内有其他嵌套元素,.string会返回None,改用.get_text()更稳定。

修复方案

方案1:Selenium + 优化BeautifulSoup匹配逻辑

增加显式等待确保页面加载完成,调整元素匹配规则,代码如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

driver = webdriver.Chrome()
driver.get("https://shop-aventa.ru/search?q=+%D0%A0%D0%B0%D0%B7%D1%8A%D0%B5%D0%BC+220+")

# 显式等待10秒,直到商品卡片加载完成
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "product-card"))
    )
except Exception as e:
    print("页面加载超时:", e)
    driver.quit()
    exit()

soup = BeautifulSoup(driver.page_source, "html.parser")
# 多类名用列表匹配,规避顺序问题
productCards = soup.find_all('li', class_=["products-cards__item", "product-card"])

for productCard in productCards:
    # 提取标题
    title_ele = productCard.find('h3', class_="product-card__title")
    if title_ele and title_ele.find('a'):
        title = title_ele.find('a').get_text(strip=True)
        print("商品标题:", title)
    # 提取价格
    price_ele = productCard.find('span', class_="product-price__value")
    if price_ele:
        price = price_ele.get_text(strip=True)
        print("商品价格:", price)

driver.quit()

方案2:直接用Selenium原生定位(更稳定)

不需要经过BeautifulSoup二次解析,直接用Selenium的CSS选择器定位元素,减少兼容问题:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

driver = webdriver.Chrome()
driver.get("https://shop-aventa.ru/search?q=+%D0%A0%D0%B0%D0%B7%D1%8A%D0%B5%D0%BC+220+")

WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, "product-card"))
)

product_cards = driver.find_elements(By.CSS_SELECTOR, "li.products-cards__item.product-card")
for card in product_cards:
    try:
        title = card.find_element(By.CSS_SELECTOR, "h3.product-card__title a").text.strip()
        print("标题:", title)
    except:
        print("无有效标题")
    try:
        price = card.find_element(By.CSS_SELECTOR, "span.product-price__value").text.strip()
        print("价格:", price)
    except:
        print("无有效价格")

driver.quit()

如果运行后还是拿不到内容,先打印driver.page_source确认返回的源码是否包含商品列表的HTML,排查是否被反爬拦截,可尝试添加请求头、调整请求间隔、使用代理解决。

内容的提问来源于stack exchange,提问作者himynameissergey

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 17:06:07