You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium爬取Nguyen Kim长名称商品时无法获取名称和价格

解决Selenium爬取Nguyen Kim网站长名称商品信息为空的问题

问题场景

爬取Nguyen Kim搜索页(https://www.nguyenkim.com/tim-kiem.html?tu-khoa=may+tinh)的商品信息时,短名称商品能正常返回完整的名称、价格、链接和图片,但长名称商品(如"Máy tính để bàn HP 205 Pro G4 AIO R5-4500U/8GB/256GB/Win10 31Y21PA")的名称和价格会返回空字符串,仅链接和图片正常。

正常返回示例:

Item(
    link = 'https://www.nguyenkim.com/chuot-logitech-m100r.html',
    name = 'Chuột máy tính Logitech M100R Đen',
    current_price = '109.000đ',
    place = 'Nguyen Kim',
    img = 'https://cdn.nguyenkimmall.com/images/thumbnails/210/210/detailed/177/10026584-chuot-logitech-m100r-den-1.jpg'
)

长名称商品返回示例:

Item(
    link = 'https://www.nguyenkim.com/may-tinh-bang-xiaomi-redmi-pad-64gb-xam.html',
    name = '',
    current_price = '',
    place = 'Nguyen Kim',
    img = 'https://cdn.nguyenkimmall.com/images/thumbnails/210/210/detailed/847/10053972-may-tinh-bang-xiaomi-redmi-pad-64gb-xam.jpg'
)

原代码片段:

content = driver.find_element(By.CLASS_NAME, 'result-wrapper')
items = content.find_elements(By.CLASS_NAME, 'product')
for _ in items:
    item = Item(
      link = _.find_element(By.CSS_SELECTOR, "div[class*='product-header']").get_attribute('href'),
      name = _.find_element(By.CSS_SELECTOR, "div.product-title a").text,
      current_price = _.find_element(By.CSS_SELECTOR, "p[class*='final-price']").text,
      place = "Nguyen Kim",
      img = _.find_element(By.CSS_SELECTOR, "img").get_attribute('src')
    )
    print(item)

问题原因

网站对长名称商品采用了CSS文本截断(如overflow: hidden),导致Selenium的.text方法只能获取可视区域内的文本,而截断后的可视文本为空;同时价格元素可能因布局原因,可视文本也被隐藏或未正确渲染。此外,电商网站通常会把完整商品名称存储在链接的title属性中,而非仅依赖可视文本。

解决方案

  • 商品名称:放弃使用.text,改为获取商品链接的title属性,该属性通常存储完整未截断的商品名称。
  • 商品价格:使用get_attribute('innerText')替代.text,innerText能获取元素内的所有文本内容(包括被CSS截断的部分);若网站将价格存储在自定义属性(如data-price)中,也可直接读取该属性并格式化。

修改后的代码

content = driver.find_element(By.CLASS_NAME, 'result-wrapper')
items = content.find_elements(By.CLASS_NAME, 'product')
for _ in items:
    # 获取完整商品名称(优先取title属性,兼容短名称商品)
    name_elem = _.find_element(By.CSS_SELECTOR, "div.product-title a")
    name = name_elem.get_attribute('title') or name_elem.text
    
    # 获取价格(用innerText避免截断问题)
    price_elem = _.find_element(By.CSS_SELECTOR, "p[class*='final-price']")
    current_price = price_elem.get_attribute('innerText') or price_elem.text
    
    item = Item(
      link = _.find_element(By.CSS_SELECTOR, "div[class*='product-header']").get_attribute('href'),
      name = name.strip(),
      current_price = current_price.strip(),
      place = "Nguyen Kim",
      img = _.find_element(By.CSS_SELECTOR, "img").get_attribute('src')
    )
    print(item)

补充说明

  • 添加or判断是为了兼容短名称商品(部分短名称商品的title属性可能与可视文本一致,或未设置title)。
  • 若innerText仍无法获取价格,可检查价格元素的HTML结构,例如是否存在data-price属性,此时可改为price_elem.get_attribute('data-price') + 'đ'进行格式化。

内容的提问来源于stack exchange,提问作者Dung8466

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 11:22:46