Python Selenium爬取Nguyen Kim长名称商品时无法获取名称和价格
解决Selenium爬取Nguyen Kim网站长名称商品信息为空的问题
问题场景
爬取Nguyen Kim搜索页(https://www.nguyenkim.com/tim-kiem.html?tu-khoa=may+tinh)的商品信息时,短名称商品能正常返回完整的名称、价格、链接和图片,但长名称商品(如"Máy tính để bàn HP 205 Pro G4 AIO R5-4500U/8GB/256GB/Win10 31Y21PA")的名称和价格会返回空字符串,仅链接和图片正常。
正常返回示例:
Item( link = 'https://www.nguyenkim.com/chuot-logitech-m100r.html', name = 'Chuột máy tính Logitech M100R Đen', current_price = '109.000đ', place = 'Nguyen Kim', img = 'https://cdn.nguyenkimmall.com/images/thumbnails/210/210/detailed/177/10026584-chuot-logitech-m100r-den-1.jpg' )
长名称商品返回示例:
Item( link = 'https://www.nguyenkim.com/may-tinh-bang-xiaomi-redmi-pad-64gb-xam.html', name = '', current_price = '', place = 'Nguyen Kim', img = 'https://cdn.nguyenkimmall.com/images/thumbnails/210/210/detailed/847/10053972-may-tinh-bang-xiaomi-redmi-pad-64gb-xam.jpg' )
原代码片段:
content = driver.find_element(By.CLASS_NAME, 'result-wrapper') items = content.find_elements(By.CLASS_NAME, 'product') for _ in items: item = Item( link = _.find_element(By.CSS_SELECTOR, "div[class*='product-header']").get_attribute('href'), name = _.find_element(By.CSS_SELECTOR, "div.product-title a").text, current_price = _.find_element(By.CSS_SELECTOR, "p[class*='final-price']").text, place = "Nguyen Kim", img = _.find_element(By.CSS_SELECTOR, "img").get_attribute('src') ) print(item)
问题原因
网站对长名称商品采用了CSS文本截断(如overflow: hidden),导致Selenium的.text方法只能获取可视区域内的文本,而截断后的可视文本为空;同时价格元素可能因布局原因,可视文本也被隐藏或未正确渲染。此外,电商网站通常会把完整商品名称存储在链接的title属性中,而非仅依赖可视文本。
解决方案
- 商品名称:放弃使用
.text,改为获取商品链接的title属性,该属性通常存储完整未截断的商品名称。 - 商品价格:使用
get_attribute('innerText')替代.text,innerText能获取元素内的所有文本内容(包括被CSS截断的部分);若网站将价格存储在自定义属性(如data-price)中,也可直接读取该属性并格式化。
修改后的代码
content = driver.find_element(By.CLASS_NAME, 'result-wrapper') items = content.find_elements(By.CLASS_NAME, 'product') for _ in items: # 获取完整商品名称(优先取title属性,兼容短名称商品) name_elem = _.find_element(By.CSS_SELECTOR, "div.product-title a") name = name_elem.get_attribute('title') or name_elem.text # 获取价格(用innerText避免截断问题) price_elem = _.find_element(By.CSS_SELECTOR, "p[class*='final-price']") current_price = price_elem.get_attribute('innerText') or price_elem.text item = Item( link = _.find_element(By.CSS_SELECTOR, "div[class*='product-header']").get_attribute('href'), name = name.strip(), current_price = current_price.strip(), place = "Nguyen Kim", img = _.find_element(By.CSS_SELECTOR, "img").get_attribute('src') ) print(item)
补充说明
- 添加
or判断是为了兼容短名称商品(部分短名称商品的title属性可能与可视文本一致,或未设置title)。 - 若
innerText仍无法获取价格,可检查价格元素的HTML结构,例如是否存在data-price属性,此时可改为price_elem.get_attribute('data-price') + 'đ'进行格式化。
内容的提问来源于stack exchange,提问作者Dung8466
相关产品推荐
相关产品推荐

