使用Requests爬取LG产品数据失败:无法获取名称、SKU等信息
问题描述
尝试使用requests爬取LG墨西哥官网空调页面(URL:https://www.lg.com/mx/aire-acondicionado/todos-los-aires-acondicionados/?ec_model_status_code=ACTIVE),需要获取4项数据:产品名称、SKU、原价、折扣价,但目前无法获取产品名称,出现方法调用错误。相关代码如下:
import requests from lxml import html from selenium.webdriver.common.by import By headers = { "user-agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Ubuntu Chromium/71.0.3578.80 Chrome/71.0.3578.80 Safari/537.36", } # URL SEMILLA url = 'https://www.lg.com/mx/aire-acondicionado/todos-los-aires-acondicionados/?ec_model_status_code=ACTIVE' respuesta = requests.get(url, headers=headers) print(respuesta) parser = html.fromstring(respuesta.content) lavadora = parser.find_element(By.CLASS_NAME,"c-product-item__ufn") print(lavadora)
问题分析与解决方案
核心问题
- 库方法混用错误:
lxml.html.fromstring()返回的解析对象没有find_element()方法,这个方法是Selenium WebDriver的专属方法,不能直接用于lxml解析后的HTML对象。 - 页面动态渲染问题:LG官网的产品列表大概率通过JavaScript动态加载,
requests仅能获取静态HTML源码,可能不包含目标产品元素。
修正方案
方案一:改用lxml原生选择器解析静态内容(若静态源码含目标元素)
如果页面静态源码中存在目标元素,可替换为lxml支持的CSS选择器或XPath方式:
import requests from lxml import html headers = { "user-agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Ubuntu Chromium/71.0.3578.80 Chrome/71.0.3578.80 Safari/537.36", } url = 'https://www.lg.com/mx/aire-acondicionado/todos-los-aires-acondicionados/?ec_model_status_code=ACTIVE' respuesta = requests.get(url, headers=headers) parser = html.fromstring(respuesta.content) # 使用CSS选择器批量获取产品名称 product_names = parser.cssselect(".c-product-item__ufn") for name in product_names: print(name.text_content().strip())
方案二:使用Selenium处理动态加载页面
若静态源码中无目标元素,必须用Selenium模拟浏览器加载动态内容:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options import time options = Options() options.add_argument("user-agent=Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Ubuntu Chromium/71.0.3578.80 Chrome/71.0.3578.80 Safari/537.36") options.add_argument("--headless=new") # 无头模式,不显示浏览器窗口 driver = webdriver.Chrome(options=options) driver.get('https://www.lg.com/mx/aire-acondicionado/todos-los-aires-acondicionados/?ec_model_status_code=ACTIVE') time.sleep(3) # 等待页面动态内容加载完成 # 获取产品名称 product_names = driver.find_elements(By.CLASS_NAME, "c-product-item__ufn") for name in product_names: print(name.text.strip()) # 可扩展获取其他数据,需替换为对应元素类名 # skus = driver.find_elements(By.CLASS_NAME, "目标SKU类名") # original_prices = driver.find_elements(By.CLASS_NAME, "目标原价类名") # discounted_prices = driver.find_elements(By.CLASS_NAME, "目标折扣价类名") driver.quit()
补充说明
- 使用方案二时,需提前安装Selenium库及对应浏览器驱动(如ChromeDriver)。
- 可通过浏览器开发者工具检查目标元素的实际类名,避免类名拼写错误。
内容的提问来源于stack exchange,提问作者Roberto Cravioto
相关产品推荐
相关产品推荐

