You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python Selenium点击Myntra的Show More按钮提取隐藏内容

问题原因

  • 点击「See More」按钮后没有设置显式等待,AJAX加载的内容还没渲染到DOM中就执行了抓取逻辑,导致拿到空值
  • 使用了超长的绝对路径Xpath定位元素,页面结构稍有变动就会定位失败
  • 全局异常捕获直接pass,无法定位是找不到元素还是元素本身没有文本内容
  • 你当前点击的仅为规格参数的「See More」,如果「Complete the look」模块也有独立的展开按钮,你的代码没有覆盖对应点击逻辑

依赖导入

首先补充需要用到的等待、元素定位相关依赖:

import pandas as pd
from selenium import webdriver
from selenium.common.exceptions import NoSuchElementException, TimeoutException
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

修正后代码

url = 'https://www.myntra.com/kurtas/jompers/jompers-men-yellow-printed-straight-kurta/11226756/buy'
df = pd.DataFrame(columns=['name','title','price','description','Size & fit','Material & care', 'Complete the look', 'specs'])
driver = webdriver.Chrome('chromedriver')
# 设置10秒最长等待时间
wait = WebDriverWait(driver, 10)

for _ in range(1): # 替换为实际links长度
    metadata = dict.fromkeys(['name','title','price','description','Size & fit','Material & care', 'Complete the look', 'specs'])
    specs = dict()
    driver.get(url)
    try:
        # 基础信息抓取,用类名定位比绝对Xpath稳定性更高
        metadata['title'] = wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'pdp-title'))).get_attribute("innerHTML")
        metadata['name'] = driver.find_element(By.CLASS_NAME, 'pdp-name').get_attribute("innerHTML")
        metadata['price'] = driver.find_element(By.CLASS_NAME, 'pdp-price').find_element(By.XPATH, './strong').get_attribute("innerHTML")
        metadata['description'] = driver.find_element(By.XPATH, '//div[contains(@class,"product-description")]/p').text

        # 点击规格参数的See More按钮
        try:
            see_more_spec = wait.until(EC.element_to_be_clickable((By.XPATH, '//div[text()="See More" and contains(@class,"index-showMoreText")]')))
            see_more_spec.click()
            # 等待规格参数容器加载完成
            wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'index-tableContainer')))
        except (NoSuchElementException, TimeoutException):
            pass
        # 抓取所有规格参数,无需循环拼接Xpath
        spec_rows = driver.find_elements(By.CLASS_NAME, 'index-row')
        for row in spec_rows:
            key = row.find_element(By.CLASS_NAME, 'index-rowKey').text
            value = row.find_element(By.CLASS_NAME, 'index-rowValue').text
            specs[key] = value
        metadata['specs'] = specs

        # 抓取Complete the look内容
        try:
            # 先定位Complete the look模块
            complete_look_section = wait.until(EC.presence_of_element_located((By.XPATH, '//div[contains(@class,"completeLook") or h2[text()="Complete The Look"]]')))
            # 检查模块是否有独立展开按钮,有则点击
            try:
                complete_look_more = complete_look_section.find_element(By.XPATH, './/div[text()="See More"]')
                complete_look_more.click()
                # 等待内容加载完成
                wait.until(EC.visibility_of_element_located((By.XPATH, './/div[contains(@class,"look-description")]')))
            except NoSuchElementException:
                pass
            # 用相对路径抓取文本
            metadata['Complete the look'] = complete_look_section.find_element(By.TAG_NAME, 'p').text.strip()
        except (NoSuchElementException, TimeoutException):
            metadata['Complete the look'] = ''

    except Exception as e:
        print(f"抓取出错: {str(e)}")
        
    df = df.append(metadata, ignore_index=True)

driver.quit()

核心修改说明

  • 增加显式等待逻辑,确保点击展开后内容完全渲染再执行抓取
  • 全部替换为相对路径+类名/文本特征定位元素,不受页面整体结构变动影响
  • 拆分不同模块的异常捕获,方便定位问题
  • 直接遍历规格行元素抓取所有参数,无需写死循环次数
  • 增加「Complete the look」模块自身的展开按钮点击逻辑,确保隐藏内容完全可见

内容的提问来源于stack exchange,提问作者user13233820

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 15:15:00