如何使用Python Selenium点击Myntra的Show More按钮提取隐藏内容
问题原因
- 点击「See More」按钮后没有设置显式等待,AJAX加载的内容还没渲染到DOM中就执行了抓取逻辑,导致拿到空值
- 使用了超长的绝对路径Xpath定位元素,页面结构稍有变动就会定位失败
- 全局异常捕获直接
pass,无法定位是找不到元素还是元素本身没有文本内容 - 你当前点击的仅为规格参数的「See More」,如果「Complete the look」模块也有独立的展开按钮,你的代码没有覆盖对应点击逻辑
依赖导入
首先补充需要用到的等待、元素定位相关依赖:
import pandas as pd from selenium import webdriver from selenium.common.exceptions import NoSuchElementException, TimeoutException from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC
修正后代码
url = 'https://www.myntra.com/kurtas/jompers/jompers-men-yellow-printed-straight-kurta/11226756/buy' df = pd.DataFrame(columns=['name','title','price','description','Size & fit','Material & care', 'Complete the look', 'specs']) driver = webdriver.Chrome('chromedriver') # 设置10秒最长等待时间 wait = WebDriverWait(driver, 10) for _ in range(1): # 替换为实际links长度 metadata = dict.fromkeys(['name','title','price','description','Size & fit','Material & care', 'Complete the look', 'specs']) specs = dict() driver.get(url) try: # 基础信息抓取,用类名定位比绝对Xpath稳定性更高 metadata['title'] = wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'pdp-title'))).get_attribute("innerHTML") metadata['name'] = driver.find_element(By.CLASS_NAME, 'pdp-name').get_attribute("innerHTML") metadata['price'] = driver.find_element(By.CLASS_NAME, 'pdp-price').find_element(By.XPATH, './strong').get_attribute("innerHTML") metadata['description'] = driver.find_element(By.XPATH, '//div[contains(@class,"product-description")]/p').text # 点击规格参数的See More按钮 try: see_more_spec = wait.until(EC.element_to_be_clickable((By.XPATH, '//div[text()="See More" and contains(@class,"index-showMoreText")]'))) see_more_spec.click() # 等待规格参数容器加载完成 wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'index-tableContainer'))) except (NoSuchElementException, TimeoutException): pass # 抓取所有规格参数,无需循环拼接Xpath spec_rows = driver.find_elements(By.CLASS_NAME, 'index-row') for row in spec_rows: key = row.find_element(By.CLASS_NAME, 'index-rowKey').text value = row.find_element(By.CLASS_NAME, 'index-rowValue').text specs[key] = value metadata['specs'] = specs # 抓取Complete the look内容 try: # 先定位Complete the look模块 complete_look_section = wait.until(EC.presence_of_element_located((By.XPATH, '//div[contains(@class,"completeLook") or h2[text()="Complete The Look"]]'))) # 检查模块是否有独立展开按钮,有则点击 try: complete_look_more = complete_look_section.find_element(By.XPATH, './/div[text()="See More"]') complete_look_more.click() # 等待内容加载完成 wait.until(EC.visibility_of_element_located((By.XPATH, './/div[contains(@class,"look-description")]'))) except NoSuchElementException: pass # 用相对路径抓取文本 metadata['Complete the look'] = complete_look_section.find_element(By.TAG_NAME, 'p').text.strip() except (NoSuchElementException, TimeoutException): metadata['Complete the look'] = '' except Exception as e: print(f"抓取出错: {str(e)}") df = df.append(metadata, ignore_index=True) driver.quit()
核心修改说明
- 增加显式等待逻辑,确保点击展开后内容完全渲染再执行抓取
- 全部替换为相对路径+类名/文本特征定位元素,不受页面整体结构变动影响
- 拆分不同模块的异常捕获,方便定位问题
- 直接遍历规格行元素抓取所有参数,无需写死循环次数
- 增加「Complete the look」模块自身的展开按钮点击逻辑,确保隐藏内容完全可见
内容的提问来源于stack exchange,提问作者user13233820
相关产品推荐
相关产品推荐

